For AI agents: a documentation index is available at /llms.txt. Markdown versions of all pages can be requested by appending `.md` to the URL, or by setting the `Accept` header to `text/markdown`.
Skip to main content
Integrations and SDKsPipecat

Pipecat speech to text

Transcribe live audio in your Pipecat voice bots with the Speechmatics agent STT service.

SpeechmaticsSTTService transcribes live audio using agent STT, the turn-based interaction pattern on the Realtime API built for applications that hold a conversation. It runs on the Linden 1 model.

Features​

  • Real-time transcription — transcript segments stream as the speaker talks, with optional partial results
  • Turn detection — Pipecat's own voice activity detector (VAD) closes each turn, or Speechmatics closes it server-side
  • Speaker diarization — attribute each segment to a speaker
  • Stable speaker labels — carry the same labels across sessions
  • Custom vocabulary — bias recognition towards domain-specific words
  • Speaker formatting — template speaker labels for your language model
  • Runtime updates — change most settings mid-conversation

Installation​

Install Pipecat with the Speechmatics extra:

uv add "pipecat-ai[speechmatics]~=1.11.0"

Authentication​

The service reads your API key from the SPEECHMATICS_API_KEY environment variable.

.env
# Or pass api_key= to SpeechmaticsSTTService() directly
SPEECHMATICS_API_KEY=your_speechmatics_key

Endpoint​

The service connects to wss://eu2.rt.speechmatics.com/v2/agent. Set base_url, or the SPEECHMATICS_RT_URL environment variable, to reach a different region or a self-hosted deployment. Agent STT sessions are served on the /v2/agent path, and the SDK appends /agent when your URL omits it.

stt = SpeechmaticsSTTService(
# The /v2/agent path is required; /v2 is the streaming endpoint
base_url="wss://eu2.rt.speechmatics.com/v2/agent",
)

For the hostname of each region, see supported endpoints.

Basic usage​

Place the service after transport.input() in your pipeline. The defaults transcribe English on Linden 1.

from pipecat.services.speechmatics.stt import SpeechmaticsSTTService
from pipecat.transcriptions.language import Language

stt = SpeechmaticsSTTService(
settings=SpeechmaticsSTTService.Settings(language=Language.EN),
)

# The service belongs directly after the transport input
pipeline = Pipeline([transport.input(), stt, user_aggregator, llm, tts, transport.output()])

Constructor parameters​

SpeechmaticsSTTService accepts the following parameters. Everything else is a setting.

ParameterTypeDefaultDescription
api_keystringenv varSpeechmatics API key. Falls back to SPEECHMATICS_API_KEY
base_urlstringwss://eu2.rt.speechmatics.com/v2/agentEndpoint to connect to. Falls back to SPEECHMATICS_RT_URL
sample_ratenumber | nullpipeline defaultAudio sample rate in Hz. Agent STT requires 16000
encodingAudioEncodingPCM_S16LEAudio encoding. Set at construction only, not a runtime setting
settingsSettings | nullnullRuntime-configurable settings
should_interruptbooleantrueInterrupt bot output when Speechmatics detects user speech. Applies to VAD turn detection only
ttfs_p99_latencynumber0.74P99 latency in seconds from speech end to final transcript. Override for your deployment

Transcription settings​

Settings are passed as settings=SpeechmaticsSTTService.Settings(...) and can be changed mid-conversation with an STTUpdateSettingsFrame.

SettingTypeDefaultDescription
modelstring"linden-1"Transcription model. linden-1 is the only model agent STT accepts
languageLanguage | stringLanguage.ENLanguage code for the input audio. See Languages
domainstring | nullnullDomain-specific language pack, for example "finance", or a bilingual pack such as "bilingual-en"
enable_partialsboolean | nulltrueEmit interim segments that update until the final segment arrives
additional_vocabarray[]AdditionalVocabEntry objects biasing recognition towards specific words

Every language is global: whichever you select, the model recognizes a range of accents and dialects.

A setting you leave unset is not sent, and agent STT applies its own default rather than the one in the Realtime API schema. Partials and diarization are both on unless you turn them off.

Custom vocabulary​

Pass additional_vocab entries to boost recognition of names and domain terms. Each entry takes a content value and optional sounds_like pronunciation variants.

stt = SpeechmaticsSTTService(
settings=SpeechmaticsSTTService.Settings(
additional_vocab=[
# sounds_like helps when the spelling is unusual
SpeechmaticsSTTService.AdditionalVocabEntry(
content="Pipecat", sounds_like=["pipe cat"]
),
],
),
)

Turn detection​

Agent STT reports turn boundaries as StartOfTurn and EndOfTurn events. The turn_detection_mode setting sets which component decides where those boundaries fall.

Turn detection modes​

Two modes are available.

ModeDetection method
TurnDetectionMode.EXTERNALDefault. Pipecat closes the turn. A VADUserStoppedSpeakingFrame calls finalize() on the service
TurnDetectionMode.VADSpeechmatics closes the turn server-side using its own voice activity detection

The two modes fail in opposite directions. EXTERNAL returns the final transcript sooner and with a tighter tail, but can clip a word at a turn boundary and drops short backchannels such as "mm-hmm" that never cross your VAD's speech threshold. VAD is slightly more accurate and keeps backchannels, because Speechmatics applies its own, more patient endpointing.

External turn detection​

EXTERNAL is the default. Pipecat's VAD drives it, so put a vad_analyzer on the transport, or a vad_analyzer on the user aggregator. Either position works: the VADUserStoppedSpeakingFrame is broadcast both upstream and downstream, so it reaches the service even from a VAD that sits after it in the pipeline.

transport_params = TransportParams(
audio_in_enabled=True,
audio_out_enabled=True,
# Signals the end of each turn to the STT service
vad_analyzer=SileroVADAnalyzer(),
)

In EXTERNAL mode something must close each turn. With no VAD anywhere in the pipeline, nothing calls finalize() and the service produces no transcripts at all rather than failing.

Server-side turn detection​

Setting turn_detection_mode to VAD hands turn handling to Speechmatics, which runs server-side VAD to drive turn detection. The service requests ExternalUserTurnStrategies when it starts, proposes each turn boundary, and lets those strategies resolve it, so should_interrupt becomes the control for barge-in.

stt = SpeechmaticsSTTService(
settings=SpeechmaticsSTTService.Settings(
# Speechmatics endpoints server-side; no Pipecat VAD needed
turn_detection_mode=SpeechmaticsSTTService.TurnDetectionMode.VAD,
),
)

In VAD mode, remove any competing vad_analyzer or turn_analyzer from your transport. A transport VAD is still useful if you want STT metrics, but it must not drive turns.

End-of-utterance timing in VAD mode belongs to the service and is not configurable.

Speaker diarization​

Diarization (attributing speech to individual speakers) is enabled by default. It labels each transcript segment with one speaker, such as S1. The label arrives on the TranscriptionFrame as user_id. Segments the model cannot attribute are labelled UU, and with diarization disabled every segment is labelled UU.

Agent STT reports one speaker per segment rather than per word, so there is no per-word speaker data.

Diarization settings​

SettingTypeDefaultDescription
enable_diarizationboolean | nulltrueAttribute each segment to a speaker
speaker_sensitivitynumber | nullnullSpeaker detection sensitivity, between 0 and 1. Higher values help separate similar voices
max_speakersnumber | nullnullMaximum speakers to detect, between 2 and 100. Unrestricted when unset
prefer_current_speakerboolean | nullnullReduce switching between similar-sounding speakers
known_speakersarray[]SpeakerIdentifier objects from an earlier session
speaker_active_formatstring"{text}"Template for speaker output

Format speaker labels​

Set speaker_active_format to fold the speaker label into the transcript text your language model receives.

stt = SpeechmaticsSTTService(
settings=SpeechmaticsSTTService.Settings(
enable_diarization=True,
# Renders as <S1>Good morning.</S1>
speaker_active_format="<{speaker_id}>{text}</{speaker_id}>",
),
)

Include the format in your bot's system prompt so the model interprets the labels consistently. The template accepts the following attributes.

AttributeTypeDescriptionExample
{speaker_id}stringLabel of the speakerS1
{text}stringTranscribed textGood morning.
{ts}floatStart time of the segment, in seconds from the start of the session12.34
{lang}stringLanguage of the transcriptionen

Reuse speaker labels across sessions​

Speechmatics assigns speaker labels per session, so the same person is S1 in one session and something else in the next. Handle the on_speakers_result event to collect identifiers, then pass them as known_speakers when you build the service for the next session.

@stt.event_handler("on_speakers_result")
async def on_speakers_result(service, speakers):
# Persist these to reuse the same labels next session
print(f"Speaker result: {speakers}")

Speaker identifiers are unique to your Speechmatics account.

Connection handling​

The service manages its own WebSocket connection. Transient failures trigger reconnection with exponential backoff while audio is buffered. Permanent errors, such as a rejected API key or a configuration agent STT refuses, mark the service unusable rather than retrying.

Deprecated and unsupported parameters​

The params=SpeechmaticsSTTService.InputParams(...) pattern is deprecated as of Pipecat 0.0.105. Use settings=SpeechmaticsSTTService.Settings(...) instead, which is also what STTUpdateSettingsFrame updates at runtime.

operating_point is a deprecated alias for model and is removed in Pipecat 2.0.0. Passing both raises a ValueError unless they name the same value.

The following are no longer settings. Passing any of them to Settings(...) raises a TypeError. Passing them as a constructor argument is silently ignored, except where noted.

ParameterReason
max_delayTranscript timing is the service's own. Warns as a constructor argument
end_of_utterance_silence_triggerEnd-of-utterance timing is the service's own. Warns as a constructor argument
end_of_utterance_max_delayEnd-of-utterance timing is the service's own
split_sentencesNot exposed by agent STT
focus_speakers, ignore_speakers, focus_modeAgent STT does not support speaker focus
speaker_passive_formatA segment has no passive case to format
extra_paramsAgent STT rejects configuration fields outside its fixed schema

The ADAPTIVE, FIXED, and SMART_TURN turn detection modes are gone. Each selected a server-side endpointing strategy that agent STT does not expose. Use EXTERNAL or VAD, as described in Turn detection modes.

The update_params() method no longer changes speaker focus. Use an STTUpdateSettingsFrame to change settings at runtime.

Full example​

This example combines transcription, turn detection, diarization, and speaker formatting in one service.

from pipecat.services.speechmatics.stt import SpeechmaticsSTTService
from pipecat.transcriptions.language import Language

stt = SpeechmaticsSTTService(
settings=SpeechmaticsSTTService.Settings(
# Transcription
language=Language.EN,
enable_partials=True,

# Turn detection: a Pipecat VAD closes each turn
turn_detection_mode=SpeechmaticsSTTService.TurnDetectionMode.EXTERNAL,

# Diarization
enable_diarization=True,
speaker_sensitivity=0.6,
max_speakers=4,
prefer_current_speaker=True,

# Renders each segment as <S1>Good morning.</S1>
speaker_active_format="<{speaker_id}>{text}</{speaker_id}>",

# Custom vocabulary
additional_vocab=[
SpeechmaticsSTTService.AdditionalVocabEntry(content="Speechmatics"),
SpeechmaticsSTTService.AdditionalVocabEntry(
content="Pipecat", sounds_like=["pipe cat"]
),
],
),
)

Next steps​