# Pipecat speech to text

Transcribe live audio in your Pipecat voice bots with the Speechmatics agent STT service.

`SpeechmaticsSTTService` transcribes live audio using agent STT, the turn-based interaction pattern on the Realtime API built for applications that hold a conversation. It runs on the Linden 1 model.

## Features[​](#features "Direct link to Features")

* **Real-time transcription** — transcript segments stream as the speaker talks, with optional partial results
* **Turn detection** — Pipecat's own voice activity detector (VAD) closes each turn, or Speechmatics closes it server-side
* **Speaker diarization** — attribute each segment to a speaker
* **Stable speaker labels** — carry the same labels across sessions
* **Custom vocabulary** — bias recognition towards domain-specific words
* **Speaker formatting** — template speaker labels for your language model
* **Runtime updates** — change most settings mid-conversation

## Installation[​](#installation "Direct link to Installation")

Install Pipecat with the Speechmatics extra:

```
uv add "pipecat-ai[speechmatics]~=1.11.0"
```

## Authentication[​](#authentication "Direct link to Authentication")

The service reads your [API key](/get-started/authentication.md) from the `SPEECHMATICS_API_KEY` environment variable.

.env

```
# Or pass api_key= to SpeechmaticsSTTService() directly
SPEECHMATICS_API_KEY=your_speechmatics_key
```

## Endpoint[​](#endpoint "Direct link to Endpoint")

The service connects to `wss://eu2.rt.speechmatics.com/v2/agent`. Set `base_url`, or the `SPEECHMATICS_RT_URL` environment variable, to reach a different region or a self-hosted deployment. Agent STT sessions are served on the `/v2/agent` path, and the SDK appends `/agent` when your URL omits it.

```
stt = SpeechmaticsSTTService(
    # The /v2/agent path is required; /v2 is the streaming endpoint
    base_url="wss://eu2.rt.speechmatics.com/v2/agent",
)
```

For the hostname of each region, see [supported endpoints](/get-started/authentication.md#supported-endpoints).

## Basic usage[​](#basic-usage "Direct link to Basic usage")

Place the service after `transport.input()` in your pipeline. The defaults transcribe English on Linden 1.

```
from pipecat.services.speechmatics.stt import SpeechmaticsSTTService
from pipecat.transcriptions.language import Language

stt = SpeechmaticsSTTService(
    settings=SpeechmaticsSTTService.Settings(language=Language.EN),
)

# The service belongs directly after the transport input
pipeline = Pipeline([transport.input(), stt, user_aggregator, llm, tts, transport.output()])
```

## Constructor parameters[​](#constructor-parameters "Direct link to Constructor parameters")

`SpeechmaticsSTTService` accepts the following parameters. Everything else is a [setting](#transcription-settings).

| Parameter          | Type             | Default                                  | Description                                                                                      |
| ------------------ | ---------------- | ---------------------------------------- | ------------------------------------------------------------------------------------------------ |
| `api_key`          | string           | env var                                  | Speechmatics API key. Falls back to `SPEECHMATICS_API_KEY`                                       |
| `base_url`         | string           | `wss://eu2.rt.speechmatics.com/v2/agent` | Endpoint to connect to. Falls back to `SPEECHMATICS_RT_URL`                                      |
| `sample_rate`      | number \| null   | pipeline default                         | Audio sample rate in Hz. Agent STT requires `16000`                                              |
| `encoding`         | AudioEncoding    | `PCM_S16LE`                              | Audio encoding. Set at construction only, not a runtime setting                                  |
| `settings`         | Settings \| null | `null`                                   | Runtime-configurable settings                                                                    |
| `should_interrupt` | boolean          | `true`                                   | Interrupt bot output when Speechmatics detects user speech. Applies to `VAD` turn detection only |
| `ttfs_p99_latency` | number           | `0.74`                                   | P99 latency in seconds from speech end to final transcript. Override for your deployment         |

## Transcription settings[​](#transcription-settings "Direct link to Transcription settings")

Settings are passed as `settings=SpeechmaticsSTTService.Settings(...)` and can be changed mid-conversation with an `STTUpdateSettingsFrame`.

| Setting            | Type               | Default       | Description                                                                                          |
| ------------------ | ------------------ | ------------- | ---------------------------------------------------------------------------------------------------- |
| `model`            | string             | `"linden-1"`  | Transcription model. `linden-1` is the only model agent STT accepts                                  |
| `language`         | Language \| string | `Language.EN` | Language code for the input audio. See [Languages](/speech-to-text/languages.md)                     |
| `domain`           | string \| null     | `null`        | Domain-specific language pack, for example `"finance"`, or a bilingual pack such as `"bilingual-en"` |
| `enable_partials`  | boolean \| null    | `true`        | Emit interim segments that update until the final segment arrives                                    |
| `additional_vocab` | array              | `[]`          | `AdditionalVocabEntry` objects biasing recognition towards specific words                            |

Every language is global: whichever you select, the model recognizes a range of accents and dialects.

A setting you leave unset is not sent, and agent STT applies its own default rather than the one in the Realtime API schema. Partials and diarization are both on unless you turn them off.

### Custom vocabulary[​](#custom-vocabulary "Direct link to Custom vocabulary")

Pass `additional_vocab` entries to boost recognition of names and domain terms. Each entry takes a `content` value and optional `sounds_like` pronunciation variants.

```
stt = SpeechmaticsSTTService(
    settings=SpeechmaticsSTTService.Settings(
        additional_vocab=[
            # sounds_like helps when the spelling is unusual
            SpeechmaticsSTTService.AdditionalVocabEntry(
                content="Pipecat", sounds_like=["pipe cat"]
            ),
        ],
    ),
)
```

## Turn detection[​](#turn-detection "Direct link to Turn detection")

Agent STT reports turn boundaries as `StartOfTurn` and `EndOfTurn` events. The `turn_detection_mode` setting sets which component decides where those boundaries fall.

### Turn detection modes[​](#turn-detection-modes "Direct link to Turn detection modes")

Two modes are available.

| Mode                         | Detection method                                                                                    |
| ---------------------------- | --------------------------------------------------------------------------------------------------- |
| `TurnDetectionMode.EXTERNAL` | Default. Pipecat closes the turn. A `VADUserStoppedSpeakingFrame` calls `finalize()` on the service |
| `TurnDetectionMode.VAD`      | Speechmatics closes the turn server-side using its own voice activity detection                     |

The two modes fail in opposite directions. `EXTERNAL` returns the final transcript sooner and with a tighter tail, but can clip a word at a turn boundary and drops short backchannels such as "mm-hmm" that never cross your VAD's speech threshold. `VAD` is slightly more accurate and keeps backchannels, because Speechmatics applies its own, more patient endpointing.

### External turn detection[​](#external-turn-detection "Direct link to External turn detection")

`EXTERNAL` is the default. Pipecat's VAD drives it, so put a `vad_analyzer` on the transport, or a `vad_analyzer` on the user aggregator. Either position works: the `VADUserStoppedSpeakingFrame` is broadcast both upstream and downstream, so it reaches the service even from a VAD that sits after it in the pipeline.

```
transport_params = TransportParams(
    audio_in_enabled=True,
    audio_out_enabled=True,
    # Signals the end of each turn to the STT service
    vad_analyzer=SileroVADAnalyzer(),
)
```

In `EXTERNAL` mode something must close each turn. With no VAD anywhere in the pipeline, nothing calls `finalize()` and the service produces no transcripts at all rather than failing.

### Server-side turn detection[​](#server-side-turn-detection "Direct link to Server-side turn detection")

Setting `turn_detection_mode` to `VAD` hands turn handling to Speechmatics, which runs server-side VAD to drive turn detection. The service requests `ExternalUserTurnStrategies` when it starts, proposes each turn boundary, and lets those strategies resolve it, so `should_interrupt` becomes the control for barge-in.

```
stt = SpeechmaticsSTTService(
    settings=SpeechmaticsSTTService.Settings(
        # Speechmatics endpoints server-side; no Pipecat VAD needed
        turn_detection_mode=SpeechmaticsSTTService.TurnDetectionMode.VAD,
    ),
)
```

In `VAD` mode, remove any competing `vad_analyzer` or `turn_analyzer` from your transport. A transport VAD is still useful if you want STT metrics, but it must not drive turns.

End-of-utterance timing in `VAD` mode belongs to the service and is not configurable.

## Speaker diarization[​](#speaker-diarization "Direct link to Speaker diarization")

[Diarization](/speech-to-text/features/diarization.md) (attributing speech to individual speakers) is enabled by default. It labels each transcript segment with one speaker, such as `S1`. The label arrives on the `TranscriptionFrame` as `user_id`. Segments the model cannot attribute are labelled `UU`, and with diarization disabled every segment is labelled `UU`.

Agent STT reports one speaker per segment rather than per word, so there is no per-word speaker data.

### Diarization settings[​](#diarization-settings "Direct link to Diarization settings")

| Setting                  | Type            | Default    | Description                                                                                    |
| ------------------------ | --------------- | ---------- | ---------------------------------------------------------------------------------------------- |
| `enable_diarization`     | boolean \| null | `true`     | Attribute each segment to a speaker                                                            |
| `speaker_sensitivity`    | number \| null  | `null`     | Speaker detection sensitivity, between `0` and `1`. Higher values help separate similar voices |
| `max_speakers`           | number \| null  | `null`     | Maximum speakers to detect, between `2` and `100`. Unrestricted when unset                     |
| `prefer_current_speaker` | boolean \| null | `null`     | Reduce switching between similar-sounding speakers                                             |
| `known_speakers`         | array           | `[]`       | `SpeakerIdentifier` objects from an earlier session                                            |
| `speaker_active_format`  | string          | `"{text}"` | Template for speaker output                                                                    |

### Format speaker labels[​](#format-speaker-labels "Direct link to Format speaker labels")

Set `speaker_active_format` to fold the speaker label into the transcript text your language model receives.

```
stt = SpeechmaticsSTTService(
    settings=SpeechmaticsSTTService.Settings(
        enable_diarization=True,
        # Renders as <S1>Good morning.</S1>
        speaker_active_format="<{speaker_id}>{text}</{speaker_id}>",
    ),
)
```

Include the format in your bot's system prompt so the model interprets the labels consistently. The template accepts the following attributes.

| Attribute      | Type   | Description                                                         | Example         |
| -------------- | ------ | ------------------------------------------------------------------- | --------------- |
| `{speaker_id}` | string | Label of the speaker                                                | `S1`            |
| `{text}`       | string | Transcribed text                                                    | `Good morning.` |
| `{ts}`         | float  | Start time of the segment, in seconds from the start of the session | `12.34`         |
| `{lang}`       | string | Language of the transcription                                       | `en`            |

### Reuse speaker labels across sessions[​](#reuse-speaker-labels-across-sessions "Direct link to Reuse speaker labels across sessions")

Speechmatics assigns speaker labels per session, so the same person is `S1` in one session and something else in the next. Handle the `on_speakers_result` event to collect identifiers, then pass them as `known_speakers` when you build the service for the next session.

```
@stt.event_handler("on_speakers_result")
async def on_speakers_result(service, speakers):
    # Persist these to reuse the same labels next session
    print(f"Speaker result: {speakers}")
```

Speaker identifiers are unique to your Speechmatics account.

## Connection handling[​](#connection-handling "Direct link to Connection handling")

The service manages its own WebSocket connection. Transient failures trigger reconnection with exponential backoff while audio is buffered. Permanent errors, such as a rejected API key or a configuration agent STT refuses, mark the service unusable rather than retrying.

## Deprecated and unsupported parameters[​](#deprecated-and-unsupported-parameters "Direct link to Deprecated and unsupported parameters")

The `params=SpeechmaticsSTTService.InputParams(...)` pattern is deprecated as of Pipecat 0.0.105. Use `settings=SpeechmaticsSTTService.Settings(...)` instead, which is also what `STTUpdateSettingsFrame` updates at runtime.

`operating_point` is a deprecated alias for `model` and is removed in Pipecat 2.0.0. Passing both raises a `ValueError` unless they name the same value.

The following are no longer settings. Passing any of them to `Settings(...)` raises a `TypeError`. Passing them as a constructor argument is silently ignored, except where noted.

| Parameter                                         | Reason                                                                        |
| ------------------------------------------------- | ----------------------------------------------------------------------------- |
| `max_delay`                                       | Transcript timing is the service's own. Warns as a constructor argument       |
| `end_of_utterance_silence_trigger`                | End-of-utterance timing is the service's own. Warns as a constructor argument |
| `end_of_utterance_max_delay`                      | End-of-utterance timing is the service's own                                  |
| `split_sentences`                                 | Not exposed by agent STT                                                      |
| `focus_speakers`, `ignore_speakers`, `focus_mode` | Agent STT does not support speaker focus                                      |
| `speaker_passive_format`                          | A segment has no passive case to format                                       |
| `extra_params`                                    | Agent STT rejects configuration fields outside its fixed schema               |

The `ADAPTIVE`, `FIXED`, and `SMART_TURN` turn detection modes are gone. Each selected a server-side endpointing strategy that agent STT does not expose. Use `EXTERNAL` or `VAD`, as described in [Turn detection modes](#turn-detection-modes).

The `update_params()` method no longer changes speaker focus. Use an `STTUpdateSettingsFrame` to change settings at runtime.

## Full example[​](#full-example "Direct link to Full example")

This example combines transcription, turn detection, diarization, and speaker formatting in one service.

```
from pipecat.services.speechmatics.stt import SpeechmaticsSTTService
from pipecat.transcriptions.language import Language

stt = SpeechmaticsSTTService(
    settings=SpeechmaticsSTTService.Settings(
        # Transcription
        language=Language.EN,
        enable_partials=True,

        # Turn detection: a Pipecat VAD closes each turn
        turn_detection_mode=SpeechmaticsSTTService.TurnDetectionMode.EXTERNAL,

        # Diarization
        enable_diarization=True,
        speaker_sensitivity=0.6,
        max_speakers=4,
        prefer_current_speaker=True,

        # Renders each segment as <S1>Good morning.</S1>
        speaker_active_format="<{speaker_id}>{text}</{speaker_id}>",

        # Custom vocabulary
        additional_vocab=[
            SpeechmaticsSTTService.AdditionalVocabEntry(content="Speechmatics"),
            SpeechmaticsSTTService.AdditionalVocabEntry(
                content="Pipecat", sounds_like=["pipe cat"]
            ),
        ],
    ),
)
```

## Next steps[​](#next-steps "Direct link to Next steps")

* [Quickstart](/integrations-and-sdks/pipecat/.md) — build a complete bot with Pipecat
* [Text to speech](/integrations-and-sdks/pipecat/tts.md) — add Speechmatics voices
* [Diarization](/speech-to-text/features/diarization.md) — tune speaker separation
* [Pipecat documentation](https://docs.pipecat.ai/server/services/stt/speechmatics) — full Speechmatics STT reference
