# Overview

Turn-based transcription for building voice agents.

Agent STT is transcription designed for building voice agents. Stream audio in and receive speaker-labelled, turn-based transcription back — punctuated and ready to pass to an LLM.

Agent STT is available on SaaS only.

There are three ways to use it:

#### [Integrations](/integrations-and-sdks/.md)

[The fastest path to a production voice agent, on LiveKit, Pipecat or Vapi](/integrations-and-sdks/.md)

#### [Agent STT SDK](/speech-to-text/agent-stt/quickstart.md)

[A Python client that handles the connection, audio streaming and message handling](/speech-to-text/agent-stt/quickstart.md)

#### [Agent STT API](/api-ref/agent-stt-websocket.md)

[Connect over WebSocket from any language, for custom pipelines or unsupported platforms](/api-ref/agent-stt-websocket.md)

## Features[​](#features "Direct link to Features")

* **Turn detection**: detect when a speaker has finished talking. See [Turn detection](/speech-to-text/agent-stt/turn-detection.md).
* **Intelligent segmentation**: group words into clean, speaker-attributed segments. See [Segmentation](/speech-to-text/agent-stt/segmentation.md).
* **Diarization**: identify and label different speakers.
* **Speech signals**: receive signals for start and end of speech.

## Integrations[​](#integrations "Direct link to Integrations")

Use an integration to handle audio transport and wiring, so you can focus on your agent logic:

[![Vapi logo](/img/integration-logos/vapi.png)](/integrations-and-sdks/vapi.md)

#### [Vapi](/integrations-and-sdks/vapi.md)

[Turnkey voice agent platform. Deploy fast with no code.](/integrations-and-sdks/vapi.md)

[![LiveKit logo](/img/integration-logos/livekit.png)](/integrations-and-sdks/livekit/.md)

#### [LiveKit](/integrations-and-sdks/livekit/.md)

[Open-source framework for building agents with WebRTC infrastructure.](/integrations-and-sdks/livekit/.md)

[![Pipecat logo](/img/integration-logos/pipecat.png)](/integrations-and-sdks/pipecat/.md)

#### [Pipecat](/integrations-and-sdks/pipecat/.md)

[Open-source framework with full control of the voice pipeline in code.](/integrations-and-sdks/pipecat/.md)

## Agent STT SDK[​](#agent-stt-sdk "Direct link to Agent STT SDK")

The Agent STT SDK is a Python client for the Agent STT service. It manages the WebSocket session, streams your audio, and delivers segments and turn events to your callbacks.

See the [Quickstart](/speech-to-text/agent-stt/quickstart.md) to get started.

## Using the API directly[​](#using-the-api-directly "Direct link to Using the API directly")

Connect to the WebSocket directly from any language when there's no SDK for your stack, or when you need control over the session that a client library doesn't expose.

```
wss://global.rt.speechmatics.com/v2/agent
```

Use the global endpoint to route to the nearest region, or pin to a region for data residency. See [Supported endpoints](/get-started/authentication.md#supported-endpoints) for regional endpoints.

Open the connection with your API key, send `StartRecognition` with your audio format and transcription config, then stream binary audio in `AddAudio` messages. Segments and turn events arrive as JSON while you send.

See the [Agent STT Reference](/api-ref/agent-stt-websocket.md) for the session flow, every message schema and the error codes.
