# Models

Compare the Speechmatics models for Batch, Realtime, and Agent STT, and choose the right one for your audio.

Speechmatics offers models in two groups. The Batch and Realtime APIs share a set of models you choose between with the `model` property. The [Agent STT API](/speech-to-text/agent-stt/.md) has its own models, which are not available on Batch or Realtime.

| API       | Models                                     |
| --------- | ------------------------------------------ |
| Batch     | `enhanced`, `standard`, `melia-1`, `oak-1` |
| Realtime  | `enhanced`, `standard`                     |
| Agent STT | `linden-1`                                 |

## Batch and Realtime models[​](#batch-and-realtime-models "Direct link to Batch and Realtime models")

### Compare the models[​](#compare-the-models "Direct link to Compare the models")

| Capability             | Enhanced                                                                                                                              | Standard                                                                                                                              | Melia 1                            | Oak 1                            |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------- | -------------------------------- |
| Model                  | `enhanced`                                                                                                                            | `standard`                                                                                                                            | `melia-1`                          | `oak-1`                          |
| Medical specialization | Medical variant available                                                                                                             | —                                                                                                                                     | —                                  | Medical native model             |
| Accuracy               | Highest                                                                                                                               | High                                                                                                                                  | High                               | High                             |
| Turnaround             | Fast                                                                                                                                  | Fastest                                                                                                                               | Fastest                            | Fastest                          |
| Processing modes       | Batch and Realtime                                                                                                                    | Batch and Realtime                                                                                                                    | Batch                              | Batch                            |
| Regions                | EU, US, AUS                                                                                                                           | EU, US, AUS                                                                                                                           | EU, US                             | EU, US                           |
| Language handling      | [Selected language or pack](/speech-to-text/languages.md) ([auto-detect](/speech-to-text/batch/language-identification.md) available) | [Selected language or pack](/speech-to-text/languages.md) ([auto-detect](/speech-to-text/batch/language-identification.md) available) | [Automatic multilingual](#melia-1) | [Automatic multilingual](#oak-1) |
| Diarization            | Speaker and channel                                                                                                                   | Speaker and channel                                                                                                                   | Speaker and channel                | Speaker and channel              |
| Language labeling      | —                                                                                                                                     | —                                                                                                                                     | Per word                           | Per word                         |
| Custom dictionary      | ✅                                                                                                                                    | ✅                                                                                                                                    | Not yet                            | Not yet                          |
| Confidence scores      | ✅                                                                                                                                    | ✅                                                                                                                                    | Not yet                            | Not yet                          |
| Speaker identification | ✅                                                                                                                                    | ✅                                                                                                                                    | Not yet                            | Not yet                          |
| Speech intelligence    | ✅                                                                                                                                    | ✅                                                                                                                                    | Not yet                            | Not yet                          |

Enhanced and Standard are feature-identical and differ only in accuracy and speed: Enhanced delivers the highest accuracy, and Standard prioritizes throughput. Melia 1 and Oak 1 add automatic multilingual transcription, but are available for Batch only and support a reduced feature set. Speech intelligence covers translation, summarization, topic detection, chapters, sentiment, and audio events.

Melia 1 and Oak 1 match the Enhanced and Standard models for core transcription features, including diarization, word timings, punctuation, notifications, and output locale. They do not yet support the following features, which are available with the Enhanced and Standard models:

* Custom vocabulary and formatting: custom dictionary, find and replace, spoken form output, profanity tagging
* Output detail: confidence scores, entity detection, audio filtering
* Speaker identification
* Speech intelligence: audio events, translation, summarization, chapters, topics, sentiment

Check the [release notes](https://speechmatics.featurebase.app/en/changelog) for the latest feature support.

### Choose a model[​](#choose-a-model "Direct link to Choose a model")

Use Enhanced for the highest accuracy on single-language audio, such as medical, legal, or subtitling work.

Use Standard when throughput or latency matter more than maximum accuracy, such as archival transcription, content indexing, or large-scale captioning.

Use Melia 1 for audio that contains more than one language, including speakers who switch language mid-conversation.

Use Oak 1 for multilingual healthcare audio, such as ambient scribes and dictation tools where speakers switch languages mid-conversation. For single-language healthcare audio that needs full feature support, such as custom dictionary or confidence scores, use the [Enhanced Medical model](#healthcare-domain) instead.

### Specify a model[​](#specify-a-model "Direct link to Specify a model")

Set the `model` property in your transcription config. If you do not set it, the `standard` model is used.

This config selects the `enhanced` model:

```
{
  "type": "transcription",
  "transcription_config": {
    "model": "enhanced",
    "language": "en"
  }
}
```

Enhanced and Standard are available for Realtime and Batch transcription. Melia 1 and Oak 1 are currently available for Batch transcription.

### Melia 1[​](#melia-1 "Direct link to Melia 1")

Melia 1 is a multilingual model. It transcribes audio that contains more than one language, including speakers who switch language mid-conversation, and returns a single continuous transcript. It does not require you to select a language pack, and its accuracy is on par with the Standard model.

Set `"model": "melia-1"` and `"language": "multi"`:

```
{
  "type": "transcription",
  "transcription_config": {
    "model": "melia-1",
    "language": "multi"
  }
}
```

Melia 1 does not support the `auto` language value, which returns an error. Set `language` to `multi`.

Melia 1 is available for Batch transcription in the EU and US regions only. It is not available in the Australia (AU1) region.

| Region       | Endpoint                       |
| ------------ | ------------------------------ |
| EU1 (Europe) | `eu1.asr.api.speechmatics.com` |
| US1 (USA)    | `us1.asr.api.speechmatics.com` |

For the full list of Batch endpoints, refer to [Authentication](/get-started/authentication.md#supported-endpoints).

Melia 1 supports [output locale](/speech-to-text/formatting.md#output-locale), including an array of locales for multilingual jobs. Refer to [Output locale for multilingual transcription](/speech-to-text/formatting.md#output-locale-for-multilingual-transcription).

For the features Melia 1 does not yet support, refer to [Compare the models](#compare-the-models).

To configure language hints and read the per-language output metadata, refer to [Input](/speech-to-text/batch/input.md#language-hints) and [Output](/speech-to-text/batch/output.md#multilingual-transcript-output).

### Oak 1[​](#oak-1 "Direct link to Oak 1")

Oak 1 is a multilingual model tuned for healthcare audio, such as ambient scribes and dictation tools. Like Melia 1, it transcribes audio that contains more than one language, including speakers who switch language mid-conversation, and returns a single continuous transcript. It does not require you to select a language pack.

Set `"model": "oak-1"` and `"language": "multi"`:

```
{
  "type": "transcription",
  "transcription_config": {
    "model": "oak-1",
    "language": "multi"
  }
}
```

Oak 1 does not support the `auto` language value, which returns an error. Set `language` to `multi`.

Oak 1 is available for Batch transcription in the EU and US regions only. It is not available in the Australia (AU1) region.

| Region       | Endpoint                       |
| ------------ | ------------------------------ |
| EU1 (Europe) | `eu1.asr.api.speechmatics.com` |
| US1 (USA)    | `us1.asr.api.speechmatics.com` |

For the full list of Batch endpoints, refer to [Authentication](/get-started/authentication.md#supported-endpoints).

Oak 1 supports [output locale](/speech-to-text/formatting.md#output-locale), including an array of locales for multilingual jobs. Refer to [Output locale for multilingual transcription](/speech-to-text/formatting.md#output-locale-for-multilingual-transcription).

For the features Oak 1 does not yet support, refer to [Compare the models](#compare-the-models).

To configure language hints and read the per-language output metadata, refer to [Input](/speech-to-text/batch/input.md#language-hints) and [Output](/speech-to-text/batch/output.md#multilingual-transcript-output).

Oak 1 is a standalone model for healthcare audio. It is separate from the [Enhanced Medical model](#healthcare-domain): use Oak 1 for multilingual healthcare audio, and Enhanced Medical for single-language healthcare audio that needs full feature support. For the languages Oak 1 supports, refer to [Medical languages](/speech-to-text/languages.md#medical-languages).

### Enhanced Medical model[​](#healthcare-domain "Direct link to Enhanced Medical model")

Speechmatics offers the Enhanced Medical model that delivers high accuracy on healthcare audio such as ambient scribes and dictation tools.

The Enhanced Medical model is kept up to date using officially maintained data sources. This improves recognition of medical terminology such as procedures, medications, conditions, and anatomy.

For languages outside the Enhanced Medical language list, the Enhanced model still delivers high accuracy on healthcare audio without the medical domain. For multilingual healthcare audio, use [Oak 1](#oak-1) instead.

To use the Enhanced Medical model, set the `domain` property to `medical`:

```
{
  "type": "transcription",
  "transcription_config": {
    "model": "enhanced",
    "language": "en",
    "domain": "medical"
  }
}
```

For the languages the Enhanced Medical model supports, refer to [Medical languages](/speech-to-text/languages.md#medical-languages).

### Operating points[​](#operating-points "Direct link to Operating points")

The `model` property replaces the `operating_point` property. Existing configs that use `operating_point` continue to transcribe without changes.

In SaaS (cloud) deployments, `operating_point` is deprecated. It maps to `model` and accepts the same `enhanced` and `standard` values. Use [`model`](#specify-a-model) going forward.

## Agent STT models[​](#agent-stt-models "Direct link to Agent STT models")

### Linden 1[​](#linden-1 "Direct link to Linden 1")

Linden 1 is the model powering [Agent STT](/speech-to-text/agent-stt/.md). It returns transcription as speaker-attributed segments, with turn messages, rather than a running word stream, so the output is ready to pass to an LLM.

To use it, connect to the Agent STT endpoint at `/v2/agent` rather than the Realtime `/v2` path, and set `"model": "linden-1"` in the `transcription_config` of your `StartRecognition` message:

```
{
  "message": "StartRecognition",
  "transcription_config": {
    "model": "linden-1",
    "language": "en"
  }
}
```

If using the Agent STT SDK, it sets the model for you if you do not.

For the full configuration, refer to the [Agent STT Reference](/api-ref/agent-stt-websocket.md).
