For AI agents: a documentation index is available at /llms.txt. Markdown versions of all pages can be requested by appending `.md` to the URL, or by setting the `Accept` header to `text/markdown`.
Skip to main content
Speech to Text

Models

Compare the Speechmatics models for Batch, Realtime, and Agent STT, and choose the right one for your audio.

Speechmatics offers models in two groups. The Batch and Realtime APIs share a set of models you choose between with the model property. The Agent STT API has its own models, which are not available on Batch or Realtime.

APIModels
Batchenhanced, standard, melia-1, oak-1
Realtimeenhanced, standard
Agent STTlinden-1

Batch and Realtime models​

Compare the models​

CapabilityEnhancedStandardMelia 1Oak 1
Modelenhancedstandardmelia-1oak-1
Medical specializationMedical variant available——Medical native model
AccuracyHighestHighHighHigh
TurnaroundFastFastestFastestFastest
Processing modesBatch and RealtimeBatch and RealtimeBatchBatch
RegionsEU, US, AUSEU, US, AUSEU, USEU, US
Language handlingSelected language or pack (auto-detect available)Selected language or pack (auto-detect available)Automatic multilingualAutomatic multilingual
DiarizationSpeaker and channelSpeaker and channelSpeaker and channelSpeaker and channel
Language labeling——Per wordPer word
Custom dictionary✅✅Not yetNot yet
Confidence scores✅✅Not yetNot yet
Speaker identification✅✅Not yetNot yet
Speech intelligence✅✅Not yetNot yet

Enhanced and Standard are feature-identical and differ only in accuracy and speed: Enhanced delivers the highest accuracy, and Standard prioritizes throughput. Melia 1 and Oak 1 add automatic multilingual transcription, but are available for Batch only and support a reduced feature set. Speech intelligence covers translation, summarization, topic detection, chapters, sentiment, and audio events.

Melia 1 and Oak 1 match the Enhanced and Standard models for core transcription features, including diarization, word timings, punctuation, notifications, and output locale. They do not yet support the following features, which are available with the Enhanced and Standard models:

  • Custom vocabulary and formatting: custom dictionary, find and replace, spoken form output, profanity tagging
  • Output detail: confidence scores, entity detection, audio filtering
  • Speaker identification
  • Speech intelligence: audio events, translation, summarization, chapters, topics, sentiment

Check the release notes for the latest feature support.

Choose a model​

Use Enhanced for the highest accuracy on single-language audio, such as medical, legal, or subtitling work.

Use Standard when throughput or latency matter more than maximum accuracy, such as archival transcription, content indexing, or large-scale captioning.

Use Melia 1 for audio that contains more than one language, including speakers who switch language mid-conversation.

Use Oak 1 for multilingual healthcare audio, such as ambient scribes and dictation tools where speakers switch languages mid-conversation. For single-language healthcare audio that needs full feature support, such as custom dictionary or confidence scores, use the Enhanced Medical model instead.

Specify a model​

Set the model property in your transcription config. If you do not set it, the standard model is used.

This config selects the enhanced model:

{
"type": "transcription",
"transcription_config": {
"model": "enhanced",
"language": "en"
}
}

Enhanced and Standard are available for Realtime and Batch transcription. Melia 1 and Oak 1 are currently available for Batch transcription.

Melia 1​

Melia 1 is a multilingual model. It transcribes audio that contains more than one language, including speakers who switch language mid-conversation, and returns a single continuous transcript. It does not require you to select a language pack, and its accuracy is on par with the Standard model.

Set "model": "melia-1" and "language": "multi":

{
"type": "transcription",
"transcription_config": {
"model": "melia-1",
"language": "multi"
}
}

Melia 1 does not support the auto language value, which returns an error. Set language to multi.

Melia 1 is available for Batch transcription in the EU and US regions only. It is not available in the Australia (AU1) region.

RegionEndpoint
EU1 (Europe)eu1.asr.api.speechmatics.com
US1 (USA)us1.asr.api.speechmatics.com

For the full list of Batch endpoints, refer to Authentication.

Melia 1 supports output locale, including an array of locales for multilingual jobs. Refer to Output locale for multilingual transcription.

For the features Melia 1 does not yet support, refer to Compare the models.

To configure language hints and read the per-language output metadata, refer to Input and Output.

Oak 1​

Oak 1 is a multilingual model tuned for healthcare audio, such as ambient scribes and dictation tools. Like Melia 1, it transcribes audio that contains more than one language, including speakers who switch language mid-conversation, and returns a single continuous transcript. It does not require you to select a language pack.

Set "model": "oak-1" and "language": "multi":

{
"type": "transcription",
"transcription_config": {
"model": "oak-1",
"language": "multi"
}
}

Oak 1 does not support the auto language value, which returns an error. Set language to multi.

Oak 1 is available for Batch transcription in the EU and US regions only. It is not available in the Australia (AU1) region.

RegionEndpoint
EU1 (Europe)eu1.asr.api.speechmatics.com
US1 (USA)us1.asr.api.speechmatics.com

For the full list of Batch endpoints, refer to Authentication.

Oak 1 supports output locale, including an array of locales for multilingual jobs. Refer to Output locale for multilingual transcription.

For the features Oak 1 does not yet support, refer to Compare the models.

To configure language hints and read the per-language output metadata, refer to Input and Output.

Oak 1 is a standalone model for healthcare audio. It is separate from the Enhanced Medical model: use Oak 1 for multilingual healthcare audio, and Enhanced Medical for single-language healthcare audio that needs full feature support. For the languages Oak 1 supports, refer to Medical languages.

Enhanced Medical model​

Speechmatics offers the Enhanced Medical model that delivers high accuracy on healthcare audio such as ambient scribes and dictation tools.

The Enhanced Medical model is kept up to date using officially maintained data sources. This improves recognition of medical terminology such as procedures, medications, conditions, and anatomy.

For languages outside the Enhanced Medical language list, the Enhanced model still delivers high accuracy on healthcare audio without the medical domain. For multilingual healthcare audio, use Oak 1 instead.

To use the Enhanced Medical model, set the domain property to medical:

{
"type": "transcription",
"transcription_config": {
"model": "enhanced",
"language": "en",
"domain": "medical"
}
}

For the languages the Enhanced Medical model supports, refer to Medical languages.

Operating points​

The model property replaces the operating_point property. Existing configs that use operating_point continue to transcribe without changes.

In SaaS (cloud) deployments, operating_point is deprecated. It maps to model and accepts the same enhanced and standard values. Use model going forward.

Agent STT models​

Linden 1​

Linden 1 is the model powering Agent STT. It returns transcription as speaker-attributed segments, with turn messages, rather than a running word stream, so the output is ready to pass to an LLM.

To use it, connect to the Agent STT endpoint at /v2/agent rather than the Realtime /v2 path, and set "model": "linden-1" in the transcription_config of your StartRecognition message:

{
"message": "StartRecognition",
"transcription_config": {
"model": "linden-1",
"language": "en"
}
}

If using the Agent STT SDK, it sets the model for you if you do not.

For the full configuration, refer to the Agent STT Reference.