Models
Speechmatics offers models in two groups. The Batch and Realtime APIs share a set of models you choose between with the model property. The Agent STT API has its own models, which are not available on Batch or Realtime.
Batch and Realtime models
Compare the models
Enhanced and Standard are feature-identical and differ only in accuracy and speed: Enhanced delivers the highest accuracy, and Standard prioritizes throughput. Melia 1 and Oak 1 add automatic multilingual transcription, but are available for Batch only and support a reduced feature set. Speech intelligence covers translation, summarization, topic detection, chapters, sentiment, and audio events.
Melia 1 and Oak 1 match the Enhanced and Standard models for core transcription features, including diarization, word timings, punctuation, notifications, and output locale. They do not yet support the following features, which are available with the Enhanced and Standard models:
- Custom vocabulary and formatting: custom dictionary, find and replace, spoken form output, profanity tagging
- Output detail: confidence scores, entity detection, audio filtering
- Speaker identification
- Speech intelligence: audio events, translation, summarization, chapters, topics, sentiment
Check the release notes for the latest feature support.
Choose a model
Use Enhanced for the highest accuracy on single-language audio, such as medical, legal, or subtitling work.
Use Standard when throughput or latency matter more than maximum accuracy, such as archival transcription, content indexing, or large-scale captioning.
Use Melia 1 for audio that contains more than one language, including speakers who switch language mid-conversation.
Use Oak 1 for multilingual healthcare audio, such as ambient scribes and dictation tools where speakers switch languages mid-conversation. For single-language healthcare audio that needs full feature support, such as custom dictionary or confidence scores, use the Enhanced Medical model instead.
Specify a model
Set the model property in your transcription config. If you do not set it, the standard model is used.
This config selects the enhanced model:
{
"type": "transcription",
"transcription_config": {
"model": "enhanced",
"language": "en"
}
}
Enhanced and Standard are available for Realtime and Batch transcription. Melia 1 and Oak 1 are currently available for Batch transcription.
Melia 1
Melia 1 is a multilingual model. It transcribes audio that contains more than one language, including speakers who switch language mid-conversation, and returns a single continuous transcript. It does not require you to select a language pack, and its accuracy is on par with the Standard model.
Set "model": "melia-1" and "language": "multi":
{
"type": "transcription",
"transcription_config": {
"model": "melia-1",
"language": "multi"
}
}
Melia 1 does not support the auto language value, which returns an error. Set language to multi.
Melia 1 is available for Batch transcription in the EU and US regions only. It is not available in the Australia (AU1) region.
For the full list of Batch endpoints, refer to Authentication.
Melia 1 supports output locale, including an array of locales for multilingual jobs. Refer to Output locale for multilingual transcription.
For the features Melia 1 does not yet support, refer to Compare the models.
To configure language hints and read the per-language output metadata, refer to Input and Output.
Oak 1
Oak 1 is a multilingual model tuned for healthcare audio, such as ambient scribes and dictation tools. Like Melia 1, it transcribes audio that contains more than one language, including speakers who switch language mid-conversation, and returns a single continuous transcript. It does not require you to select a language pack.
Set "model": "oak-1" and "language": "multi":
{
"type": "transcription",
"transcription_config": {
"model": "oak-1",
"language": "multi"
}
}
Oak 1 does not support the auto language value, which returns an error. Set language to multi.
Oak 1 is available for Batch transcription in the EU and US regions only. It is not available in the Australia (AU1) region.
For the full list of Batch endpoints, refer to Authentication.
Oak 1 supports output locale, including an array of locales for multilingual jobs. Refer to Output locale for multilingual transcription.
For the features Oak 1 does not yet support, refer to Compare the models.
To configure language hints and read the per-language output metadata, refer to Input and Output.
Oak 1 is a standalone model for healthcare audio. It is separate from the Enhanced Medical model: use Oak 1 for multilingual healthcare audio, and Enhanced Medical for single-language healthcare audio that needs full feature support. For the languages Oak 1 supports, refer to Medical languages.
Enhanced Medical model
Speechmatics offers the Enhanced Medical model that delivers high accuracy on healthcare audio such as ambient scribes and dictation tools.
The Enhanced Medical model is kept up to date using officially maintained data sources. This improves recognition of medical terminology such as procedures, medications, conditions, and anatomy.
For languages outside the Enhanced Medical language list, the Enhanced model still delivers high accuracy on healthcare audio without the medical domain. For multilingual healthcare audio, use Oak 1 instead.
To use the Enhanced Medical model, set the domain property to medical:
{
"type": "transcription",
"transcription_config": {
"model": "enhanced",
"language": "en",
"domain": "medical"
}
}
For the languages the Enhanced Medical model supports, refer to Medical languages.
Operating points
The model property replaces the operating_point property. Existing configs that use operating_point continue to transcribe without changes.
In SaaS (cloud) deployments, operating_point is deprecated. It maps to model and accepts the same enhanced and standard values. Use model going forward.
Agent STT models
Linden 1
Linden 1 is the model powering Agent STT. It returns transcription as speaker-attributed segments, with turn messages, rather than a running word stream, so the output is ready to pass to an LLM.
To use it, connect to the Agent STT endpoint at /v2/agent rather than the Realtime /v2 path, and set "model": "linden-1" in the transcription_config of your StartRecognition message:
{
"message": "StartRecognition",
"transcription_config": {
"model": "linden-1",
"language": "en"
}
}
If using the Agent STT SDK, it sets the model for you if you do not.
For the full configuration, refer to the Agent STT Reference.