# Speechmatics Docs > Developer documentation for Speechmatics APIs - [Speechmatics Docs](/index.md) ## search - [Search the documentation](/search.md) ## administration ### accounts Sign up, manage sign-in methods, update your profile and security settings, and delete your account. - [Accounts](/administration/accounts.md): Sign up, manage sign-in methods, update your profile and security settings, and delete your account. ### api-keys Create, view, and revoke the API keys that authenticate your applications to Speechmatics. - [API keys](/administration/api-keys.md): Create, view, and revoke the API keys that authenticate your applications to Speechmatics. ### billing Manage your payment card, understand how credits are used, and view and download invoices. - [Billing](/administration/billing.md): Manage your payment card, understand how credits are used, and view and download invoices. ### domain-verification Verify a domain so users from your organization join your workspace automatically. - [Domain verification](/administration/domain-verification.md): Verify a domain so users from your organization join your workspace automatically. ### manage-members Invite users to your workspace, change their roles, and remove access. - [Manage members](/administration/manage-members.md): Invite users to your workspace, change their roles, and remove access. ### management-tokens Automate workspace administration with tokens that create projects and manage API keys programmatically. - [Management tokens](/administration/management-tokens.md): Automate workspace administration with tokens that create projects and manage API keys programmatically. ### plans Understand the Free, Pro, and Enterprise plans and how each one is billed. - [Plans](/administration/plans.md): Understand the Free, Pro, and Enterprise plans and how each one is billed. ### projects Organize work inside a workspace, scope API keys to projects, and track usage per project. - [Projects](/administration/projects.md): Organize work inside a workspace, scope API keys to projects, and track usage per project. ### regions Choose where Speechmatics processes your audio across the available regions. - [Regions](/administration/regions.md): Choose where Speechmatics processes your audio across the available regions. ### security-and-compliance Find Speechmatics security certifications and request compliance documentation. - [Security and compliance](/administration/security-and-compliance.md): Find Speechmatics security certifications and request compliance documentation. ### sso Configure single sign-on for your workspace so users authenticate through your identity provider. - [Single sign-on (SSO)](/administration/sso.md): Configure single sign-on for your workspace so users authenticate through your identity provider. ### usage Track transcription usage by processing mode and model in the portal and through the API. - [Usage](/administration/usage.md): Track transcription usage by processing mode and model in the portal and through the API. ### workspaces-concepts Understand how workspaces organize users, projects, billing, and access on the Speechmatics platform. - [Workspaces concepts](/administration/workspaces-concepts.md): Understand how workspaces organize users, projects, billing, and access on the Speechmatics platform. ## api-ref Browse API reference for Speechmatics APIs - [API reference](/api-ref.md): Browse API reference for Speechmatics APIs ### batch - [Create a new job](/api-ref/batch/create-a-new-job.md): Create a new job - [Delete a job](/api-ref/batch/delete-a-job.md): Delete a job and remove all associated resources. - [Get job details](/api-ref/batch/get-job-details.md): Get job details, including progress and any error reports. - [Get the log file for a job.](/api-ref/batch/get-the-log-file-for-a-job.md): Get the log file for a job. - [Get the transcript for a transcription job](/api-ref/batch/get-the-transcript-for-a-transcription-job.md): Get the transcript for a transcription job - [Get usage statistics](/api-ref/batch/get-usage-statistics.md): Get usage statistics - [List all jobs](/api-ref/batch/list-all-jobs.md): List all jobs - [Speechmatics ASR REST API](/api-ref/batch/speechmatics-asr-rest-api.md): The Speechmatics Automatic Speech Recognition REST API is used to submit ASR jobs and receive the results. The supported job type is transcription of audio files. ### management - [Create an API key](/api-ref/management/create-an-api-key.md): Create an API key in your workspace, authenticated with a [management token](/administration/management-tokens). - [Delete a project](/api-ref/management/delete-a-project.md): Delete a project - [Delete an API key](/api-ref/management/delete-an-api-key.md): Delete an API key - [Get a project by ID](/api-ref/management/get-a-project-by-id.md): Get a project by ID - [Get all API keys](/api-ref/management/get-all-api-keys.md): Get all API keys - [Get all projects](/api-ref/management/get-all-projects.md): Get all projects - [Management API](/api-ref/management/management-api.md): The Management API lets you manage projects and API keys in your workspace programmatically. Authenticate each request with a [management token](/administration/management-tokens), which is scoped to the workspace. - [Post a new project](/api-ref/management/post-a-new-project.md): Post a new project - [Update a project](/api-ref/management/update-a-project.md): Update a project ### realtime-transcription-websocket API Reference for the Realtime Websocket API - [Realtime API Reference](/api-ref/realtime-transcription-websocket.md): API Reference for the Realtime Websocket API ## deployments Learn about the different ways to use our APIs, including cloud services and on-prem containers. - [Overview](/deployments.md): Learn about the different ways to use our APIs, including cloud services and on-prem containers. ### container - [Accessing images](/deployments/container/accessing-images.md): Learn how to access images in the Speechmatics Container system - [Additional security features](/deployments/container/additional-security.md): Learn about the Speechmatics container system security - [Batch persistent worker](/deployments/container/batch-persistent-worker.md): Run a long-lived HTTP transcription worker that accepts multiple jobs without restarting, reducing turnaround time and improving CPU/GPU utilisation. - [CPU Speech to text container](/deployments/container/cpu-speech-to-text.md): Learn about the Speechmatics CPU container system - [GPU Speech to text container (Standard and Enhanced)](/deployments/container/gpu-speech-to-text.md): Learn about the Speechmatics Transcription GPU container system - [GPU Speech to text container (Melia 1)](/deployments/container/gpu-speech-to-text-melia-1.md): Learn about the Speechmatics Melia 1 GPU container system - [Translation GPU inference container](/deployments/container/gpu-translation.md): Learn about the Speechmatics Translation GPU container system - [Language ID container](/deployments/container/language-id.md): Learn about the Speechmatics language ID Container - [Licensing](/deployments/container/licensing.md): Learn about the licensing for Speechmatics containers - [Performance and cost](/deployments/container/performance-and-cost.md): Get an overview of the performance and cost of Speechmatics container deployments - [Speaker identification secrets](/deployments/container/speaker-identification.md): Prepare and manage Speaker Identification secrets for Speechmatics deployments - [Troubleshooting](/deployments/container/troubleshooting.md): Troubleshooting for Speechmatics containers ### kubernetes Learn about the Kubernetes deployment options for Speechmatics - [Kubernetes](/deployments/kubernetes.md): Learn about the Kubernetes deployment options for Speechmatics - [Prerequisites](/deployments/kubernetes/prerequisites.md): Prerequisites for deploying Realtime on Kubernetes - [Realtime](/deployments/kubernetes/realtime.md): Learn about the Kubernetes deployment options for Realtime ### usage-reporting Learn about the usage reporting for on-prem deployments - [Usage reporting](/deployments/usage-reporting.md): Learn about the usage reporting for on-prem deployments - [Automatic usage reporting](/deployments/usage-reporting/automatic.md): Learn about automatic usage reporting for on-prem deployments - [Offline usage reporting](/deployments/usage-reporting/offline.md): Learn about offline usage reporting for on-prem deployments ### virtual-appliance Deploy Speechmatics to your own hardware. - [Virtual Appliance](/deployments/virtual-appliance.md): Deploy Speechmatics to your own hardware. - [Adding Languages](/deployments/virtual-appliance/administration/adding-languages.md): Add languages to a Virtual Appliance deployment - [Language Identification](/deployments/virtual-appliance/administration/language-identification.md): Configure Language Identification on a Virtual Appliance deployment - [Logcli Help](/deployments/virtual-appliance/administration/logcli-help.md): Usage manual for the `logcli` command-line tool. - [Monitoring](/deployments/virtual-appliance/administration/monitoring.md): Monitor the appliance resources. - [Networking](/deployments/virtual-appliance/administration/networking.md): Configure the appliance's network settings. - [Remote Access](/deployments/virtual-appliance/administration/remote-access.md): Configure remote access to the appliance. - [Virtual appliance scaling](/deployments/virtual-appliance/administration/scaling.md): Increase appliance performance by scaling the number of threads. - [Security](/deployments/virtual-appliance/administration/security.md): Configure security settings for the appliance. - [Services](/deployments/virtual-appliance/administration/services.md): Configure services for the appliance. - [SSL configuration](/deployments/virtual-appliance/administration/ssl-configuration.md): Configure SSL settings for the appliance. - [Using a GPU](/deployments/virtual-appliance/administration/using-a-gpu.md): Enable GPU processing for the appliance - [Download and import](/deployments/virtual-appliance/installation/download-and-import.md): Get started with your Virtual Appliance deployment. - [Licensing features](/deployments/virtual-appliance/installation/license-features.md): Ensure the appliance is licensed for your use case. - [Licensing](/deployments/virtual-appliance/installation/licensing.md): Ensure you have a valid license for your deployment. - [Network configuration](/deployments/virtual-appliance/installation/network-config.md): Set up the appliance's network configuration. - [System requirements](/deployments/virtual-appliance/installation/system-requirements.md): Ensure your system is ready to run the appliance. - [Verify and go](/deployments/virtual-appliance/installation/verify-and-go.md): Confirm the appliance is working properly. ## get-started ### authentication Learn about how the Speechmatics API handles authentication - [Authentication](/get-started/authentication.md): Learn about how the Speechmatics API handles authentication ### quickstart Take your first steps with the Speechmatics API. - [Quickstart](/get-started/quickstart.md): Take your first steps with the Speechmatics API. ## integrations-and-sdks Discover which integrations and SDKs to add Speechmatics' STT, TTS or voice agents to your applications. - [Overview](/integrations-and-sdks.md): Discover which integrations and SDKs to add Speechmatics' STT, TTS or voice agents to your applications. ### livekit Build a voice AI agent with Speechmatics STT and TTS using LiveKit Agents. - [LiveKit quickstart](/integrations-and-sdks/livekit.md): Build a voice AI agent with Speechmatics STT and TTS using LiveKit Agents. - [LiveKit speech to text](/integrations-and-sdks/livekit/stt.md): Transcribe live audio in your LiveKit voice agents with Speechmatics STT. - [LiveKit text to speech](/integrations-and-sdks/livekit/tts.md): Use Speechmatics text-to-speech voices in your LiveKit voice agents. ### pipecat Build a local voice bot with Speechmatics STT and TTS using Pipecat. - [Pipecat quickstart](/integrations-and-sdks/pipecat.md): Build a local voice bot with Speechmatics STT and TTS using Pipecat. - [Pipecat speech to text](/integrations-and-sdks/pipecat/stt.md): Transcribe live audio in your Pipecat voice bots with Speechmatics STT. - [Pipecat text to speech](/integrations-and-sdks/pipecat/tts.md): Use Speechmatics text to speech voices in your Pipecat voice bots. ### sdks Learn how to use the Speechmatics SDKs - [SDKs](/integrations-and-sdks/sdks.md): Learn how to use the Speechmatics SDKs ### vapi Learn how to integrate Speechmatics STT with Vapi. - [Vapi integration](/integrations-and-sdks/vapi.md): Learn how to integrate Speechmatics STT with Vapi. ### zapier Transcribe audio in your Zapier workflow - no code required. - [Zapier integration](/integrations-and-sdks/zapier.md): Transcribe audio in your Zapier workflow - no code required. ## private ### melia-1-realtime Melia 1 delivers code-switching for 56 languages, in Realtime. - [Melia 1 Realtime Preview - Multilingual transcription](/private/melia-1-realtime.md): Melia 1 delivers code-switching for 56 languages, in Realtime. ### next-gen-model Get started with our next-generation model - [Next-generation Model](/private/next-gen-model.md): Get started with our next-generation model ### preview-mode Get early access to features - [Preview Mode](/private/preview-mode.md): Get early access to features ### voice-agent-api Early access to the Voice Agent API — a turn-based API built for voice agents - [Voice Agent API](/private/voice-agent-api.md): Early access to the Voice Agent API — a turn-based API built for voice agents ## speech-to-text Learn how to turn audio into text. - [Speech to Text overview](/speech-to-text.md): Learn how to turn audio into text. ### accuracy-benchmarking How to calculate Word Error Rate - [Accuracy benchmarking](/speech-to-text/accuracy-benchmarking.md): How to calculate Word Error Rate ### app-analytics Track usage by adding an application ID to your requests. - [App analytics](/speech-to-text/app-analytics.md): Track usage by adding an application ID to your requests. ### batch - [Alignment](/speech-to-text/batch/alignment.md): Learn about the Speechmatics transcription alignment product - [Batch diarization](/speech-to-text/batch/batch-diarization.md): Learn how to use the Speechmatics API to separate speakers in Batch - [Input](/speech-to-text/batch/input.md): Learn about configuration and supported input audio formats for the Speechmatics Batch API - [Language identification (SaaS)](/speech-to-text/batch/language-identification.md): Learn about Speechmatics Language ID - [Limits – Batch](/speech-to-text/batch/limits.md): Learn about rate limiting and usage limits for the Speechmatics Batch API - [Notifications](/speech-to-text/batch/notifications.md): Learn how Speechmatics notifications work - [Output](/speech-to-text/batch/output.md): Learn about the supported output formats for the Speechmatics Batch API - [Quickstart](/speech-to-text/batch/quickstart.md): Learn how to transcribe pre-recorded audio and video files. - [Batch speaker identification](/speech-to-text/batch/speaker-identification.md): Learn how to use the Speechmatics API to identify speakers in Batch - [Chapters](/speech-to-text/batch/speech-intelligence/auto-chapters.md): Learn how to use Speechmatics' Chapters. - [Sentiment analysis](/speech-to-text/batch/speech-intelligence/sentiment-analysis.md): Learn about the sentiment analysis offering for the Speechmatics Batch API - [Summarization](/speech-to-text/batch/speech-intelligence/summarization.md): Learn how to use Speechmatics's summarization feature. - [Topics](/speech-to-text/batch/speech-intelligence/topic-detection.md): Learn how to use Speechmatics' Topics. - [SRT formatting](/speech-to-text/batch/srt-format.md): Learn how to get Speechmatics transcriptions in an SRT format - [Synchronous transcription](/speech-to-text/batch/synchronous.md): Wait for a transcription job to finish synchronously via HTTP - [Troubleshooting](/speech-to-text/batch/troubleshooting.md): Guides to help with troubleshooting the Speechmatics API - [Usage reporting](/speech-to-text/batch/usage.md): Learn how to get information about your API usage ### features - [Audio events](/speech-to-text/features/audio-events.md): Learn how to utilize the Audio Events feature in your media processing workflows - [Audio filtering](/speech-to-text/features/audio-filtering.md): Learn how to utilize Audio Filtering to remove background speech - [Custom dictionary](/speech-to-text/features/custom-dictionary.md): Learn how to use the Speechmatics custom dictionary - [Diarization](/speech-to-text/features/diarization.md): Learn how Speechmatics diarization separates speakers in audio - [Feature discovery](/speech-to-text/features/feature-discovery.md): Learn how to use Speechmatics' Discovery API. - [Speaker identification](/speech-to-text/features/speaker-identification.md): Learn how Speechmatics identifies speakers in audio - [Translation](/speech-to-text/features/translation.md): Translate your audio into multiple languages with a single API call. ### formatting Control how numbers, punctuation, and special text appear in your transcripts. - [Formatting](/speech-to-text/formatting.md): Control how numbers, punctuation, and special text appear in your transcripts. ### languages See which languages Speechmatics supports for transcription and translation, including bilingual packs. - [Languages](/speech-to-text/languages.md): See which languages Speechmatics supports for transcription and translation, including bilingual packs. ### models Compare the Enhanced, Standard, and Melia 1 models and choose the right one for your audio. - [Models](/speech-to-text/models.md): Compare the Enhanced, Standard, and Melia 1 models and choose the right one for your audio. ### realtime - [Python using FFMPEG](/speech-to-text/realtime/guides/python-using-ffmpeg.md): Use ffmpeg to pipe microphone input into the Speechmatics Realtime API - [Python microphone input](/speech-to-text/realtime/guides/python-using-microphone.md): Use the Speechmatics Python library to transcribe your voice using a microphone. - [Input](/speech-to-text/realtime/input.md): Learn about the supported input audio formats for the Speechmatics Realtime API - [Limits – Realtime](/speech-to-text/realtime/limits.md): Learn about the limits for the Speechmatics Realtime API - [Output](/speech-to-text/realtime/output.md): Learn about latency in the Speechmatics Realtime server - [Quickstart](/speech-to-text/realtime/quickstart.md): Learn how to transcribe streaming audio to text in real-time. - [Realtime diarization](/speech-to-text/realtime/realtime-diarization.md): Learn how to use the Speechmatics API to separate speakers in real-time - [Realtime speaker identification](/speech-to-text/realtime/speaker-identification.md): Learn how to use the Speechmatics API to identify speakers in real-time - [Turn detection](/speech-to-text/realtime/turn-detection.md): Learn how to detect the end of speech ## text-to-speech ### quickstart Learn how to convert text to speech using our API. - [Quickstart](/text-to-speech/quickstart.md): Learn how to convert text to speech using our API. ## voice-agents ### overview Learn how to build voice agents with Speechmatics integrations and the Voice SDK. - [Voice agents overview](/voice-agents/overview.md): Learn how to build voice agents with Speechmatics integrations and the Voice SDK. ### voice-sdk Learn how to use the Voice SDK. - [Voice SDK](/voice-agents/voice-sdk.md): Learn how to use the Voice SDK. --- # Full Documentation Content For AI agents: a documentation index is available at /llms.txt. Markdown versions of all pages can be requested by appending \`.md\` to the URL, or by setting the \`Accept\` header to \`text/markdown\`. [Skip to main content](#__docusaurus_skipToContent_fallback) [![Speechmatics Logo](/img/logo-text.svg)![Speechmatics Logo](/img/logo-text-dark.svg)](/index.md) Search [Docs](/index.md)[API Reference](/api-ref/.md) [Sign up](https://portal.speechmatics.com/signup) # Search the documentation Type your search here Powered by[](https://www.algolia.com/) Copyright © 2026 Speechmatics. Built with Docusaurus. --- # Accounts Sign up, manage sign-in methods, update your profile and security settings, and delete your account. An account is your individual identity on the Speechmatics platform. Accounts belong to one or more [workspaces](/administration/workspaces-concepts.md), where you act as either an Admin or a Member. ## Sign up and verification[​](#sign-up-and-verification "Direct link to Sign up and verification") 1. Go to [portal.speechmatics.com/signup](https://portal.speechmatics.com/signup). 2. Sign up with email and password, or with a supported OAuth provider (Google or GitHub). 3. If you signed up with email, verify your email address using the link sent to your inbox. If your email address matches a [verified domain](/administration/domain-verification.md), you join that workspace automatically. Otherwise, a new workspace is created for you. ## Sign-in methods[​](#sign-in-methods "Direct link to Sign-in methods") You can sign in with email and password, or with a connected OAuth account. Speechmatics supports Google and GitHub. Manage connected accounts under **[Settings > Profile > User](https://portal.speechmatics.com/settings/profile/user)**. The available sign-in methods are determined by your workspace. If your workspace has [SSO](/administration/sso.md) enabled, sign in through your identity provider instead. ## Profile and security[​](#profile-and-security "Direct link to Profile and security") Go to **Settings > Profile** in the portal to manage your account. * **User**: update your name, view your email address, and manage connected OAuth accounts. * **Security**: change your password and manage multi-factor authentication (MFA). * **Signed-in devices**: view active sessions and sign out of individual devices or all other devices at once. ## Delete your account[​](#delete-your-account "Direct link to Delete your account") Deleting your account is permanent. You will no longer be able to create an account with the same email address. 1. Go to **Settings > Profile > User**. 2. In the **Delete account** section, click **Delete account**. 3. Confirm the deletion. You cannot delete your account while you are an Admin in multiple workspaces. Transfer Admin access to another member first, or contact . ## Next steps[​](#next-steps "Direct link to Next steps") * [Workspaces concepts](/administration/workspaces-concepts.md): how workspaces organize members, projects, and billing. * [Single sign-on (SSO)](/administration/sso.md): sign in through your organization's identity provider. --- # API keys Create, view, and revoke the API keys that authenticate your applications to Speechmatics. An API key authenticates requests to the Speech to Text and Text to Speech APIs. Each key is scoped to a single [project](/administration/projects.md): it can access only the transcripts produced within that project. For how keys are used in requests, see [Authentication](/get-started/authentication.md). API keys differ from [management tokens](/administration/management-tokens.md). API keys authenticate transcription and synthesis requests and are scoped to a project. Management tokens authenticate workspace administration and are scoped to the workspace. ## Create an API key[​](#create-an-api-key "Direct link to Create an API key") API keys are created within your active project. To create a key in a different project, switch projects first. See [Projects](/administration/projects.md#switch-between-projects). 1. Go to **API keys** in the sidebar under **Projects**. 2. Click **Create new key**. 3. Enter a descriptive name for the key. 4. Click **Generate new key**. 5. Copy the key value from the dialog. You cannot access the key value again after closing this dialog. Store it in a secure location immediately. ## Revoke an API key[​](#revoke-an-api-key "Direct link to Revoke an API key") Revoking a key disables it immediately. Requests made with the key are rejected, which may break any application that depends on it. 1. Go to **API keys** in the sidebar under **Projects**. 2. Click the delete button next to the key. 3. Confirm the removal. ## Manage API keys with the API[​](#manage-api-keys-with-the-api "Direct link to Manage API keys with the API") You can manage API keys programmatically with the Management API, authenticated with a [management token](/administration/management-tokens.md). Available operations include: * [Get all API keys](/api-ref/management/get-all-api-keys.md) * [Create an API key](/api-ref/management/create-an-api-key.md) * [Delete an API key](/api-ref/management/delete-an-api-key.md) ## API key best practices[​](#api-key-best-practices "Direct link to API key best practices") * **Use descriptive names.** Name keys after the application or environment that uses them, so you can audit and revoke the right key later. * **Scope keys to projects.** Create keys in the project that matches their purpose to keep transcripts and usage isolated. * **Rotate keys periodically.** Create a replacement key, update your application, then revoke the old one. * **Revoke unused keys.** Remove any key that is no longer in use. ## Next steps[​](#next-steps "Direct link to Next steps") * [Authentication](/get-started/authentication.md): how to use API keys in requests. * [Projects](/administration/projects.md): scope keys by creating them in the right project. * [Management tokens](/administration/management-tokens.md): automate key management programmatically. --- # Billing Manage your payment card, understand how credits are used, and view and download invoices. On the Billing page you add or manage your payment card, review your credits, and see your invoices. Free users add a card here to upgrade to Pro. Managing billing requires the Admin role. Enterprise billing is handled through a contract, not a payment card. The card and invoices described here apply to the Pro plan only. See [Enterprise](/administration/plans.md#enterprise). ## Credits and pay-as-you-go[​](#credits-and-pay-as-you-go "Direct link to Credits and pay-as-you-go") Speechmatics bills in credits. One credit is worth one US dollar. New sign-ups receive a $100 credit grant. Accounts created before 1 August 2026 receive a one-time $25 transition credit. Pro accounts pay through pay-as-you-go (PAYG): you keep a payment card on file, and usage draws down your granted credits first. Once your granted credits are used up, further usage is charged to your card on the first of each month for the previous month's usage. ## Add or edit a payment card[​](#add-or-edit-a-payment-card "Direct link to Add or edit a payment card") Adding a card upgrades a Free plan to Pro. See [Plans](/administration/plans.md). 1. Go to **Billing** in the sidebar under **Workspace**. 2. In the **Payment method** section, click **Add payment card**, or **Edit payment information** if you already have a card. 3. Enter your card details. 4. Save your changes. ## Remove a payment card[​](#remove-a-payment-card "Direct link to Remove a payment card") Without a payment card, you can keep using the API for as long as you have credits. Once your credits are used up, add a card again to carry on. 1. In the **Payment method** section, click the delete icon next to your card. 2. Confirm the removal. ## Invoices and payments[​](#invoices-and-payments "Direct link to Invoices and payments") The **Payments** section lists each billing period with its usage, cost, and status. Each period covers one month, and any charge for usage beyond your granted credits is raised on the first of the following month. For how charges are calculated, see [Plans](/administration/plans.md#upgrade-to-pro). A billing period has one of the following statuses: * **Paid.** The charge has been settled. * **Due.** The charge is outstanding. * **Skipped.** No charge was raised for the period, for example when usage was fully covered by your granted credits. To view and download the invoice for a period, click the link icon at the end of its row. ## Next steps[​](#next-steps "Direct link to Next steps") * [Plans](/administration/plans.md): how Free, Pro, and Enterprise billing works. * [Usage](/administration/usage.md): see a detailed breakdown of usage by mode and model. --- # Domain verification Verify a domain so users from your organization join your workspace automatically. Once a domain is verified, anyone signing up with an email address on that domain is added to your workspace as a Member. Domain verification requires the Admin role. ## Verify a domain[​](#verify-a-domain "Direct link to Verify a domain") 1. Go to **Manage workspace** in the sidebar under **Workspace**. 2. In the **Domain** section, click **Add domain**. 3. Enter the domain you want to verify. 4. Add the DNS TXT record shown to your domain's DNS settings. 5. Return to the portal and click **Verify**. DNS propagation can take up to 24 hours. The domain status changes to **Verified** once the record is detected. ## Remove a verified domain[​](#remove-a-verified-domain "Direct link to Remove a verified domain") 1. In the **Domain** section, click the delete icon next to the verified domain. 2. Confirm the removal. Removing a domain stops new automatic joins. Existing members keep their access. ## Next steps[​](#next-steps "Direct link to Next steps") * [Manage members](/administration/manage-members.md): invite users manually or change their roles. * [Single sign-on (SSO)](/administration/sso.md): configure SSO for your workspace. --- # Manage members Invite users to your workspace, change their roles, and remove access. Member management requires the Admin role. See [Workspaces concepts](/administration/workspaces-concepts.md#roles) for what each role can do. ## Invite a user[​](#invite-a-user "Direct link to Invite a user") 1. Go to **Manage workspace** in the sidebar under **Workspace**. 2. In the **Users** section, click **Invite user**. 3. Enter the user's email address. 4. Select a role: **Admin** or **Member**. 5. Click **Invite**. The invited user receives an email with a link to complete their account. Until they accept, they will not appear as an active member. ## Change a user's role[​](#change-a-users-role "Direct link to Change a user's role") 1. In the **Users** table, click the options menu (⋯) next to the user. 2. Select **Edit role**. 3. Choose the new role and confirm. ## Remove a user[​](#remove-a-user "Direct link to Remove a user") Removing a user revokes their access to the workspace immediately. Any API keys they created stay in the workspace and continue to work, because keys belong to the workspace rather than the user. Revoke them separately if needed. See [API keys](/administration/api-keys.md). 1. In the **Users** table, click the options menu (⋯) next to the user. 2. Select **Remove user**. 3. Confirm the removal. ## Next steps[​](#next-steps "Direct link to Next steps") * [Domain verification](/administration/domain-verification.md): let users from your organization join automatically. * [Single sign-on (SSO)](/administration/sso.md): configure SSO for your workspace. --- # Management tokens Automate workspace administration with tokens that create projects and manage API keys programmatically. ## How management tokens work[​](#how-management-tokens-work "Direct link to How management tokens work") Management tokens authenticate programmatic access to workspace administration. Use them to automate project creation, API key management, and other workspace operations through the Management API. Tokens are scoped to the workspace and do not expire unless you revoke them. Each token can be assigned a subset of permissions, so you can limit access to only what your automation requires. ### Permissions[​](#permissions "Direct link to Permissions") | Permission | Description | | --------------- | ---------------------------------------------------- | | View projects | View projects created within your workspace. | | Manage projects | Manage, edit, and delete projects in your workspace. | | View API keys | View API keys generated in your workspace. | | Delete API keys | Delete API keys in your workspace. | | Create API key | Create API keys in your workspace. | ## Create a management token[​](#create-a-management-token "Direct link to Create a management token") 1. Go to **Manage workspace** in the sidebar under **Workspace**. 2. In the **Management tokens** section, click **+ Create management token**. 3. Enter a descriptive name for the token. 4. Select the permissions your automation needs. 5. Click **Create management token**. 6. Copy the token value from the **Save your key** dialog. You cannot access the token value again after closing this dialog. Store it in a secure location immediately. ## View token details[​](#view-token-details "Direct link to View token details") 1. In the **Management tokens** table, click the options menu (⋯) next to the token. 2. Select **View details**. The details panel displays the token name, key prefix, last used timestamp, assigned permissions, and creation date. ## Revoke a management token[​](#revoke-a-management-token "Direct link to Revoke a management token") Revoking a token disables it immediately. API requests made with the token are rejected, which may break any systems that depend on it. Revoking a management token cannot be undone. 1. In the **Management tokens** table, click the options menu (⋯) next to the token. 2. Select **Revoke key**. 3. Review the token name and key in the confirmation dialog. 4. Click **Revoke key** to confirm. ## Management token best practices[​](#management-token-best-practices "Direct link to Management token best practices") * **Limit permissions.** Assign only the permissions each token needs. A token that creates API keys does not need permission to delete them. * **Use descriptive names.** Name tokens after their purpose or the system that uses them, such as `ci-pipeline` or `key-rotation-script`. This makes it easier to audit and revoke the right token later. * **Rotate tokens periodically.** Create a replacement token, update your systems, then revoke the old one. * **Revoke unused tokens.** Check the **Last used** column regularly. Revoke any token that has never been used or has not been used recently. ## Next steps[​](#next-steps "Direct link to Next steps") * [Create an API key](/api-ref/management/create-an-api-key.md): Management API reference for creating keys programmatically. * [Projects](/administration/projects.md): create and manage projects. * [API keys](/administration/api-keys.md): how API keys differ from management tokens. --- # Plans Understand the Free, Pro, and Enterprise plans and how each one is billed. Speechmatics offers two self-serve plans, Free and Pro, plus Enterprise for contract-based usage. * **Free.** Get started with a $100 credit grant at no cost, with no payment card required. * **Pro.** A self-serve plan you pay for with pay-as-you-go billing, with volume and model training discounts. * **Enterprise.** A contract-based plan with dedicated support and deployment options beyond the cloud. See [Enterprise](#enterprise). Billing is in credits: one credit is worth one US dollar. New sign-ups receive a $100 credit grant. On Free you use these credits at no cost; on Pro you add a payment card and pay for usage beyond your granted credits through pay-as-you-go (PAYG). See [Billing](/administration/billing.md). Accounts created before 1 August 2026 receive a one-time $25 transition credit. Prices are not listed here. For current rates, see the [Speechmatics pricing page](https://www.speechmatics.com/pricing). ## Upgrade to Pro[​](#upgrade-to-pro "Direct link to Upgrade to Pro") To upgrade from Free to Pro, add a payment card on the Billing page. See [Billing](/administration/billing.md). On Pro, usage draws down your granted credits first. Once your granted credits are used up, further usage is charged to your payment card on the first of each month for the previous month's usage. Charges are calculated to the exact second, based on the per-hour rate for each product type. ## Discounts[​](#discounts "Direct link to Discounts") Pro plans can access two discounts, which can be applied at the same time. For current discount rates, see the [pricing page](https://www.speechmatics.com/pricing). ### Model training discount[​](#model-training-discount "Direct link to Model training discount") Enable **Model training** to receive a discount on your usage. Model training, also called data logging, lets Speechmatics use anonymized data to improve its models and services. You can change this setting at any time, and it applies to future usage only. Enable **Model training** in **Manage workspace**. ### Volume discount[​](#volume-discount "Direct link to Volume discount") A volume discount applies automatically to billable usage above 500 hours per month, for each Speech to Text product type (a combination of processing mode and model, such as Realtime Enhanced). For example, if in one month you use 800 hours of Realtime Enhanced and 400 hours of Realtime Standard, you are billed as: * 500 hours of Realtime Enhanced at the base rate * 300 hours of Realtime Enhanced at the discounted rate * 400 hours of Realtime Standard at the base rate ## Enterprise[​](#enterprise "Direct link to Enterprise") Enterprise is a contract-based plan for organizations that need dedicated support or deployment options beyond the cloud. It is separate from the self-serve Free and Pro plans: billing is handled through a contract rather than pay-as-you-go, with no payment card or monthly portal billing. Enterprise includes: * **High-touch support.** Dedicated customer support, onboarding, and solutions engineering. * **On-prem and on-device deployment.** These deployment options are available only on Enterprise. See [Deployments](/deployments/.md). * The best available discounts and early access to new features. To discuss an Enterprise plan, [contact sales](https://www.speechmatics.com/speak-to-sales). ## Next steps[​](#next-steps "Direct link to Next steps") * [Billing](/administration/billing.md): add a payment card, and view invoices and payments. * [Usage](/administration/usage.md): track the hours that determine your charges. --- # Projects Organize work inside a workspace, scope API keys to projects, and track usage per project. A project is an organizational unit inside a [workspace](/administration/workspaces-concepts.md). API keys, the transcripts they produce, and any voice agent configurations belong to a project. Any member of the workspace can access any project within it. Each workspace starts with a default project, which cannot be deleted. A workspace can have up to 1000 projects. If you need more, use [temporary keys](/get-started/authentication.md) or contact . API keys are scoped to a single project: a key created in one project cannot access transcripts produced in another. Billing and usage are not segmented by project, though usage can be viewed per project. See [Usage](/administration/usage.md). ## Common use cases[​](#common-use-cases "Direct link to Common use cases") * **Development workflows.** Separate environments such as development, staging, and production. * **Organization segmentation.** Monitor usage by product line, department, or business unit. * **Multi-client service providers.** Give each client their own API keys and isolate their transcripts and usage data. ## Create a project[​](#create-a-project "Direct link to Create a project") 1. Open the projects dropdown in the top navigation. 2. Select **Create project**. 3. Enter a descriptive name. 4. Click **Create**. ## Switch between projects[​](#switch-between-projects "Direct link to Switch between projects") Your active project is shown in the top navigation. Transcribing in the portal and managing API keys both apply to the active project. 1. Click the projects dropdown. 2. Select a project. When using the API, you do not specify a project. Each API key is tied to one project, so the key determines which project a request applies to. ## Rename or delete a project[​](#rename-or-delete-a-project "Direct link to Rename or delete a project") Manage projects under **Settings > Projects**. The project ID is read-only, and the default project cannot be deleted. Deleting a project permanently removes all of its transcripts, usage history, and API keys. This cannot be undone. ## Manage projects with the API[​](#manage-projects-with-the-api "Direct link to Manage projects with the API") You can manage projects programmatically with the Management API, authenticated with a [management token](/administration/management-tokens.md). Available operations include: * [Get all projects](/api-ref/management/get-all-projects.md) * [Get a project by ID](/api-ref/management/get-a-project-by-id.md) * [Create a project](/api-ref/management/post-a-new-project.md) * [Update a project](/api-ref/management/update-a-project.md) * [Delete a project](/api-ref/management/delete-a-project.md) ## Next steps[​](#next-steps "Direct link to Next steps") * [API keys](/administration/api-keys.md): create and scope keys to a project. * [Usage](/administration/usage.md): track usage for a single project or across the workspace. --- # Regions Choose where Speechmatics processes your audio across the available regions. A region is the location where your audio is processed. Each region is reached through its own API endpoint, so the region you use is determined by the endpoint you call. For the endpoint hostnames, see [supported endpoints](/get-started/authentication.md#supported-endpoints). Whether anything is stored depends on the processing mode, not the region. See [Data handling](/administration/security-and-compliance.md#data-handling). ## Available regions[​](#available-regions "Direct link to Available regions") The following regions are available to all customers: * **EU1** (Europe) * **US1** (United States) * **AU1** (Australia) Batch and Realtime are both available in all three regions. ## Select a region[​](#select-a-region "Direct link to Select a region") In the portal, use the **Region** selector in the top navigation. Your selection sets the region for transcription started from the portal. When using the API, the endpoint you call determines the region. A job is created in the region of the endpoint used, and all requests relating to that job must use the same endpoint. ## Use multiple regions[​](#use-multiple-regions "Direct link to Use multiple regions") You can use more than one region to balance load or to fail over if one region is disrupted. Because each job belongs to the region it was created in, retrieve a job from the same region that created it. Enterprise customers may have region arrangements set in their contract. To use a different region, contact your account manager or [support](https://support.speechmatics.com). ## Next steps[​](#next-steps "Direct link to Next steps") * [Supported endpoints](/get-started/authentication.md#supported-endpoints): the endpoint hostname for each region. * [Usage](/administration/usage.md): track usage across your workspace. --- # Security and compliance Find Speechmatics security certifications and request compliance documentation. Speechmatics maintains the following certifications and compliance standards: * ISO/IEC 27001:2022 * SOC 2 Type II * GDPR * HIPAA For an overview of how Speechmatics protects your data, see the [security page](https://www.speechmatics.com/security). ## Request documentation[​](#request-documentation "Direct link to Request documentation") Security and compliance documentation, including certificates and reports, is distributed through the [Speechmatics Trust Center](https://speechmatics.safebase.us/). To access documentation, request access in the Trust Center, agree to the non-disclosure agreement, and download the documents you need. Access is reviewed and granted by the Speechmatics security team. ## Data handling[​](#data-handling "Direct link to Data handling") How long your data is stored depends on the processing mode: * **Realtime.** Audio is streamed over a WebSocket and is never recorded or stored. Transcripts are streamed back as they are produced and then discarded. Processing happens in memory. * **Batch.** Audio files, transcripts, and job configuration are stored for 7 days, then deleted automatically. You can delete them sooner with the API. See [Delete a job](/api-ref/batch/delete-a-job.md). ## Next steps[​](#next-steps "Direct link to Next steps") * [Security page](https://www.speechmatics.com/security): how Speechmatics secures your data. * [Trust Center](https://speechmatics.safebase.us/): request certificates and compliance reports. --- # Single sign-on (SSO) Configure single sign-on for your workspace so users authenticate through your identity provider. SSO lets users sign in to Speechmatics using your existing identity provider, such as Okta, Azure AD, or Google Workspace. Authentication is delegated to the provider; Speechmatics does not store user passwords for SSO users. SSO is an add-on service with a separate subscription. To enable SSO for your workspace, contact . ## Supported protocols[​](#supported-protocols "Direct link to Supported protocols") Speechmatics supports SSO through [WorkOS AuthKit](https://workos.com/docs/sso): * SAML 2.0 * OpenID Connect (OIDC) ## Configure SSO[​](#configure-sso "Direct link to Configure SSO") Once SSO is enabled for your workspace by Speechmatics support, follow the provider-specific setup instructions to exchange metadata between your identity provider and Speechmatics. Configuration requires the Admin role. ## Next steps[​](#next-steps "Direct link to Next steps") * [Domain verification](/administration/domain-verification.md): verify your domain so SSO users join your workspace automatically. * [Manage members](/administration/manage-members.md): assign roles to users who join through SSO. --- # Usage Track transcription usage by processing mode and model in the portal and through the API. Usage reporting in the portal covers cloud usage only. On-prem deployments report usage differently. See [on-prem usage reporting](/deployments/usage-reporting/.md). ## View usage in the portal[​](#view-usage-in-the-portal "Direct link to View usage in the portal") 1. Go to **Usage** in the sidebar under **Workspace**. 2. Select a time range, such as **7D**, **30D**, or **Custom**. 3. Use the **All projects** dropdown to view usage for a single [project](/administration/projects.md) or across the workspace. Usage is shown in two charts, **Realtime** and **Batch**, matching the processing mode of each request. Each chart breaks usage down by model and reports total time used and request count. ## Usage by model[​](#usage-by-model "Direct link to Usage by model") The models shown depend on the processing mode: * **Realtime**: Standard and Enhanced. * **Batch**: Standard, Enhanced, and Melia 1. Medical and agent variants are reported under their underlying model, not as separate lines. For example, Batch usage on the medical variant appears under Enhanced. ## Track usage with the API[​](#track-usage-with-the-api "Direct link to Track usage with the API") To retrieve usage data programmatically, use the Usage API. See [Get usage statistics](/api-ref/batch/get-usage-statistics.md) in the API reference for endpoint details. The Usage API currently reports Batch usage only. ## Next steps[​](#next-steps "Direct link to Next steps") * [Billing](/administration/billing.md): understand how usage translates into charges. * [Projects](/administration/projects.md): organize work to track usage separately. --- # Workspaces concepts Understand how workspaces organize users, projects, billing, and access on the Speechmatics platform. A workspace is the top-level container for your team's use of Speechmatics. Members, projects, API keys, billing, and usage all sit inside a workspace. Anything you do in the portal and API happens within one workspace at a time. ## What a workspace contains[​](#what-a-workspace-contains "Direct link to What a workspace contains") A workspace groups four things: * **Members.** Users who can sign in to the workspace, each with an assigned role. * **Projects.** Organizational units inside the workspace. API keys are created within projects, and usage can be viewed per project. See [Projects](/administration/projects.md). * **API keys and management tokens.** API keys are scoped to a project; [management tokens](/administration/management-tokens.md) are scoped to the workspace. * **Billing and usage.** Pricing plan, payment method, credits, invoices, and aggregated usage belong to the workspace. ## Switch between workspaces[​](#switch-between-workspaces "Direct link to Switch between workspaces") A user can belong to more than one workspace. Each workspace is independent: switching changes which members, projects, keys, and billing details you see. To switch, open the avatar menu in the portal sidebar and select a workspace from the list. ## Roles[​](#roles "Direct link to Roles") Every member of a workspace has a role that determines what they can do. | Role | Capabilities | | ------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Admin | Full access to the workspace, including settings, billing, members, and projects. Can invite and remove members, create and delete projects and API keys, and rename the workspace. | | Member | Access to all workspace resources, including projects and usage. Cannot manage members, billing, or workspace settings. | For instructions on inviting members and changing roles, see [Manage members](/administration/manage-members.md). ## Workspaces and authentication[​](#workspaces-and-authentication "Direct link to Workspaces and authentication") Speechmatics uses [WorkOS AuthKit](https://workos.com/docs/authkit) for identity and access. A Speechmatics workspace corresponds to a WorkOS Organization, and the Admin and Member roles correspond to the WorkOS organization roles `admin` and `member`. Sign-in methods, including OAuth and SSO, are configured per workspace. When you verify a domain, new users signing up with an email address on that domain are added to your workspace automatically. See [Domain verification](/administration/domain-verification.md). ## Workspaces and regions[​](#workspaces-and-regions "Direct link to Workspaces and regions") A workspace is not tied to a single region. The [region](/administration/regions.md) you select determines which [endpoint processes](/get-started/authentication.md#supported-endpoints) a request, not where the workspace lives. Members of the same workspace can submit work to different regions. ## Next steps[​](#next-steps "Direct link to Next steps") * [Manage members](/administration/manage-members.md): invite users, change roles, and remove access. * [Domain verification](/administration/domain-verification.md): let users from your organization join your workspace automatically. * [Single sign-on (SSO)](/administration/sso.md): configure SSO for your workspace. --- # API reference Browse API reference for Speechmatics APIs Speechmatics offers flexible REST and Websocket APIs for transcription and voice AI applications. ## Quicklinks[​](#quicklinks "Direct link to Quicklinks") #### [Realtime transcription via websocket](/api-ref/realtime-transcription-websocket.md) #### [Create a transcription job by API](/api-ref/batch/create-a-new-job.md) #### [Get an API key](https://portal.speechmatics.com/settings/api-keys/) #### [Manage projects and API keys programmatically](/api-ref/management/create-an-api-key.md) #### [Download the Batch transcription API spec](/batch.yaml) --- # Create a new job ``` POST /jobs ``` Create a new job ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 201 * 400 * 401 * 403 * 429 * 500 OK Bad request Unauthorized Forbidden Rate Limited Internal Server Error --- # Delete a job ``` DELETE /jobs/:jobid ``` Delete a job and remove all associated resources. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 401 * 404 * 410 * 423 * 429 * 500 The job that was deleted. Unauthorized Not found Gone Locked Rate Limited Internal Server Error --- # Get job details ``` GET /jobs/:jobid ``` Get job details, including progress and any error reports. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 401 * 404 * 410 * 429 * 500 OK Unauthorized Not found Gone Rate Limited Internal Server Error --- # Get the log file for a job. ``` GET /jobs/:jobid/log ``` Get the log file for a job. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 401 * 404 * 410 * 429 * 500 * 501 OK Unauthorized Not Found Gone Rate Limited Internal Server Error Not Implemented --- # Get the transcript for a transcription job ``` GET /jobs/:jobid/transcript ``` Get the transcript for a transcription job ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 401 * 404 * 410 * 429 * 500 OK Unauthorized Not found Gone Rate Limited Internal Server Error --- # Get usage statistics ``` GET /usage ``` Get usage statistics ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 401 * 403 * 429 * 500 OK Unauthorized Forbidden Rate Limited Internal Server Error --- # List all jobs ``` GET /jobs ``` List all jobs ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 401 * 422 * 429 * 500 OK Unauthorized Unprocessable Entity Rate Limited Internal Server Error --- Version: 2.0.0 # Speechmatics ASR REST API The Speechmatics Automatic Speech Recognition REST API is used to submit ASR jobs and receive the results. The supported job type is transcription of audio files. ## Authentication[​](#authentication "Direct link to Authentication") * API Key: bearerAuth Bearer token authentication. Use the format: Bearer $API\_KEY\_OR\_JWT | Security Scheme Type: | apiKey | | ---------------------- | ------------- | | Header parameter name: | Authorization | ### Contact --- # Create an API key ``` POST /api-keys ``` Create an API key in your workspace, authenticated with a [management token](/administration/management-tokens.md). This endpoint issues two kinds of key: * **Permanent key** — a long-lived key for a project. Set `project_id` (and optionally `name`) and omit `ttl`. Use this for a normal project API key. * **Temporary key** — a short-lived key that expires automatically. Set `ttl` to the lifetime in seconds. Use temporary keys to authenticate end-user requests without exposing a long-lived key. The `au` region supports Batch only, so it cannot be combined with `type=rt`. ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 201 * 403 * 404 * 410 * 500 Created Forbidden, project has maximum number of api keys already Not Found Gone Internal Server Error --- # Delete a project ``` DELETE /projects/:project_id ``` Delete a project ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 204 * 403 * 404 * 410 * 500 Project deleted. Forbidden. Most probably due to outstanding payments. Project was not found Gone Internal Server Error --- # Delete an API key ``` DELETE /api-keys/:apikey_id ``` Delete an API key ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 410 * 500 OK Not Found Gone Internal Server Error --- # Get a project by ID ``` GET /projects/:project_id ``` Get a project by ID ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 400 * 401 * 404 * 410 * 500 OK Bad Request Unauthorized Project was not found Project Gone Internal Server Error --- # Get all API keys ``` GET /api-keys ``` Get all API keys ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 403 * 404 * 410 * 500 OK Forbidden Project not found Gone Internal Server Error --- # Get all projects ``` GET /projects ``` Get all projects ## Responses[​](#responses "Direct link to Responses") * 200 * 404 * 410 * 500 OK Not Found Gone Internal Server Error --- Version: 1.0.0 # Management API The Management API lets you manage projects and API keys in your workspace programmatically. Authenticate each request with a [management token](/administration/management-tokens.md), which is scoped to the workspace. Use it to automate workspace administration, such as creating a project and issuing its API key from a CI pipeline. To get started, [create a management token](/administration/management-tokens.md), then call the endpoints below. ### Contact --- # Post a new project ``` POST /projects ``` Post a new project ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 201 * 410 * 500 Created Gone Internal Server Error --- # Update a project ``` PUT /projects/:project_id ``` Update a project ## Request[​](#request "Direct link to Request") ## Responses[​](#responses "Direct link to Responses") * 200 * 404 OK Not Found --- # Realtime API Reference ``` GET wss://eu.rt.speechmatics.com/v2/ ``` You can either pin to a specific region for data residency or connect to `wss://global.rt.speechmatics.com/v2/`, which automatically routes each connection to the nearest region for lowest latency. See [Supported endpoints](/get-started/authentication.md#supported-endpoints). ## Protocol overview[​](#protocol-overview "Direct link to Protocol overview") A basic Realtime session will have the following message exchanges: #### Browser based transcription[​](#browser-based-transcription "Direct link to Browser based transcription") When starting a Realtime transcription session **in the browser**, [temporary keys](/get-started/authentication.md#temporary-keys) should be used to avoid exposing your long-lived API key. To do so, you must provide the temporary key as a part of a query parameter. This is due to a browser limitation. For example: ``` wss://eu.rt.speechmatics.com/v2?jwt= ``` ### Handshake responses[​](#handshake-responses "Direct link to Handshake responses") **Successful Response** * `101 Switching Protocols` - Switch to WebSocket protocol Here is an example for a successful WebSocket handshake: ``` GET /v2/ HTTP/1.1 Host: eu.rt.speechmatics.com Upgrade: websocket Connection: Upgrade Sec-WebSocket-Key: ujRTbIaQsXO/0uCbjjkSZQ== Sec-WebSocket-Version: 13 Sec-WebSocket-Extensions: permessage-deflate; client_max_window_bits Authorization: Bearer wmz9fkLJM6U5NdyaG3HLHybGZj65PXp User-Agent: Python/3.8 websockets/8.1 ``` A successful response should look like: ``` HTTP/1.1 101 Switching Protocols Server: nginx/1.17.8 Date: Wed, 06 Jan 2021 11:01:05 GMT Connection: upgrade Upgrade: WebSocket Sec-WebSocket-Accept: 87kiC/LI5WgXG52nSylnfXdz260= ``` **Malformed request** A malformed handshake request will result in one of the following HTTP responses: * `400 Bad Request` * `401 Unauthorized` - when the API key is not valid * `405 Method Not Allowed` - when the request method is not GET **Client Retry** Following a successful handshake and switch to the WebSocket protocol, the client could receive an immediate error message and WebSocket close handshake from the server. For the following errors only, we recommend adding a client retry interval of at least 5-10 seconds: * `4005 quota_exceeded` * `4013 job_error` * `1011 internal_error` ## Message handling[​](#message-handling "Direct link to Message handling") Each message that the **Server** accepts is a stringified JSON object with the following fields: * `message` (String): The name of the message we are sending. Any other fields depend on the value of the `message` and are described below. The messages sent by the **Server** to a **Client** are stringified JSON objects as well. The only exception is a binary message sent from the **Client** to the **Server** containing a chunk of audio which will be referred to as `AddAudio`. The following values of the `message` field are supported: ## Sent messages[​](#sent-messages "Direct link to Sent messages") ### StartRecognition[​](#startrecognition "Direct link to startrecognition") See example Initiates a new recognition session. **message**required Constant value: `StartRecognition` **audio\_format** objectrequired oneOf * Raw * File Raw audio samples, described by the following additional mandatory fields: **type**required Constant value: `raw` **encoding**stringrequired Possible values: \[`pcm_f32le`, `pcm_s16le`, `mulaw`] **sample\_rate**integerrequired The sample rate of the audio in Hz. **Example**: `{"type":"raw","encoding":"pcm_s16le","sample_rate":44100}` Choose this option to send audio encoded in a recognized format. The AddAudio messages have to provide all the file contents, including any headers. The file is usually not accepted all at once, but segmented into reasonably sized messages. Note: Only the following formats are supported: `wav`, `mp3`, `aac`, `ogg`, `mpeg`, `amr`, `m4a`, `mp4`, `flac` **type**required Constant value: `file` **transcription\_config** objectrequired Contains configuration for this recognition session. **language**stringrequired Language model to process the audio input, normally specified as an ISO language code. The value must be consistent with the language code used in the API endpoint URL. **Example: **`en` **domain**string Request a specialized model based on 'language' but optimized for a particular field, e.g. `finance` or `medical`. **output\_locale**string Configure locale for outputted transcription. See [output formatting](https://docs.speechmatics.com/speech-to-text/formatting#output-locale). Possible values: `non-empty` **additional\_vocab** object\[] Configure [custom dictionary](https://docs.speechmatics.com/speech-to-text/features/custom-dictionary). Default is an empty list. You should be aware that there is a performance penalty (latency degradation and memory increase) from using `additional_vocab`, especially if you use a large word list. When initializing a session that uses `additional_vocab` in the config, you should expect a delay of up to 15 seconds (depending on the size of the list). * Array \[ oneOf * String * Object Pass a non-empty string to add a single word to the dictionary. ****string Pass a non-empty string to add a single word to the dictionary. Possible values: `non-empty` Pass an object to add a single word to the dictionary, with an array of words which it sounds like. **content**stringrequired Possible values: `non-empty` **sounds\_like**string\[] Possible values: `>= 1` * ] **diarization**string Set to `speaker` to apply [Speaker Diarization](https://docs.speechmatics.com/speech-to-text/features/diarization) to the audio. Possible values: \[`none`, `speaker`, `channel`, `channel_and_speaker`] **Default value: **`none` **max\_delay**number This is the delay in seconds between the end of a spoken word and returning the Final transcript results. See [Latency](https://docs.speechmatics.com/speech-to-text/realtime/output#latency) for more details Possible values: `>= 0.7` and `<= 4` **Default value: **`4` **max\_delay\_mode**string This allows some additional time for [Smart Formatting](https://docs.speechmatics.com/speech-to-text/formatting#smart-formatting). Possible values: \[`flexible`, `fixed`] **Default value: **`flexible` **speaker\_diarization\_config** object **max\_speakers**integer Configure the maximum number of speakers to detect. See [Max Speakers](http://docs.speechmatics.com/speech-to-text/features/diarization#max-speakers). Possible values: `>= 2` **prefer\_current\_speaker**boolean When set to `true`, reduces the likelihood of incorrectly switching between similar sounding speakers. See [Prefer Current Speaker](https://docs.speechmatics.com/speech-to-text/features/diarization#prefer-current-speaker). **Default value: **`false` **speaker\_sensitivity**float Possible values: `>= 0` and `<= 1` **get\_speakers**boolean If true, speaker identifiers will be returned at the end of transcript. **speakers** object\[] Use this option to provide speaker labels linked to their speaker identifiers. When passed, the transcription system will tag spoken words in the transcript with the provided speaker labels whenever any of the specified speakers is detected in the audio. A maximum of 50 speakers identifiers across all speakers can be provided. * Array \[ **label**stringrequired Speaker label, which must not match the format used internally (e.g. S1, S2, etc) Possible values: `non-empty` **speaker\_identifiers**bytes\[]required Possible values: `>= 1` * ] **audio\_filtering\_config** object Puts a lower limit on the volume of processed audio by using the `volume_threshold` setting. See [Audio Filtering](https://docs.speechmatics.com/speech-to-text/features/audio-filtering). **volume\_threshold**float Possible values: `>= 0` and `<= 100` **transcript\_filtering\_config** object **remove\_disfluencies**boolean When set to `true`, removes disfluencies from the transcript. See [Removing disfluencies](https://docs.speechmatics.com/speech-to-text/formatting#removing-disfluencies) **replacements** object\[] A list of replacement rules to apply to the transcript. Each rule consists of a pattern to match and a replacement string. See [Word replacement](https://docs.speechmatics.com/speech-to-text/formatting#word-replacement) * Array \[ **from**stringrequired **to**stringrequired * ] **enable\_partials**boolean Whether or not to send Partials (i.e. `AddPartialTranslation` messages) as well as Finals (i.e. `AddTranslation` messages) See [Partial transcripts](https://docs.speechmatics.com/speech-to-text/realtime/output#partial-transcripts). **Default value: **`false` **enable\_entities**boolean **Default value: **`true` **operating\_point**stringdeprecated **Deprecated**: Kept for backward compatibility only. Use `model` instead going forward. Possible values: \[`standard`, `enhanced`, `melia-1`] **model**string Which model you wish to use. See [Models](http://docs.speechmatics.com/speech-to-text/models) for more details. Possible values: \[`standard`, `enhanced`, `melia-1`] **Default value: **`standard` **punctuation\_overrides** object Options for controlling punctuation in the output transcripts. See [Punctuation Settings](https://docs.speechmatics.com/speech-to-text/formatting#punctuation) **permitted\_marks**string\[] The punctuation marks which the client is prepared to accept in transcription output, or the special value 'all' (the default). Unsupported marks are ignored. This value is used to guide the transcription process. Possible values: Value must match regular expression `^(.|all)$` **sensitivity**float Ranges between zero and one. Higher values will produce more punctuation. The default is 0.5. Possible values: `>= 0` and `<= 1` **conversation\_config** object This mode will detect when a speaker has stopped talking. The `end_of_utterance_silence_trigger` is the time in seconds after which the server will assume that the speaker has finished speaking, and will emit an `EndOfUtterance` message. A value of 0 disables the feature. **end\_of\_utterance\_silence\_trigger**float Possible values: `>= 0` and `<= 2` **Default value: **`0` **channel\_diarization\_labels**string\[] **translation\_config** object Specifies various configuration values for translation. All fields except `target_languages` are optional, using default values when omitted. **target\_languages**string\[]required List of languages to translate to from the source transcription `language`. Specified as an [ISO Language Code](https://docs.speechmatics.com/speech-to-text/languages). **enable\_partials**boolean Whether or not to send Partials (i.e. `AddPartialTranslation` messages) as well as Finals (i.e. `AddTranslation` messages). **Default value: **`false` **audio\_events\_config** object Contains configuration for [Audio Events](https://docs.speechmatics.com/speech-to-text/features/audio-events) **types**string\[] List of [Audio Event types](https://docs.speechmatics.com/speech-to-text/features/audio-events#supported-audio-events) to enable. ### AddAudio[​](#addaudio "Direct link to addaudio") A binary chunk of audio. The server confirms receipt by sending an AudioAdded message. **string**binary ### AddChannelAudio[​](#addchannelaudio "Direct link to addchannelaudio") Audio belonging to a specific channel. **message**required Constant value: `AddChannelAudio` **channel**stringrequired The channel identifier to which the audio belongs. **data**stringrequired The audio data in base64 format. ### EndOfStream[​](#endofstream "Direct link to endofstream") Declares that the client has no more audio to send. **message**required Constant value: `EndOfStream` **last\_seq\_no**integerrequired ### EndOfChannel[​](#endofchannel "Direct link to endofchannel") Declares that the channel has no more audio to send. **message**required Constant value: `EndOfChannel` **channel**stringrequired The channel identifier to which the audio belongs. **last\_seq\_no**integerrequired ### ForceEndOfUtterance[​](#forceendofutterance "Direct link to forceendofutterance") Requests a finalized transcript. **message**required Constant value: `ForceEndOfUtterance` **channel**string The channel to request finalized transcript from. This field is only seen in multichannel. **timestamp**float Timestamp of the audio data that corresponds to the force end of utterance request. It's the number of seconds since the beginning of the audio. Possible values: `>= 0` ### SetRecognitionConfig[​](#setrecognitionconfig "Direct link to setrecognitionconfig") Allows the client to re-configure the recognition session. The language field can be included in SetRecognitionConfig, but its value must match the language specified in the initial StartRecognition message. Attempting to change the language or any other transcription configuration parameter not listed below will result in an error. **message**required Constant value: `SetRecognitionConfig` **transcription\_config** objectrequired Contains configuration for this recognition session. **language**string Language model to process the audio input, normally specified as an ISO language code. The value must be consistent with the language code used in the API endpoint URL. **Example: **`en` **max\_delay**number This is the delay in seconds between the end of a spoken word and returning the Final transcript results. See [Latency](https://docs.speechmatics.com/speech-to-text/realtime/output#latency) for more details Possible values: `>= 0.7` and `<= 4` **Default value: **`4` **max\_delay\_mode**string This allows some additional time for [Smart Formatting](https://docs.speechmatics.com/speech-to-text/formatting#smart-formatting). Possible values: \[`flexible`, `fixed`] **Default value: **`flexible` **audio\_filtering\_config** object Puts a lower limit on the volume of processed audio by using the `volume_threshold` setting. See [Audio Filtering](https://docs.speechmatics.com/speech-to-text/features/audio-filtering). **volume\_threshold**float Possible values: `>= 0` and `<= 100` **enable\_partials**boolean Whether or not to send Partials (i.e. `AddPartialTranslation` messages) as well as Finals (i.e. `AddTranslation` messages) See [Partial transcripts](https://docs.speechmatics.com/speech-to-text/realtime/output#partial-transcripts). **Default value: **`false` **conversation\_config** object This mode will detect when a speaker has stopped talking. The `end_of_utterance_silence_trigger` is the time in seconds after which the server will assume that the speaker has finished speaking, and will emit an `EndOfUtterance` message. A value of 0 disables the feature. **end\_of\_utterance\_silence\_trigger**float Possible values: `>= 0` and `<= 2` **Default value: **`0` ### GetSpeakers[​](#getspeakers "Direct link to getspeakers") Requests any detected speaker identifiers to be returned. **message**required Constant value: `GetSpeakers` **final**boolean Optional. This flag controls when speaker identifiers are returned. Defaults to false if omitted. When false, multiple GetSpeakers requests can be sent during transcription, each returning the speaker identifiers generated so far. To reduce the chance of empty results, send requests after at least one TranscriptAdded message is received to make sure that the server has processed some audio. When true, speaker identifiers are returned only once at the end of the transcription, regardless of how many final: true requests are sent. Even with final: true requests, you can still send final: false requests to receive intermediate speaker identifier updates. ## Received messages[​](#received-messages "Direct link to Received messages") ### RecognitionStarted[​](#recognitionstarted "Direct link to recognitionstarted") See example Server response to StartRecognition, acknowledging that a recognition session has started. **message**required Constant value: `RecognitionStarted` **orchestrator\_version**string **id**string **language\_pack\_info** object Properties of the language pack. **language\_description**string Full descriptive name of the language, e.g. 'Japanese'. **word\_delimiter**stringrequired The character to use to separate words. **writing\_direction**string The direction that words in the language should be written and read in. Possible values: \[`left-to-right`, `right-to-left`] **itn**boolean Whether or not ITN (inverse text normalization) is available for the language pack. **adapted**boolean Whether or not language model adaptation has been applied to the language pack. **channel\_ids**string\[] ### AudioAdded[​](#audioadded "Direct link to audioadded") Server response to AddAudio, indicating that audio has been added successfully. When clients send audio faster than real-time, the server may read data slower than it's sent. If binary `AddAudio` messages exceed the server's internal buffer, the server will process other WebSocket messages until buffer space is available. Clients receive `AudioAdded` responses only after binary data is read. This can fill TCP buffers, potentially causing WebSocket write failures and connection closure [with prejudice](https://websockets.spec.whatwg.org#the-closeevent-interface). Clients can monitor the WebSocket's [`bufferedAmount`](https://www.w3.org/TR/websockets#dom-websocket-bufferedamount) attribute to prevent this. **message**required Constant value: `AudioAdded` **seq\_no**integerrequired ### ChannelAudioAdded[​](#channelaudioadded "Direct link to channelaudioadded") Server response to AddChannelAudio, indicating that audio has been added successfully. **message**required Constant value: `ChannelAudioAdded` **seq\_no**integerrequired **channel**stringrequired ### AddPartialTranscript[​](#addpartialtranscript "Direct link to addpartialtranscript") A partial transcript is a transcript that can be changed in a future `AddPartialTranscript` as more words are spoken until the `AddTranscript` **Final** message is sent for that audio. Partials will only be sent if `transcription_config.enable_partials` is set to `true` in the `StartRecognition` message. The message structure is the same as `AddTranscript`, with a few [limitations](https://docs.speechmatics.com/speech-to-text/realtime/output#partial-transcripts). For `AddPartialTranscript` messages the `confidence` field for `alternatives` has no meaning and should not be relied on. **message**required Constant value: `AddPartialTranscript` **format**string Speechmatics JSON output format version number. **Example: **`2.1` **metadata** objectrequired **start\_time**floatrequired **end\_time**floatrequired **transcript**stringrequired The entire transcript contained in the segment in text format. Providing the entire transcript here is designed for ease of consumption; we have taken care of all the necessary formatting required to concatenate the transcription results into a block of text. This transcript lacks the detailed information however which is contained in the `results` field of the message - such as the timings and confidences for each word. **results** object\[]required * Array \[ **type**stringrequired Possible values: \[`word`, `punctuation`, `entity`] **start\_time**floatrequired **end\_time**floatrequired **attaches\_to**string Possible values: \[`next`, `previous`, `none`, `both`] **is\_eos**boolean **alternatives** object\[] * Array \[ **content**stringrequired A word or punctuation mark. **confidence**floatrequired A confidence score assigned to the alternative. Ranges from 0.0 (least confident) to 1.0 (most confident). **language**string The language that the alternative word is assumed to be spoken in. Currently, this will always be equal to the language that was requested in the initial `StartRecognition` message. **display** object Information about how the word/symbol should be displayed. **direction**stringrequired Either `ltr` for words that should be displayed left-to-right, or `rtl` vice versa. Possible values: \[`ltr`, `rtl`] **speaker**string Label indicating who said that word. Only set if [diarization](https://docs.speechmatics.com/speech-to-text/features/diarization) is enabled. **tags**string\[] This is a set list of profanities and disfluencies respectively that cannot be altered by the end user. `[disfluency]` is present in the [supported languages](https://docs.speechmatics.com/speech-to-text/formatting#supported-languages-for-disfluencies), and `[profanity]` is present in English, Spanish, and Italian Possible values: \[`disfluency`, `profanity`] * ] **volume**float Possible values: `>= 0` and `<= 100` **entity\_class**string For 'entity' results only, the class the entity has been formatted as. Examples: 'date', 'money', 'number' **spoken\_form** object\[] For 'entity' results only, the spoken\_form is the transcript of the individual words directly spoken. * Array \[ **alternatives** object\[]required * Array \[ **content**stringrequired A word or punctuation mark. **confidence**floatrequired A confidence score assigned to the alternative. Ranges from 0.0 (least confident) to 1.0 (most confident). **language**string The language that the alternative word is assumed to be spoken in. Currently, this will always be equal to the language that was requested in the initial `StartRecognition` message. **display** object Information about how the word/symbol should be displayed. **direction**stringrequired Either `ltr` for words that should be displayed left-to-right, or `rtl` vice versa. Possible values: \[`ltr`, `rtl`] **speaker**string Label indicating who said that word. Only set if [diarization](https://docs.speechmatics.com/speech-to-text/features/diarization) is enabled. **tags**string\[] This is a set list of profanities and disfluencies respectively that cannot be altered by the end user. `[disfluency]` is present in the [supported languages](https://docs.speechmatics.com/speech-to-text/formatting#supported-languages-for-disfluencies), and `[profanity]` is present in English, Spanish, and Italian Possible values: \[`disfluency`, `profanity`] * ] **end\_time**floatrequired **start\_time**floatrequired **type**stringrequired Possible values: \[`word`, `punctuation`] * ] **written\_form** object\[] For 'entity' results only, the written\_form is a standardized form of the spoken words. It contains the formatted entity split into individual words. * Array \[ **alternatives** object\[]required * Array \[ **content**stringrequired A word or punctuation mark. **confidence**floatrequired A confidence score assigned to the alternative. Ranges from 0.0 (least confident) to 1.0 (most confident). **language**string The language that the alternative word is assumed to be spoken in. Currently, this will always be equal to the language that was requested in the initial `StartRecognition` message. **display** object Information about how the word/symbol should be displayed. **direction**stringrequired Either `ltr` for words that should be displayed left-to-right, or `rtl` vice versa. Possible values: \[`ltr`, `rtl`] **speaker**string Label indicating who said that word. Only set if [diarization](https://docs.speechmatics.com/speech-to-text/features/diarization) is enabled. **tags**string\[] This is a set list of profanities and disfluencies respectively that cannot be altered by the end user. `[disfluency]` is present in the [supported languages](https://docs.speechmatics.com/speech-to-text/formatting#supported-languages-for-disfluencies), and `[profanity]` is present in English, Spanish, and Italian Possible values: \[`disfluency`, `profanity`] * ] **end\_time**floatrequired **start\_time**floatrequired **type**stringrequired Possible values: \[`word`] * ] * ] **channel**string The channel identifier to which the audio belongs. This field is only seen in multichannel. **forced**boolean Whether this message was triggered by a `ForceEndOfUtterance` message. This field is only seen on forced messages, where its value is `true`. ### AddTranscript[​](#addtranscript "Direct link to addtranscript") Contains the final transcript of a part of the audio that the client has sent. **message**required Constant value: `AddTranscript` **format**string Speechmatics JSON output format version number. **Example: **`2.1` **metadata** objectrequired **start\_time**floatrequired **end\_time**floatrequired **transcript**stringrequired The entire transcript contained in the segment in text format. Providing the entire transcript here is designed for ease of consumption; we have taken care of all the necessary formatting required to concatenate the transcription results into a block of text. This transcript lacks the detailed information however which is contained in the `results` field of the message - such as the timings and confidences for each word. **results** object\[]required * Array \[ **type**stringrequired Possible values: \[`word`, `punctuation`, `entity`] **start\_time**floatrequired **end\_time**floatrequired **attaches\_to**string Possible values: \[`next`, `previous`, `none`, `both`] **is\_eos**boolean **alternatives** object\[] * Array \[ **content**stringrequired A word or punctuation mark. **confidence**floatrequired A confidence score assigned to the alternative. Ranges from 0.0 (least confident) to 1.0 (most confident). **language**string The language that the alternative word is assumed to be spoken in. Currently, this will always be equal to the language that was requested in the initial `StartRecognition` message. **display** object Information about how the word/symbol should be displayed. **direction**stringrequired Either `ltr` for words that should be displayed left-to-right, or `rtl` vice versa. Possible values: \[`ltr`, `rtl`] **speaker**string Label indicating who said that word. Only set if [diarization](https://docs.speechmatics.com/speech-to-text/features/diarization) is enabled. **tags**string\[] This is a set list of profanities and disfluencies respectively that cannot be altered by the end user. `[disfluency]` is present in the [supported languages](https://docs.speechmatics.com/speech-to-text/formatting#supported-languages-for-disfluencies), and `[profanity]` is present in English, Spanish, and Italian Possible values: \[`disfluency`, `profanity`] * ] **volume**float Possible values: `>= 0` and `<= 100` **entity\_class**string For 'entity' results only, the class the entity has been formatted as. Examples: 'date', 'money', 'number' **spoken\_form** object\[] For 'entity' results only, the spoken\_form is the transcript of the individual words directly spoken. * Array \[ **alternatives** object\[]required * Array \[ **content**stringrequired A word or punctuation mark. **confidence**floatrequired A confidence score assigned to the alternative. Ranges from 0.0 (least confident) to 1.0 (most confident). **language**string The language that the alternative word is assumed to be spoken in. Currently, this will always be equal to the language that was requested in the initial `StartRecognition` message. **display** object Information about how the word/symbol should be displayed. **direction**stringrequired Either `ltr` for words that should be displayed left-to-right, or `rtl` vice versa. Possible values: \[`ltr`, `rtl`] **speaker**string Label indicating who said that word. Only set if [diarization](https://docs.speechmatics.com/speech-to-text/features/diarization) is enabled. **tags**string\[] This is a set list of profanities and disfluencies respectively that cannot be altered by the end user. `[disfluency]` is present in the [supported languages](https://docs.speechmatics.com/speech-to-text/formatting#supported-languages-for-disfluencies), and `[profanity]` is present in English, Spanish, and Italian Possible values: \[`disfluency`, `profanity`] * ] **end\_time**floatrequired **start\_time**floatrequired **type**stringrequired Possible values: \[`word`, `punctuation`] * ] **written\_form** object\[] For 'entity' results only, the written\_form is a standardized form of the spoken words. It contains the formatted entity split into individual words. * Array \[ **alternatives** object\[]required * Array \[ **content**stringrequired A word or punctuation mark. **confidence**floatrequired A confidence score assigned to the alternative. Ranges from 0.0 (least confident) to 1.0 (most confident). **language**string The language that the alternative word is assumed to be spoken in. Currently, this will always be equal to the language that was requested in the initial `StartRecognition` message. **display** object Information about how the word/symbol should be displayed. **direction**stringrequired Either `ltr` for words that should be displayed left-to-right, or `rtl` vice versa. Possible values: \[`ltr`, `rtl`] **speaker**string Label indicating who said that word. Only set if [diarization](https://docs.speechmatics.com/speech-to-text/features/diarization) is enabled. **tags**string\[] This is a set list of profanities and disfluencies respectively that cannot be altered by the end user. `[disfluency]` is present in the [supported languages](https://docs.speechmatics.com/speech-to-text/formatting#supported-languages-for-disfluencies), and `[profanity]` is present in English, Spanish, and Italian Possible values: \[`disfluency`, `profanity`] * ] **end\_time**floatrequired **start\_time**floatrequired **type**stringrequired Possible values: \[`word`] * ] * ] **channel**string The channel identifier to which the audio belongs. This field is only seen in multichannel. **forced**boolean Whether this message was triggered by a `ForceEndOfUtterance` message. This field is only seen on forced messages, where its value is `true`. ### AddPartialTranslation[​](#addpartialtranslation "Direct link to addpartialtranslation") Contains a work-in-progress translation of a part of the audio that the client has sent. **message**required Constant value: `AddPartialTranslation` **format**string Speechmatics JSON output format version number. **Example: **`2.1` **language**stringrequired Language translation relates to given as an ISO language code. **results** object\[]required * Array \[ **content**stringrequired **start\_time**floatrequired The start time (in seconds) of the original transcribed audio segment **end\_time**floatrequired The end time (in seconds) of the original transcribed audio segment **speaker**string The speaker that uttered the speech if speaker diarization is enabled * ] ### AddTranslation[​](#addtranslation "Direct link to addtranslation") Contains the final translation of a part of the audio that the client has sent. **message**required Constant value: `AddTranslation` **format**string Speechmatics JSON output format version number. **Example: **`2.1` **language**stringrequired Language translation relates to given as an ISO language code. **results** object\[]required * Array \[ **content**stringrequired **start\_time**floatrequired The start time (in seconds) of the original transcribed audio segment **end\_time**floatrequired The end time (in seconds) of the original transcribed audio segment **speaker**string The speaker that uttered the speech if speaker diarization is enabled * ] ### EndOfTranscript[​](#endoftranscript "Direct link to endoftranscript") Server response to `EndOfStream`, after the server has finished sending all AddTranscript messages. **message**required Constant value: `EndOfTranscript` ### AudioEventStarted[​](#audioeventstarted "Direct link to audioeventstarted") Start of an audio event detected. **message**required Constant value: `AudioEventStarted` **event** objectrequired **type**stringrequired The type of audio event that has started or ended. See our list of [supported Audio Event types](https://docs.speechmatics.com/speech-to-text/features/audio-events#supported-audio-events). **start\_time**floatrequired The time (in seconds) of the audio corresponding to the beginning of the audio event. **confidence**floatrequired A confidence score assigned to the audio event. Ranges from 0.0 (least confident) to 1.0 (most confident). Possible values: `>= 0` and `<= 1` **channel**string The channel identifier to which the audio belongs. This field is only seen in multichannel. ### AudioEventEnded[​](#audioeventended "Direct link to audioeventended") End of an audio event detected. **message**required Constant value: `AudioEventEnded` **event** objectrequired **type**stringrequired The type of audio event that has started or ended. See our list of [supported Audio Event types](https://docs.speechmatics.com/speech-to-text/features/audio-events#supported-audio-events). **end\_time**floatrequired **channel**string The channel identifier to which the audio belongs. This field is only seen in multichannel. ### EndOfUtterance[​](#endofutterance "Direct link to endofutterance") Indicates the end of an utterance, triggered by a configurable period of non-speech. The message is sent when no speech has been detected for a short period of time, configurable by the `end_of_utterance_silence_trigger` parameter in `conversation_config` (see [End Of Utterance](https://docs.speechmatics.com/speech-to-text/realtime/turn-detection#configuration)). Like punctuation, an `EndOfUtterance` has zero duration. **message**required Constant value: `EndOfUtterance` **metadata** objectrequired **start\_time**float The time (in seconds) that the end of utterance was detected. **end\_time**float The time (in seconds) that the end of utterance was detected. **channel**string The channel identifier to which the EndOfUtterance message belongs. This field is only seen in multichannel. **forced**boolean Whether this message was triggered by a `ForceEndOfUtterance` message. This field is only seen on forced messages, where its value is `true`. ### Info[​](#info "Direct link to info") Additional information sent from the server to the client. **message**required Constant value: `Info` **type**stringrequired The following are the possible info types: | Info Type | Description | | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `recognition_quality` | Informs the client what particular quality-based model is used to handle the recognition. Sent to the client immediately after the WebSocket handshake is completed. | | `concurrent_session_usage` | Informs the client of their quota for concurrent sessions and how much of it they are using. Sent to the client immediately after the WebSocket handshake is completed. | Possible values: \[`recognition_quality`, `concurrent_session_usage`] **reason**stringrequired **code**integer **seq\_no**integer **quality**string Only set when `type` is `recognition_quality`. Quality-based model name. It is one of "telephony", "broadcast". The model is selected automatically, for high-quality audio (12kHz+) the broadcast model is used, for lower quality audio the telephony model is used. **usage**number Only set when `type` is `concurrent_session_usage`. Indicates the current usage (number of active concurrent sessions). **quota**number Only set when `type` is `concurrent_session_usage`. Indicates the current quota (maximum number of concurrent sessions allowed). **last\_updated**string Only set when `type` is `concurrent_session_usage`. Indicates the timestamp of the most recent usage update, in the format `YYYY-MM-DDTHH:MM:SSZ` (UTC). This value is updated even when usage exceeds the quota, as it represents the most recent known data. In some cases, it may be empty or outdated due to internal errors preventing successful update. **Example: **`2025-03-25T08:45:31Z` **region**string Only set when `type` is `concurrent_session_usage`. The region that reported this usage, for example `eu`. **Example: **`eu` ### Warning[​](#warning "Direct link to warning") Warning messages sent from the server to the client. **message**required Constant value: `Warning` **type**stringrequired The following are the possible warning types: | Warning Type | Description | | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `duration_limit_exceeded` | The maximum allowed duration of a single utterance to process has been exceeded. Any `AddAudio` messages received that exceed this limit are confirmed with `AudioAdded`, but are ignored by the transcription engine. Exceeding the limit triggers the same mechanism as receiving an `EndOfStream` message, so the Server will eventually send an `EndOfTranscript` message and suspend. | | `unsupported_translation_pair` | One of the requested translation target languages is unsupported (given the source audio language). The error message specifies the unsupported language pair. | | `idle_timeout` | Informs that the session is approaching the idle duration limit (no audio data sent within the last hour), with a `reason` of the form:`Session will timeout in {time_remaining}m due to inactivity, no audio sent within the last {time_elapsed}m`Currently the server will send messages at 15, 10 and 5m prior to timeout, and will send a final error message on timeout, before closing the connection with the code 1008. (see [Realtime limits](https://docs.speechmatics.com/speech-to-text/realtime/limits) for more information). | | `session_timeout` | Informs that the session is approaching the max session duration limit (maximum session duration of 48 hours), with a `reason` of the form:`Session will timeout in {time_remaining}m due to max duration, session has been active for {time_elapsed}m`Currently the server will send messages at 45, 30 and 15m prior to timeout, and will send a final error message on timeout, before closing the connection with the code 1008. (see [Realtime limits](https://docs.speechmatics.com/speech-to-text/realtime/limits) for more information). | | `empty_translation_target_list` | No supported translation target languages specified. Translation will not run. | | `add_audio_after_eos` | Protocol specification doesn't allow adding audio after `EndOfStream` has been received. Any \`AddAudio messages after this, will be ignored. | | `speaker_id` | Informs the client about any speaker ID related issues. | Possible values: \[`duration_limit_exceeded`, `unsupported_translation_pair`, `idle_timeout`, `session_timeout`, `empty_translation_target_list`, `add_audio_after_eos`, `speaker_id`] **reason**stringrequired **code**integer **seq\_no**integer **duration\_limit**number Only set when `type` is `duration_limit_exceeded`. Indicates the limit that was exceeded (in seconds). ### Error[​](#error "Direct link to error") Error messages sent from the server to the client. After any error, the transcription is terminated and the connection is closed. **message**required Constant value: `Error` **type**stringrequired The following are the possible error types: | Error Type | Description | | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `invalid_message` | The message received was not understood. | | `invalid_model` | Unable to use the model for the recognition. This can happen if the language is not supported at all, or is not available for the user. | | `invalid_language` | The requested language is not valid or is not supported. | | `invalid_config` | The config received contains some wrong or unsupported fields, or too many translation target languages were requested. | | `invalid_audio_type` | Audio type is not supported, is deprecated, or the `audio_type` is malformed. | | `invalid_output_format` | Output format is not supported, is deprecated, or the `output_format` is malformed. | | `not_authorised` | User was not recognised, or the API key provided is not valid. | | `not_allowed` | User is not allowed to use this message (is not allowed to perform the action the message would invoke). | | `job_error` | Unable to do any work on this job, the server might have timed out etc. | | `protocol_error` | Message received was syntactically correct, but could not be accepted due to protocol limitations. This is usually caused by messages sent in the wrong order. | | `quota_exceeded` | Maximum number of concurrent connections allowed for the contract has been reached | | `timelimit_exceeded` | Usage quota for the contract has been reached | | `idle_timeout` | Idle duration limit was reached (no audio data sent within the last hour), a closing handshake with code 1008 follows this in-band error. | | `session_timeout` | Max session duration was reached (maximum session duration of 48 hours), a closing handshake with code 1008 follows this in-band error. | | `unknown_error` | An error that did not fit any of the types above. | `invalid_message`, `protocol_error` and `unknown_error` can be triggered as a response to any type of messages. Possible values: \[`invalid_message`, `invalid_model`, `invalid_language`, `invalid_config`, `invalid_audio_type`, `invalid_output_format`, `not_authorised`, `not_allowed`, `job_error`, `protocol_error`, `quota_exceeded`, `timelimit_exceeded`, `idle_timeout`, `session_timeout`, `unknown_error`] **reason**stringrequired **code**integer **seq\_no**integer ### SpeakersResult[​](#speakersresult "Direct link to speakersresult") Server response to GetSpeakers request returning any detected speaker identifiers. **message**required Constant value: `SpeakersResult` **speakers** object\[]required * Array \[ **label**stringrequired Speaker label. Possible values: `non-empty` **speaker\_identifiers**bytes\[]required Possible values: `>= 1` * ] ## Websocket errors[​](#websocket-errors "Direct link to Websocket errors") In the Realtime SaaS, an in-band error message can be followed by a WebSocket close message. The table below shows the possible WebSocket close codes and associated error types. The error types are provided in the payload of the close message. | WebSocket Close Code | WebSocket Close Payload | | -------------------- | ----------------------- | | 1003 | `protocol_error` | | 1008 | `policy_violation` | | 1011 | `internal_error` | | 4001 | `not_authorised` | | 4003 | `not_allowed` | | 4004 | `invalid_model` | | 4005 | `quota_exceeded` | | 4006 | `timelimit_exceeded` | | 4013 | `job_error` | --- # Overview Learn about the different ways to use our APIs, including cloud services and on-prem containers. ## Cloud[​](#cloud "Direct link to Cloud") Leverage Speechmatics’ cloud services for easy, scalable, and fully managed speech-to-text and translation capabilities. The best way to get started using Speechmatics' cloud services is: * Create an account in our [Portal](https://portal.speechmatics.com/) * Check out our [Realtime transcription](/speech-to-text/realtime/quickstart.md) * Check out our [Batch transcription](/speech-to-text/batch/quickstart.md) ## On-prem[​](#on-prem "Direct link to On-prem") Deploy Speechmatics services in your own environment using containers. This option provides maximum control over your deployment and data. * [CPU speech-to-text container](/deployments/container/cpu-speech-to-text.md): Deploy the Speechmatics speech-to-text engine as a CPU based containerized service on your own hardware. * [GPU speech-to-text container](/deployments/container/gpu-speech-to-text.md): Deploy the Speechmatics speech-to-text engine as a GPU based containerized service on your own hardware. * [Kubernetes](/deployments/kubernetes/.md): Deploy the Speechmatics application as a Kubernetes service on your own hardware or your choosen cloud vendor. * [Language ID container](/deployments/container/language-id.md): Identify the language spoken in your audio using the Language ID container. * [Translation container](/deployments/container/gpu-translation.md): Translate audio from one language to another using the Translation container. ## Feature availability[​](#feature-availability "Direct link to Feature availability") Feature availability varies depending on the deployment method you choose. Below is a table summarizing the speech to text feature availability for each deployment method and processing mode. | Feature | Modes | Deployments | | ---------------------------------------------------------------------------------------------- | --------------- | ------------------------------------ | | [Multilingual speech to text](/speech-to-text/languages.md#bilingual-and-multi-language-packs) | Batch, Realtime | SaaS, On-prem | | [Alignment](/speech-to-text/batch/alignment.md) | Batch | SaaS | | [Audio events](/speech-to-text/features/audio-events.md) | Batch, Realtime | SaaS, On-prem | | [Audio filtering](/speech-to-text/features/audio-filtering.md) | Batch, Realtime | SaaS, On-prem | | [Auto chapters](/speech-to-text/batch/speech-intelligence/auto-chapters.md) | Batch | SaaS | | [Custom dictionary](/speech-to-text/features/custom-dictionary.md) | Batch, Realtime | SaaS, On-prem | | [Diarization](/speech-to-text/features/diarization.md) | Batch, Realtime | SaaS, On-prem | | [Disfluencies and word replacement](/speech-to-text/formatting.md#disfluencies) | Batch, Realtime | SaaS, On-prem | | [Feature discovery](/speech-to-text/features/feature-discovery.md) | Batch, Realtime | SaaS | | [Fetch URL](/speech-to-text/batch/input.md#fetch-url) | Batch | SaaS, On-Prem | | [Language identification](/speech-to-text/batch/language-identification.md) | Batch | SaaS | | [Notifications](/speech-to-text/batch/notifications.md) | Batch | SaaS, On-prem | | [Punctuation settings](/speech-to-text/formatting.md#punctuation) | Batch, Realtime | SaaS, On-prem | | [Sentiment analysis](/speech-to-text/batch/speech-intelligence/sentiment-analysis.md) | Batch | SaaS, On-prem | | [Smart formatting](/speech-to-text/formatting.md#smart-formatting) | Batch, Realtime | SaaS, On-prem | | [Speaker identification](/speech-to-text/features/speaker-identification.md) | Batch, Realtime | SaaS, On-prem[1](#user-content-fn-1) | | [Summarization](/speech-to-text/batch/speech-intelligence/summarization.md) | Batch | SaaS | | [Topic detection](/speech-to-text/batch/speech-intelligence/topic-detection.md) | Batch | SaaS | | [Tracking](/speech-to-text/batch/output.md#tracking-metadata) | Batch, Realtime | SaaS, On-prem | | [Translation](/speech-to-text/features/translation.md) | Batch, Realtime | SaaS, On-prem | | [Turn detection](/speech-to-text/realtime/turn-detection.md) | Realtime | SaaS, On-prem | ## Footnotes[​](#footnote-label "Direct link to Footnotes") 1. On an on-prem deployment, batch speaker identification requires the [GPU speech-to-text container](/deployments/container/gpu-speech-to-text.md). Realtime speaker identification is supported on both CPU and GPU containers. See [speaker identification secrets](/deployments/container/speaker-identification.md). [↩](#user-content-fnref-1) --- # Accessing images Learn how to access images in the Speechmatics Container system The Speechmatics Docker images are obtained from the Speechmatics Docker Repository. If you do not have a Speechmatics Docker Repository account or have lost your details, please reach out to [Support](https://support.speechmatics.com). The latest information about the Containers can be found in the knowledge base section of the [Support Portal](https://support.speechmatics.com). If a support account is not available or the *Containers* section is not visible in the Support Portal, please reach out to [Support](https://support.speechmatics.com) for help. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * Speechmatics Docker repository credentials – speak to [Support](https://support.speechmatics.com) to get these * Language Code – the ISO language code (for example `fr` for French) * `LICENSE_TOKEN` - The value of the signed claims token which is used to validate the license file. This is required to run the Container. Speechmatics Support will provide this within the license file generated for each customer * `TAG` – which is used to identify the image version ## Docker repository login[​](#docker-repository-login "Direct link to Docker repository login") Using the credentials supplied by Speechmatics, login to our Docker repository: ``` docker login https://speechmaticspublic.azurecr.io ``` You will be prompted for your username and password that was provided to you. If successful, you will see the response: ``` Login Succeeded ``` If unsuccessful, please verify your credentials and URL. If problems persist, please contact Speechmatics Support. Speechmatics require all customers to cache a copy of the Docker image(s) within their own environment. Once the first version of the container is pulled from the Speechmatics Software Repository, please re-host in a private container registry instead and reference your personal registry in your deployments. ## Pulling core speech CPU images[​](#pulling-core-speech-cpu-images "Direct link to Pulling core speech CPU images") Each supported language pack comes as a different Docker image, so the process will need to be repeated for each language pack required using the relevant language code. ``` # pulling Batch Global English (en) with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/batch-asr-transcriber-en:15.19.0 # pulling the Batch Spanish (es) model with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/batch-asr-transcriber-es:15.19.0 # pulling Realtime Global English (en) with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/rt-asr-transcriber-en:15.19.0 ``` See [how to run the Core Speech CPU container here.](/deployments/container/cpu-speech-to-text.md) ## Pulling transcription GPU images[​](#pulling-transcription-gpu-images "Direct link to Pulling transcription GPU images") The Transcription GPU images are required to use the most accurate models. See [how to run the Transcription GPU container here.](/deployments/container/gpu-speech-to-text.md) To access additional language configurations for containers, please [speak to our Support Team](https://support.speechmatics.com). ### Standard model[​](#standard-model "Direct link to Standard model") There is a single image available that supports all languages for the Standard model. There are language specific images available that support the Enhanced and Standard models. ``` # pulling the Standard model Transcription GPU inference server which supports all languages with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/sm-gpu-inference-server-standard-all:15.19.0 # pulling language specific Transcription GPU inference servers available for en, es, de, fr. Supports both Enhanced and Standard models with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/sm-gpu-inference-server-en:15.19.0 ``` ### Enhanced model[​](#enhanced-model "Direct link to Enhanced model") Depending on which Enhanced model languages are required, you can pull specific images. Language Pack 1 Bashkir, Basque, Belarusian, English, Esperanto, Irish, Malay & English bilingual, Mandarin & English bilingual, Mandarin Malay Tamil & English multilingual, Marathi, Mongolian, Tamil, Tamil & English bilingual, Turkish, Ukrainian, Uyghur, Welsh ``` docker pull speechmaticspublic.azurecr.io/sm-gpu-inference-server-enhanced-recipe1:15.19.0 ``` Language Pack 2 Bulgarian, Croatian, Estonian, Galician, Indonesian, Interlingua, Latvian, Lithuanian, Persian, Romanian, Slovakian, Slovenian, Spanish, Tagalog, Urdu ``` docker pull speechmaticspublic.azurecr.io/sm-gpu-inference-server-enhanced-recipe2:15.19.0 ``` Language Pack 3 Arabic & English bilingual, Catalan, Czech, Danish, Finnish, German, Greek, Hebrew, Hindi, Hungarian, Italian, Korean, Malay, Swahili, Swedish ``` docker pull speechmaticspublic.azurecr.io/sm-gpu-inference-server-enhanced-recipe3:15.19.0 ``` Language Pack 4 Arabic, Bengali, Cantonese, Dutch, French, Japanese, Maltese, Mandarin, Norwegian, Polish, Portuguese, Russian, Thai, Vietnamese ``` docker pull speechmaticspublic.azurecr.io/sm-gpu-inference-server-enhanced-recipe4:15.19.0 ``` ### Standard and Enhanced model[​](#standard-and-enhanced-model "Direct link to Standard and Enhanced model") A single image supports all languages for both the Standard and Enhanced models, so you do not need to pull individual language pack images. The `sm-gpu-inference-server-standard-all` image above covers all languages for the Standard model only. ``` # pulling the Transcription GPU inference server supporting all languages for both Enhanced and Standard models with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/sm-gpu-inference-server-all-lang:15.19.0 ``` This image contains both models. To load only one of them, see [Running only one model](/deployments/container/gpu-speech-to-text.md#running-only-one-model). To load only some of the image's languages, see [Loading only selected languages](/deployments/container/gpu-speech-to-text.md#loading-only-selected-languages). ### Melia 1 model[​](#melia-1-model "Direct link to Melia 1 model") Melia 1 is a multilingual, GPU-only model, currently available in Batch. For an overview, see [Models](/speech-to-text/models.md#melia-1). ``` # pulling the Melia 1 Batch transcriber with the 1.3.0 tag: docker pull speechmaticspublic.azurecr.io/sm-asr-transcriber-melia-1:1.3.0 # pulling the Melia 1 GPU inference server with the 1.3.0 tag: docker pull speechmaticspublic.azurecr.io/sm-asr-inference-server-melia-1:1.3.0 ``` See [how to run the Melia 1 GPU container here.](/deployments/container/gpu-speech-to-text-melia-1.md) ## Pulling translation GPU image[​](#pulling-translation-gpu-image "Direct link to Pulling translation GPU image") This GPU image is required to use Translation in Batch or Realtime. ``` # pulling the Translation GPU inference server which supports all translation pairs with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/sm-translation-inference-server:15.19.0 ``` See [how to run the Translation GPU container here.](/deployments/container/gpu-translation.md) ## Pulling bilingual images[​](#pulling-bilingual-images "Direct link to Pulling bilingual images") To use Spanish and English bilingual transcription you need to pull the Core Speech CPU image below for the client and use the Spanish [GPU Inference Server](#pulling-transcription-gpu-images). ``` # pulling Batch bilingual Spanish and English with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/batch-asr-transcriber-es-bilingual-en:15.19.0 # pulling Realtime bilingual Spanish and English with the 15.19.0 tag: docker pull speechmaticspublic.azurecr.io/rt-asr-transcriber-es-bilingual-en:15.19.0 ``` Bilingual is only available for GPU deployments. Batch and Realtime are supported. ## Pulling language ID Image[​](#pulling-language-id-image "Direct link to Pulling language ID Image") This image is required to use Language ID. ``` # pulling the latest Language ID image: docker pull speechmaticspublic.azurecr.io/langid:2.2.1 ``` See [how to run the Language ID container here.](/deployments/container/language-id.md) --- # Additional security features Learn about the Speechmatics container system security This section documents additional measures you can take to run the Batch Container where there are restrictive requirements on data storage or user access. ## Custom mapping temporary directories to run the batch container[​](#custom-mapping-temporary-directories-to-run-the-batch-container "Direct link to Custom mapping temporary directories to run the batch container") Users may wish to run the Batch Container in an environment where they cannot or do not want to write anything to disk, and instead use temporary storage like `tmpfs` or `ramfs` to ensure regulatory compliance. The Batch Container supports mounting temporary directories for the storage of all intermediate files created during transcription, as well as mounting the directories where input, output and job configuration files are placed. Files can also be locally retrieved by using the `fetch_url` functionality in the configuration object. Speechmatics also supports the `--job-config` option to specify the location of the configuration object. The job config location must specify the location in the container at which the config file can be found. If this also needs to be in a temporary directory (e.g. `tmp`), rather than `tmpfs` this must be a volume from a host machine in which the configuration object can be found. Below is an example, where the intermediate files and configuration object are in temporary storage. Please note that the `--job-config` argument must come after the image name ``` docker run --rm -i \ --read-only --tmpfs /home/smuser \ -v :/tmp \ -e LICENSE_TOKEN=$TOKEN_VALUE \ batch-asr-transcriber-en:15.19.0 \ --job-config /tmp/config.json ``` This example sets up a `tmpfs` for intermediate files created by transcription, so that all such files are written to transient storage, rather than to disk. The configuration object is mounted in a retrievable folder in `tmp`. An alternative is to use `tmp` as `tmpfs` and then mount an additional read-only volume on a path inside the Container in which the config can be found. ``` docker run --rm -i \ --read-only --tmpfs /home/smuser --tmpfs /tmp \ -v :/configs_dir:ro \ -e LICENSE_TOKEN=$TOKEN_VALUE \ batch-asr-transcriber-en:15.19.0 \ --job-config /example_configs_dir/config.json ``` If the Container is run using Kubernetes, users can use the [emptyDir](https://kubernetes.io/docs/concepts/storage/volumes/#emptydir) to mount `tmpfs` in the needed directories (/home/smuser and /tmp). Configuration files can also be stored in an emptyDir if any of the Containers in the pod is able to put it there. This could be achieved in deployment software like Kubernetes by using an [initContainer](https://kubernetes.io/docs/concepts/workloads/pods/init-containers/) or using the [sidecar pattern](https://www.oreilly.com/library/view/designing-distributed-systems/9781491983638/ch02.html) or by fetching the configuration from its original location and storing it in the `emptyDir` volume. Then the transcriber should be called with the `--job-config argument` pointing to the path in the emptyDir volume in which the config was stored. Users can also pull files from temporary locations using `fetch_url` functionality Below is a configuration example: ``` { "type": "transcription", "transcription_config": { "language": "en" }, "fetch_data": { "url": "file:///tmp/$FILENAME.wav" } } ``` ## Running a batch container as a non-root user[​](#running-a-batch-container-as-a-non-root-user "Direct link to Running a batch container as a non-root user") There are some use cases where you may not be able to run the Batch Container as a root user. This may be because you are working in a hosting environment that mandates the use of a named user rather than root. You must start the Container with the flag `-–user $USERNUMBER:$GROUPID`. User number and group ID are non-zero numerical values from a value of **1** up to a value of **65535**. Here is a working example: ``` docker run --user 1000:3000 ubuntu echo hello world ``` ### Getting transcription output as a non-root user[​](#getting-transcription-output-as-a-non-root-user "Direct link to Getting transcription output as a non-root user") If you take transcription via the default STDOUT, then this will not change as a non-root user. Below is an example: ``` docker run -u 1020:4000 \ -v /Users/$USER/work/pipeline/mydev/config.json:/config.json \ -v /Users/$USER/work/pipeline/mydev/input.audio:/input.audio \ ${IMAGE_NAME} ``` If you want to write the output to a specific directory, you must volume map a directory to which the non-root user would have access. ### Running a batch container as a non-root user on kubernetes[​](#running-a-batch-container-as-a-non-root-user-on-kubernetes "Direct link to Running a batch container as a non-root user on kubernetes") The examples below **do not** constitute an explicit recommendation to run as non-root user, merely a guideline on how to do so with Kubernetes only when this is an unavoidable requirement. If you require named users to be deployed on Kubernetes Pods, you must set the following Security Config. The user and group **must** correspond to the user and group you use when starting the container ``` securityContext: runAsUser: { non-zero numerical value between 0 and 65535 } runAsGroup: { non-zero numerical value between 0 and 65535 } ``` There is more information on how to configure security settings on Kubernetes pods [here](https://kubernetes.io/docs/tasks/configure-pod-container/security-context/) Some Kubernetes deployments may mandate the use of PodSecurity Admissions Controllers. These provide stricter security requirements. More information on them can be found [here](https://kubernetes.io/docs/reference/access-authn-authz/admission-controllers/). If your deployment requires this set up, here is an example configuration that allows you to carry out transcription as a non-root user. ``` apiVersion: policy/v1beta1 kind: PodSecurityPolicy metadata: name: restricted annotations: seccomp.security.alpha.kubernetes.io/allowedProfileNames: 'docker/default,runtime/default' apparmor.security.beta.kubernetes.io/allowedProfileNames: 'runtime/default' seccomp.security.alpha.kubernetes.io/defaultProfileName: 'runtime/default' apparmor.security.beta.kubernetes.io/defaultProfileName: 'runtime/default' spec: privileged: false # Required to prevent escalations to root. allowPrivilegeEscalation: false requiredDropCapabilities: - ALL # Allow core volume types. volumes: - 'configMap' - 'emptyDir' - 'projected' - 'secret' - 'downwardAPI' # Assume that persistentVolumes set up by the cluster admin are safe to use. - 'persistentVolumeClaim' hostNetwork: false hostIPC: false hostPID: false runAsUser: # Require the container to run without root privileges. rule: 'MustRunAsNonRoot' seLinux: # This policy assumes the nodes are using AppArmor rather than SELinux. rule: 'RunAsAny' supplementalGroups: rule: 'MustRunAs' ranges: # Forbid adding the root group. - min: 1 max: 65535 fsGroup: rule: 'MustRunAs' ranges: # Forbid adding the root group. - min: 1 max: 65535 readOnlyRootFilesystem: false ``` --- # Batch persistent worker Run a long-lived HTTP transcription worker that accepts multiple jobs without restarting, reducing turnaround time and improving CPU/GPU utilisation. Available from version 15.7.0 A batch persistent worker (also called as **HTTP batch worker**) is a long-running transcription service that loads the ASR models once at startup and then accepts jobs over an HTTP API for the lifetime of the container. Unlike standard batch containers — which start up, process a single job, and exit — a persistent worker stays alive indefinitely, serving jobs as they arrive. This gives you: * **No per-job cold start.** The models are loaded into memory once. Every subsequent job skips the startup cost entirely. * **Concurrent processing.** The `--parallel` flag controls how many processing units the worker handles simultaneously. Individual jobs can also be assigned multiple processing units (referred as `engines` in this document) to reduce their own turnaround time. The worker exposes an HTTP API for submitting jobs, polling status, fetching transcripts, and checking availability. ## Why use a persistent worker?[​](#why-use-a-persistent-worker "Direct link to Why use a persistent worker?") | | Standard batch | Persistent worker | | ------------------- | ------------------------ | --------------------------------- | | Startup cost | Per job | Once | | Memory usage | One container per job | Multiple jobs share one container | | CPU/GPU utilisation | Interrupted between jobs | Continuous | | Best for | Large, infrequent files | High throughput or smaller files | **Cold start overhead is significant for short audio.** Loading the ASR models especially onto GPU takes several seconds. For a five minute file this cost is negligible. For a ten second clip, startup can take longer than transcription itself. The persistent worker eliminates this by loading the models once. **High-throughput workloads benefit from a single long-lived container.** Routing many jobs to one worker is more efficient than launching a container per job. The `--parallel` setting lets you tune concurrency to your workload. **GPU utilisation is maximised.** On GPU deployments, a standard batch container leaves the GPU idle between jobs. A persistent worker keeps the GPU warm and available, reducing wasted capacity across back-to-back requests. When processing long audio jobs the benefits on RTF of the persistent batch worker is negligible, and the resultant RTF is similar to that of a standard batch job. ## Deploying the worker[​](#deploying-the-worker "Direct link to Deploying the worker") ### Docker[​](#docker "Direct link to Docker") ``` docker run -it -e LICENSE_TOKEN=$TOKEN_VALUE -e SM_INFERENCE_ENDPOINT=: -p PORT:18000 batch-asr-transcriber-en:15.19.0 --run-mode http --parallel=4 --all-formats /output_dir_name ``` #### Parameters[​](#parameters "Direct link to Parameters") | Parameter | Description | | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `--parallel` | Number of parallel engines (each engine maps to one GPU connection when on GPU container). | | `--all-formats` | Directory where all job outputs and logs are saved. If omitted, defaults to `/tmp/jobs`. See [generating multiple transcript formats](https://docs.speechmatics.com/deployments/container/cpu-speech-to-text#generating-multiple-transcript-formats) for details. | | `PORT` | The local port forwarded to the container's internal port (`18000`). | #### Environment variables[​](#environment-variables "Direct link to Environment variables") | Variable | Description | | --------------------------------- | ------------------------------------------------------------ | | `SM_BATCH_WORKER_LISTEN_PORT` | Override the default internal port (`18000`). | | `SM_BATCH_WORKER_MAX_JOB_HISTORY` | Maximum number of completed job records to retain in memory. | ## Submitting a job[​](#submitting-a-job "Direct link to Submitting a job") Once the worker is running and is available, submit jobs by making a `POST` request to `/v2/jobs` with an audio file and transcription config. The worker queues the job and returns a `job_id` immediately. You can poll [`GET /v2/jobs/{job_id}`](#job-api-reference) for the job status, and fetch the transcript when the status changes to `DONE`. curlcurlPython SDKPython SDK ``` curl -X POST address.of.container:PORT/v2/jobs \ -H 'X-SM-Processing-Data: {"parallel_engines": 2, "user_id": "MY_USER_ID"}' \ -F 'config={ "type": "transcription", "transcription_config": { "language": "en", "diarization": "speaker", "model": "enhanced" } }' \ -F 'data_file=@~/audio_file.mp3' ``` ``` import asyncio import os from dotenv import load_dotenv from speechmatics.batch import AsyncClient load_dotenv() async def main(): client = AsyncClient( api_key=os.getenv("SPEECHMATICS_API_KEY"), url="address.of.container:PORT/v2" ) result = await client.transcribe( "audio.wav", parallel_engines=2, user_id="MY_USER_ID" ) print(result.transcript_text) await client.close() asyncio.run(main()) ``` ### Fetching audio from a URL[​](#fetching-audio-from-a-url "Direct link to Fetching audio from a URL") As well as uploading audio files directly, you can have the worker fetch them from a remote URL by adding a `fetch_data` section to the job config. You can find an example of this below. The full config options are documented [here](/speech-to-text/batch/input.md#fetch-url). ``` curl -X POST address.of.container:PORT/v2/jobs \ -F 'config={ "type": "transcription", "transcription_config": { "language": "en" }, "fetch_data": { "url": "https://example.com/audio_file.mp3" } }' ``` ## Managing capacity[​](#managing-capacity "Direct link to Managing capacity") The worker processes multiple jobs concurrently, up to the `--parallel` limit you set at startup. To check available capacity before submitting, query the `/jobs` endpoint. ### `GET /jobs`[​](#get-jobs "Direct link to get-jobs") Returns current engine usage and a list of active jobs. The `unused_engines` field tells you how many engines are free, and you can use it to determine how many engines you can request for the next job. **Example response:** ``` { "active_jobs": [ { "job_id": "f8a564954b334eecb823", "parallel_engines": 1 }, { "job_id": "29351ae8cf2c4e8694f0", "parallel_engines": 1 } ], "max_engines": 8, "unused_engines": 6 } ``` ### Requesting parallel engines[​](#requesting-parallel-engines "Direct link to Requesting parallel engines") Each job can request multiple engines using the `parallel_engines` value in the `X-SM-Processing-Data` header. More engines per job means faster turnaround for that job, at the cost of reduced concurrency for others. ``` curl -X POST address.of.container:PORT/v2/jobs \ -H 'X-SM-Processing-Data: {"parallel_engines": 2}' \ -F 'config={"type": "transcription", "transcription_config": {"language": "en"}}' \ -F 'data_file=@~/audio_file.mp3' ``` If a job requests more engines than are currently available, it will be rejected: ``` HTTP 503: {"detail": "Server busy: 8 engines not available (2 engines in use, 5 parallel allowed)"} ``` ## Speaker identification[​](#speaker-identification "Direct link to Speaker identification") To enable the speaker identification feature you can use the same logic used for the one shot [batch container](/speech-to-text/features/speaker-identification.md). Speaker identification requires the [GPU speech-to-text container](/deployments/container/gpu-speech-to-text.md) — it is not supported in batch mode on CPU. To enable per-customer encrypted identifiers (as used in our SaaS offering), pass a `user_id` in the `X-SM-Processing-Data` header. ``` curl -X POST address.of.container:PORT/v2/jobs \ -H 'X-SM-Processing-Data: {"user_id": "MY_USER_ID"}' \ -F 'config={ "type": "transcription", "transcription_config": { "language": "en", "diarization": "speaker", "model": "enhanced" } }' \ -F 'data_file=@~/audio_file.mp3' ``` For details on secrets management, refer to the [Speaker identification documentation](/deployments/container/speaker-identification.md). ## Job API reference[​](#job-api-reference "Direct link to Job API reference") The HTTP batch worker API is similar to our [V2 SaaS API](/api-ref/batch/create-a-new-job.md). This makes it easy to use our SaaS and on-prem offerings interchangeably. The only differences between the SaaS API and our HTTP workers are: We don't support below calls: * The `include_deleted` parameter in the `GET /v2/jobs` * `GET /v2/usage` * `DELETE /v2/jobs:jobid` For the API call `GET /v2/jobs/{job_id}`, we also return the `request_id` as part of the response. ## Health endpoints[​](#health-endpoints "Direct link to Health endpoints") The worker exposes two health endpoints on the same port as job submission. These endpoints are designed to work as [liveness and readiness probes](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/) in a Kubernetes cluster. ### `GET /live`[​](#get-live "Direct link to get-live") Liveness probe. Returns `200` when all container services are running and healthy. ``` { "live": true } ``` ### `GET /ready`[​](#get-ready "Direct link to get-ready") Readiness probe. Returns `200` when at least one engine slot is free, `503` when all engines are occupied. The response includes the current number of engines in use. ``` { "ready": true, "engines_used": 0 } ``` --- # CPU Speech to text container Learn about the Speechmatics CPU container system The Core Speech CPU container is a single container that provides transcription. It should be used for deployments where GPUs are not available. To use our latest and most accurate models, please refer to the [Transcription GPU Container](/deployments/container/gpu-speech-to-text.md) deployment. ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [A license file or a license token](/deployments/container/licensing.md) * [Access to our Docker repository](/deployments/container/accessing-images.md) ## System requirements[​](#system-requirements "Direct link to System requirements") CPU containers are split by language. Each running container will require the following resources: * 1 vCPU * 2-5GB RAM * 100MB hard disk space * 3GB storage * If you are using the Enhanced model, it is recommended to use the upper limit of the RAM recommendations * The host machine should have an AMD or Intel CPU with modern AVX instructions as this generally improves transcription processing speed. Exact impact on transcription speed varies on brand, generation, which instruction sets are available and resource allocation * If you are using a hypervisor, you should pass through to the VM all AVX related instruction sets When using the [parallel processing](#parallel-processing-guide) functionality of the Batch container, this will require more resource due to the intensive memory required. When using parallel processing, we recommend using (`N * RAM` requirements) where `N` is the number of vCPUs intended to be used for parallel processing. So if 2 vCPUs were required for parallel processing, the RAM requirements would be up to 10GB See [Performance and cost](/deployments/container/performance-and-cost.md) for more information on the performance and cost of the container. ## Batch transcription[​](#batch-transcription "Direct link to Batch transcription") Each Batch container processes one input file and outputs a resulting transcript in a predefined language in a number of supported outputs. All data is transitory. Once a container completes its transcription it removes all record of the operation. Input file sizes up to 2 hours in length or 4GB in size. ### Input methods[​](#input-methods "Direct link to Input methods") There are two different methods for passing an audio file into a container. Stream the audio through the container via standard input (STDIN): ``` cat ~/example.wav | docker run -i \ -e LICENSE_TOKEN=$TOKEN_VALUE \ batch-asr-transcriber-en:15.19.0 ``` Pull an audio file from a mapped directory into the `input.audio` file within the container: ``` docker run \ -v ~/example.wav:/input.audio \ -e LICENSE_TOKEN=$TOKEN_VALUE \ batch-asr-transcriber-en:15.19.0 ``` See [Docker docs](https://docs.docker.com/engine/reference/commandline/run/) for a full list of the available options. Both the methods will produce the same transcribed outcome and will write a JSON response to standard output (stdout) and any other logs to standard error (stderr). The intermediate files created during the transcription are stored in `/home/smuser/work`. This is the case whether running the container as a root or non-root user. Here is an example output: ``` { "format": "2.9", "metadata": { "created_at": "2023-08-02T15:43:50.871Z", "type": "transcription", "language_pack_info": { "adapted": false, "itn": true, "language_description": "English", "word_delimiter": " ", "writing_direction": "left-to-right" }, "transcription_config": { "language": "en", "diarization": "none" } }, "results": [ { "alternatives": [ { "confidence": 1.0, "content": "Are", "language": "en", "speaker": "UU" } ], "end_time": 3.61, "start_time": 3.49, "type": "word" }, { "alternatives": [ { "confidence": 1.0, "content": "on", "language": "en", "speaker": "UU" } ], "end_time": 3.73, "start_time": 3.61, "type": "word" } ] } ``` The exit code of the Container will also determine if the transcription was successful. There are two exit code possibilities: * `Exit Code == 0` : The transcript was a success; the output will contain a JSON output defining the transcript (more info below) * `Exit Code != 0` : the output will contain a stack trace and other useful information. This output should be used in any communication with Speechmatics Support to aid understanding and resolution of any problems that may occur If you encounter any issues refer to the [troubleshooting](/deployments/container/troubleshooting.md) documentation, which includes more detailed exit codes. Now that you have successfully run a job, you can use the above APIs in your workflow. In the following section we will show ways you can modify the container to create simple ways of orchestrating the it within your deployments. ### Modifying the image[​](#modifying-the-image "Direct link to Modifying the image") #### Building an image[​](#building-an-image "Direct link to Building an image") Using STDIN to pass files in and obtain the transcription may not be sufficient for all use cases. It is possible to build a new Docker Image that will use the Speechmatics Image as a layer if required for your specific workflow. To include the Speechmatics Docker Image inside another image, ensure to add the pulled Docker Image into the Dockerfile for the new application. #### Requirements for a custom image[​](#requirements-for-a-custom-image "Direct link to Requirements for a custom image") To ensure the Speechmatics Docker Image works as expected inside the custom image, please consider the following: * Any audio that needs to be transcribed must to be copied to a file called `/input.audio` inside the running Container * To initiate transcription, call the application `pipeline`. The `pipeline` will start the transcription service and use `/input.audio` as the audio source * When running `pipeline`, the working directory must be set to `/opt/orchestrator`, using either the Dockerfile `WORKDIR` directive, the `cd` command or similar means * Once `pipeline` finishes transcribing, ensure you move the transcription data outside the Container #### Dockerfile[​](#dockerfile "Direct link to Dockerfile") To add a Speechmatics Docker Image into a custom one, the Dockerfile must be modified to include the full image name of the locally available image. Example: Adding English (en) with tag 15.19.0 to the Dockerfile Dockerfile ``` FROM batch-asr-transcriber-en:15.19.0 ADD download_audio.sh /usr/local/bin/download_audio.sh RUN chmod +x /usr/local/bin/download_audio.sh CMD ["/usr/local/bin/download_audio.sh"] ``` Once the above image is built, and a Container instantiated from it, a script called `download_audio.sh` will be executed (this could do something like pulling a file from a webserver and copying it to `/input.audio` before starting the pipeline application). This is a very basic Dockerfile to demonstrate a way of orchestrating the Speechmatics Docker Image. For support purposes, it is assumed the Docker Image provided by Speechmatics has been unmodified. If you experience issues, Speechmatics support will require you to replicate the issues with the unmodified Docker image e.g. `batch-asr-transcriber-en:15.19.0` ### Parallel processing guide[​](#parallel-processing-guide "Direct link to Parallel processing guide") For customers who are looking to improve job turnaround time and who are able to assign sufficient resources, it is possible to pass a parallel transcription parameter to the container to take advantage of multiple CPUs. The parameter is called parallel and the following example shows how it can be used. In this case to use 4 cores to process the audio you would run the Container like this: ``` docker run -i \ -v ~/example.wav:/input.audio \ batch-asr-transcriber-en:15.19.0 \ --parallel=4 ``` Depending on your hardware, you may need to experiment to find the optimum performance. We've noticed significant improvement in turnaround time for jobs by using this approach. If you limit or are limited on the number of CPUs you can use (for example your platform places restrictions on the number of cores you can use, or you use the --cpu flag in your docker run command), then you should ensure that you do not set the parallel value to be more than the number of available cores. If you attempt to use a setting in excess of your free resources, then the Container will only use the available cores. If you simply increase the parallel setting to a large number you will see diminishing returns. Moreover, because files are split into 5 minute chunks for parallel processing, if your files are shorter than 5 minutes then you will see no parallelization (in general the longer your audio files the more speedup you will see by using parallel processing). If you are running the container on a shared resource you may experience different results depending on what other processes are running at the same time. The optimum number of cores is `N / 5`, where `N` is the length of the audio in minutes. Values higher than this will deliver little to no value, as there will be more cores than chunks of work. A typical approach will be to increment the parallel setting to a point where performance plateaus, and leave it at that (all else being equal). For large files and large numbers of cores, the time taken by the first and last stages of processing (which cannot be parallelized) will start to dominate, with diminishing returns. ### Generating multiple transcript formats[​](#generating-multiple-transcript-formats "Direct link to Generating multiple transcript formats") In addition to our primary JSON format, the Speechmatics container can output transcripts in the plain text (TXT) and SubRip (SRT) subtitle format. This can be done by using
`--all-formats` command and then specifying a directory parameter within the transcription request. This is where all supported transcript formats will be saved. You can also use
`--allformats` to generate the same response. This directory must be mounted into the container so the transcripts can be retrieved after container finishes. You will receive a transcript in all currently supported formats: JSON, TXT, and SRT. The following example shows how to use `--all-formats` parameter. In this scenario, after processing the file, three separate transcripts would be found in the `~/tmp/output` directory. These transcripts would be in JSON, TXT, and SRT format. ``` docker run \ -v ~/example.wav:/input.audio \ -v ~/tmp/output:/output_dir_name \ -e LICENSE_TOKEN=$TOKEN_VALUE \ batch-asr-transcriber-en:15.19.0 \ --all-formats /output_dir_name ``` ## Realtime transcription[​](#realtime-transcription "Direct link to Realtime transcription") The Realtime container provides the ability to transcribe speech data in a predefined language from a live stream or a recorded audio file. * Multiple instances of the container can be run on the same Docker host. This enables scaling of a single language or multiple languages as required * All data is transitory, once a container completes its transcription it removes all record of the operation, no data is persisted Here's an example of how to start the Container from the command line: ``` docker run \ -p 9000:9000 \ -p 8001:8001 \ -e LICENSE_TOKEN=$TOKEN_VALUE \ rt-asr-transcriber-en:15.19.0 ``` See [Docker docs](https://docs.docker.com/engine/reference/commandline/run/) for a full list of the available options. ### Multi-session containers[​](#multi-session-containers "Direct link to Multi-session containers") By default the real-time container will accept only one websocket connection at a time. To enable multiple connections, set the environment variable `SM_MAX_CONCURRENT_CONNECTIONS` to the maximum number of sessions to allow. When this is set, the `/ready` [health check endpoint](#health-service) will return true if there is a free connection available. CPU usage scales with the number of active sessions, whereas most memory usage is shared between connections. When the transcription container is linked to a GPU inference server, the amount of memory which is shared is further increased. ### Reducing initial connection time[​](#reducing-initial-connection-time "Direct link to Reducing initial connection time") The first-session loading time can be reduced down to several hundred milliseconds by prewarming the transcriber. You can enable this feature by setting the `SM_PREWARM_ENGINE_MODES` environment variable, with a semicolon separated list describing the required engine modes. For example, to prewarm 1 English GPU Standard and 2 English GPU Enhanced: `SM_PREWARM_ENGINE_MODES='en_general_gpu_standard:1;en_general_gpu_enhanced:2'` In general, the format is: `{language}_{domain}_{processor}_{model}:{prewarm_connections}`. The parameters are: * `language` - One of the supported [language codes](/speech-to-text/languages.md) * `domain` - One of `general` or a domain used for some [multi-lingual transcription](/speech-to-text/languages.md#bilingual-and-multi-language-packs) use cases. For example: `SM_PREWARM_ENGINE_MODES='es_bilingual-en_gpu_standard:1'` * `processor` - One of `cpu` or `gpu`. Note that selecting `gpu` requires a [GPU Inference Container](/deployments/container/gpu-speech-to-text.md) * `model` - One of `standard` or `enhanced`. The [model](/speech-to-text/models.md) you want to prewarm * `prewarm_connections` - Integer. The number of engine instances of the specific mode you want to pre-warm. The total number of `prewarm_connections` cannot be greater than `SM_MAX_CONCURRENT_CONNECTIONS`. After the pre-warming is complete, this parameter does not limit the types of connections the engine can start. ### Input modes[​](#input-modes "Direct link to Input modes") The supported method for passing audio to a Realtime Container is to use a WebSocket. A session is setup with configuration parameters passed in using a `StartRecognition` message, and thereafter audio is sent to the container in binary chunks, with transcripts being returned in an `AddTranscript` message. In the `AddTranscript` message individual result segments are returned, corresponding to audio segments defined by pauses (and other latency measurements). #### Output[​](#output "Direct link to Output") The results list are sorted by increasing `start_time`, with a supplementary rule to sort by decreasing `end_time`. See below for an example: ``` { "message": "AddTranscript", "format": "2.9", "metadata": { "transcript": "full tell radar", "start_time": 0.11, "end_time": 1.07 }, "results": [ { "type": "word", "start_time": 0.11, "end_time": 0.4, "alternatives": [{ "content": "full", "confidence": 0.7 }] }, { "type": "word", "start_time": 0.41, "end_time": 0.62, "alternatives": [{ "content": "tell", "confidence": 0.6 }] }, { "type": "word", "start_time": 0.65, "end_time": 1.07, "alternatives": [{ "content": "radar", "confidence": 1.0 }] } ] } ``` #### Transcription duration information[​](#transcription-duration-information "Direct link to Transcription duration information") The Container will output a log message after every transcription session to indicate the duration of speech transcribed during that session. This duration only includes speech, and not any silence or background noise which was present in the audio. This data can be used to report usage back to us, or simply for your own records. The format of the log messages produced should match the following example: ``` 2020-04-13 22:48:05.312 INFO sentryserver Transcribed 52 seconds of speech ``` Consider using the following regular expression to extract just the seconds part from the line if you are parsing it: ``` ^.+ .+ INFO sentryserver Transcribed (\d+) seconds of speech$ ``` ### Read-only mode[​](#read-only-mode "Direct link to Read-only mode") Users may wish to run the Container in read-only mode. This may be necessary due to their regulatory environment, or a requirement not to write any media file to disk. An example of how to do this is below. ``` docker run -it --read-only \ -p 9000:9000 \ --tmpfs /tmp \ -e LICENSE_TOKEN=$TOKEN_VALUE \ rt-asr-transcriber-en:15.19.0 ``` The Container still requires a temporary directory with write permissions. Users can provide a directory (e.g `/tmp`) by using the `--tmpfs` Docker argument. A tmpfs mount is temporary, and only persisted in the host memory. When the Container stops, the tmpfs mount is removed, and files written there won’t be persisted. If customers want to use the shared Custom Dictionary Cache feature, they must also specify the location of cache and mount it as a volume ``` docker run -it --read-only \ -p 9000:9000 \ --tmpfs /tmp \ -v /cachelocation:/cache \ -e LICENSE_TOKEN=$TOKEN_VALUE \ -e SM_CUSTOM_DICTIONARY_CACHE_TYPE=shared \ rt-asr-transcriber-en:15.19.0 ``` ### Running container as a non-root user[​](#running-container-as-a-non-root-user "Direct link to Running container as a non-root user") A Realtime Container can be run as a non-root user with no impact to feature functionality. This may be required if a hosting environment or a company's internal regulations specify that a Container must be run as a named user. Users may specify the non-root command by the `docker run –-user $USERNUMBER:$GROUPID`. User number and group ID are non-zero numerical values from a value of **1** up to a value of **65535** An example is below: ``` docker run -it --user 100:100 \ -p 9000:9000 \ -e LICENSE_TOKEN=$TOKEN_VALUE \ rt-asr-transcriber-en:15.19.0 ``` ## How to use a shared custom dictionary cache[​](#how-to-use-a-shared-custom-dictionary-cache "Direct link to How to use a shared custom dictionary cache") The Speechmatics Realtime Container includes an optional [Custom Dictionary](/speech-to-text/features/custom-dictionary.md) cache mechanism to reduce session initialization times. You will see improvements when reusing an identical Custom Dictionary from the second time onwards. The cache volume is safe to use from multiple Containers concurrently if the operating system and its filesystem support file locking operations. The cache can store multiple Custom Dictionaries in any language used for transcription. It can support multiple Custom Dictionaries in the same language. If a Custom Dictionary is small enough to be stored within the cache volume, this will take place automatically if the shared cache is specified. For more information about how the shared cache storage management works, please see [Maintaining the Shared Cache](#maintaining-the-shared-cache). We highly recommend you ensure any location you use for the shared cache has enough space for the number of Custom Dictionaries you plan to allocate there. How to allocate Custom Dictionaries to the shared cache is documented below. ### How to set up the shared cache[​](#how-to-set-up-the-shared-cache "Direct link to How to set up the shared cache") The shared cache is enabled by setting the following value when running transcription: * Cache Location: You must volume map the directory location you plan to use as the shared cache to `/cache` when submitting a job * `SM_CUSTOM_DICTIONARY_CACHE_TYPE`: (mandatory if using the shared cache) This environment variable must be set to `shared` * `SM_CUSTOM_DICTIONARY_CACHE_ENTRY_MAX_SIZE`: (optional if using the shared cache). This determines the maximum size of any single Custom Dictionary that can be stored within the shared cache in **bytes** * E.G. a `SM_CUSTOM_DICTIONARY_CACHE_ENTRY_MAX_SIZE` with a value of 10000000 would set a max storage size of any Custom Dictionary at **10MB** * For reference a Custom Dictionary wordlist with 1000 words produces a cache entry of size around 200 kB, or **200000** bytes * A value of `-1` will allow **every** Custom Dictionary to be stored within the shared cache. This is the **default** assumed value * A Custom Dictionary Cache entry **larger** than the `SM_CUSTOM_DICTIONARY_CACHE_ENTRY_MAX_SIZE` will still be used in transcription, but will not be cached ### Maintaining the shared cache[​](#maintaining-the-shared-cache "Direct link to Maintaining the shared cache") If you specify the shared cache to be used and your Custom Dictionary is within the permitted size, Speechmatics Realtime Container will always try to cache the Custom Dictionary. If a Custom Dictionary cannot occupy the shared cache due to other cached Custom Dictionaries within the allocated cache, then older Custom Dictionaries will be removed from the cache to free up as much space as necessary for the new Custom Dictionary. This is carried out in order of the least recent Custom Dictionary to be used. Therefore, you must ensure your cache allocation large enough to handle the number of Custom Dictionaries you plan to store. We recommend a relatively large cache to avoid this situation if you are processing multiple Custom Dictionaries using the batch container (e.g 50 MB). If you don't allocate sufficient storage this could mean one or multiple Custom Dictionaries are deleted when you are trying to store a new Custom Dictionary. It is recommended to use a Docker volume with a dedicated filesystem with a limited size. If a user decides to use a volume that shares filesystem with the host, it is the user's responsibility to purge the cache if necessary. ### Creating the shared cache[​](#creating-the-shared-cache "Direct link to Creating the shared cache") In the example below, transcription is run where an example local docker volume is created for the shared cache. It will allow a Custom Dictionary of up to 5MB to be cached. BatchBatchRealtimeRealtime ``` docker volume create speechmatics-cache docker run -i -v /home/user/sm_audio.wav:/input.audio \ -e SM_CUSTOM_DICTIONARY_CACHE_TYPE=shared \ -e SM_CUSTOM_DICTIONARY_CACHE_ENTRY_MAX_SIZE=5000000 \ -v speechmatics-cache:/cache \ -e LICENSE_TOKEN=$TOKEN_VALUE \ batch-asr-transcriber-en:15.19.0 ``` ``` docker volume create speechmatics-cache docker run --rm -d \ -p 9000:9000 \ -e SM_CUSTOM_DICTIONARY_CACHE_TYPE=shared \ -e SM_CUSTOM_DICTIONARY_CACHE_ENTRY_MAX_SIZE=5000000 \ -v speechmatics-cache:/cache \ -e LICENSE_TOKEN=$TOKEN_VALUE \ rt-asr-transcriber-en:15.19.0 speechmatics transcribe --additional-vocab gnocchi --url ws://localhost:9000/v2 --ssl-mode=none test.mp3 ``` #### Viewing the shared cache[​](#viewing-the-shared-cache "Direct link to Viewing the shared cache") If all set correctly and the cache was used for the first time, a single entry in the cache should be present. The following example shows how to check what Custom Dictionaries are stored within the cache. This will show the **language**, the **sampling rate**, and the **checksum** value of the cached dictionary entries. ``` ls $(docker inspect -f "{{.Mountpoint}}" speechmatics-cache)/custom_dictionary en,16kHz,bef53e5bcca838a39c3707f1134bda6a09ff87aaa09203617528774734455edd ``` **Reducing the shared cache size** Cache size can be reduced by removing some or all cache entries. ``` rm -rf $(docker inspect -f "{{.Mountpoint}}" speechmatics-cache)/custom_dictionary/* ``` Before manually purging the cache, ensure that no containers have the volume mounted, otherwise an error during transcription might occur. Consider creating a new docker volume as a temporary cache while performing purging maintenance on the cache. ## Linking to a GPU inference container[​](#linking-to-a-gpu-inference-container "Direct link to Linking to a GPU inference container") The GPU Inference Container allows multiple speech recognition containers to offload heavy inference tasks to a GPU, where they can be batched and parallelized more efficiently. The CPU is run as normal, but with the additional environment variable `SM_INFERENCE_ENDPOINT` which indicates the GRPC endpoint of the Inference Server. Speech containers running in GPU mode use less local CPU and memory, so they can be packed more densely on a server. ``` docker run \ --rm \ -it \ -e SM_INFERENCE_ENDPOINT=: \ -v $PWD/license.json:/license.json \ -v $PWD/example.wav:/input.audio \ ``` #### When the inference server is not available[​](#when-the-inference-server-is-not-available "Direct link to When the inference server is not available") At start up, the Container will make a TCP connection to the `SM_INFERENCE_ENDPOINT` server to establish if it's accessible. If this test fails, the transcription will terminate with an error. #### Batch[​](#batch "Direct link to Batch") In the event of a connection error during transcription, the transcriber will retry for up to 60 seconds using an exponential back off. The length of this retry period can be configured with the `SM_SPLIT_RETRY_TIMEOUT` environment variable, which is a whole number of seconds. #### Realtime[​](#realtime "Direct link to Realtime") In Realtime mode, the transcriber will retry connection to the server for a maximum of 250ms before giving up. For more details see [GPU Inference Container](/deployments/container/gpu-speech-to-text.md) ## Health service[​](#health-service "Direct link to Health service") The container is able to expose an HTTP Health Service, which offers startup, liveness, readiness, and session listing probes. This is accessible from port 8001, and has four endpoints, `started`, `live`, `ready` and `session_status`. This may be especially helpful if you are deploying the container into a Kubernetes cluster. If you are using Kubernetes, we recommend that you also refer to the Kubernetes documentation around [liveness and readiness probes](https://kubernetes.io/docs/tasks/configure-pod-container/configure-liveness-readiness-startup-probes/). The Health Service is enabled by default and runs as a subprocess of the main entrypoint to the container. ### Endpoints[​](#endpoints "Direct link to Endpoints") The Health Service offers four endpoints: #### `/started`[​](#started "Direct link to started") This endpoint provides a startup probe. It can be queried using an HTTP GET request. You must include the relevant port, 8001, in the request. This probe indicates whether all services in the Container have successfully started. Once it returns a successful response code, it should never return an unsuccessful response code later. Possible responses: * `200` if all of the services in the container have successfully started. * `503` otherwise. A JSON object is also returned in the body of the response, indicating the status. Example: ``` $ curl -i address.of.container:8001/started HTTP/1.0 200 OK Server: BaseHTTP/0.6 Python/3.8.5 Date: Mon, 08 Feb 2021 12:46:21 GMT Content-Type: application/json { "started": true } ``` #### `/live`[​](#live "Direct link to live") This endpoint provides a liveness probe. It can be queried using an HTTP GET request. You must include the relevant port, 8001, in the request. This probe indicates whether all services in the Container are active. The services in the Container send regular updates to the Health Service, if they don't send an update for more than 10 seconds then they will be marked as 'dead' and this endpoint will return an unsuccessful response code. For example, if the WebSocket server in the Container were to crash, this endpoint should indicate that. Possible responses: * `200` if all of the services in the Container have successfully started, and have recently sent an update to the Health Service. * `503` otherwise. A JSON object is also returned in the body of the response, indicating the status. Example: ``` $ curl -i address.of.container:8001/live HTTP/1.0 200 OK Server: BaseHTTP/0.6 Python/3.8.5 Date: Mon, 08 Feb 2021 12:46:45 GMT Content-Type: application/json { "alive": true } ``` #### `/ready`[​](#ready "Direct link to ready") This endpoint provides a readiness probe. It can be queried using an HTTP GET request. The container has been designed to process multiple audio streams at a time. This probe indicates whether the container has a slot free for connections, and can be used as a scaling mechanism. **Note**: The readiness check is accurate within a 2 second resolution. If you do use this probe for load balancing, be aware that bursts of traffic within that 2 second window could all be allocated to a single Container since its readiness state will not change. Possible responses: * `200` if the container has a free connection slot. * `503` otherwise. In the body of the response there is also a JSON object with the current status. Example: ``` $ curl -i address.of.container:8001/ready HTTP/1.0 200 OK Server: BaseHTTP/0.6 Python/3.8.5 Date: Mon, 08 Feb 2021 12:47:05 GMT Content-Type: application/json { "ready": true } ``` #### `/session_status` (from 13.0.0 onwards)[​](#session_status-from-1300-onwards "Direct link to session_status-from-1300-onwards") This endpoint provides a list of the sessions being served by the container. It can be queried using an HTTP GET request. Possible responses: * `200` Successful listing of the current sessions * `503` otherwise. In the body of the response there is a JSON object listing the current sessions, which can be tied to log entries and individual client connections. The `session_id` is created by the transcriber and returned to the client on first connection, and will always be present. This endpoint only returns data for the `request_id` field if the header `x-request-id` was set in the initial websocket handshake. This is designed to support deployments where the transcriber sits behind a proxy or load balancer and the proxy adds an id to the connection when it's first created. Example: ``` $ curl -i address.of.container:8001/session_status HTTP/1.0 200 OK Server: BaseHTTP/0.6 Python/3.10.13 Date: Wed, 05 Mar 2025 11:41:51 GMT Content-Type: application/json { "max_sessions": 4, "active_sessions": [ {"session_id": "499e3f17-b9e4-4c72-b9aa-66cbbdafea53", "request_id": "499e3f17-b9e4-4c72-b9aa-66cbbdafea53"}, {"session_id": "fffc5412-d4b5-4e72-a912-663107349968", "request_id": "802da03c-77f5-4b9a-b0df-441734a3b2b0"}, ] } ``` --- # GPU Speech to text container (Standard and Enhanced) Learn about the Speechmatics Transcription GPU container system ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [A license file or a license token](/deployments/container/licensing.md) * There is no specific license for the GPU Inference Container, it will run using an existing Speechmatics license for the Realtime or Batch Container * [Access to our Docker repository](/deployments/container/accessing-images.md) ## System requirements[​](#system-requirements "Direct link to System requirements") The system must have: * Nvidia GPU(s) with at least 16GB of GPU memory * Nvidia drivers (see below for supported versions) * CUDA [compute capability](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#compute-capabilities) of 7.5-12.1 inclusive, which corresponds to the Turing, Ampere, Lovelace, Hopper, Blackwell architectures. Cards with the Volta architecture or below are not able to run the models * 24 GB RAM * The [nvidia-container-toolkit](https://github.com/NVIDIA/nvidia-docker) installed * Docker version > 19.03 The raw image size of the GPU Inference Container is around 15GB. See [Performance and cost](/deployments/container/performance-and-cost.md) for more information on the performance and cost of the container. ### Nvidia drivers[​](#nvidia-drivers "Direct link to Nvidia drivers") * 15.0.0 container version or higher: The GPU Inference Container is based on CUDA 13.0.1, which requires NVIDIA Driver release 580 or later. * 14.13.0 container version or lower: The GPU Inference Container is based on CUDA 12.4.1, which requires NVIDIA Driver release 525 or later. Driver installation can be validated by running `nvidia-smi`. This command should return the Nvidia driver version and show additional information about the GPU(s). #### Cloud instances[​](#cloud-instances "Direct link to Cloud instances") The GPU node can be provisioned in the cloud. ## Running the image[​](#running-the-image "Direct link to Running the image") Currently, each GPU Inference Container can only run on a single GPU. If a system has more than one GPU, the device must be specified using `CUDA_VISIBLE_DEVICES` or selecting the device using the `--gpus` argument. See [Nvidia/CUDA documentation for details](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#cuda-environment-variables). ``` docker run --rm -it \ -v $PWD/license.json:/license.json \ --gpus '"device=0"' \ -e CUDA_VISIBLE_DEVICES \ -p 8001:8001 \ speechmaticspublic.azurecr.io/sm-gpu-inference-server-en:15.19.0 ``` When the Container starts you should see output similar to this, indicating that the server has started and is ready to serve requests. ``` I1215 11:43:57.300390 1 server.cc:633] +----------------------+---------+--------+ | Model | Version | Status | +----------------------+---------+--------+ | am_en_enhanced | 1 | READY | | am_en_standard | 1 | READY | | body_enhanced | 1 | READY | | body_standard | 1 | READY | | diar_enhanced | 1 | READY | | diar_standard | 1 | READY | | ensemble_en_enhanced | 1 | READY | | ensemble_en_standard | 1 | READY | | lm_en_enhanced | 1 | READY | +----------------------+---------+--------+ ... I1215 11:43:57.375233 1 grpc_server.cc:4819] Started GRPCInferenceService at 0.0.0.0:8001 I1215 11:43:57.375473 1 http_server.cc:3477] Started HTTPService at 0.0.0.0:8000 I1215 11:43:57.417749 1 http_server.cc:184] Started Metrics Service at 0.0.0.0:8002 ``` ### Batch and Realtime inference[​](#batch-and-realtime-inference "Direct link to Batch and Realtime inference") The Inference server can run in two modes: *batch*, for processing whole files and returning the transcript at the end, and *real-time* for processing audio streams. The default mode is batch. To configure the GPU server for real-time, set the environment variable `SM_BATCH_MODE=false` by passing it into the `docker run` command. The modes correspond to the two types of client speech Container, which are distinguished by their name: * **rt**-asr-transcriber-en:\ * **batch**-asr-transcriber-en:\ The server can only support one of these modes at once. ### Linking to a GPU inference container[​](#linking-to-a-gpu-inference-container "Direct link to Linking to a GPU inference container") Once the GPU Server is running, follow the [Instructions for Linking a CPU Container](/deployments/container/cpu-speech-to-text.md#linking-to-a-gpu-inference-container). ### Running only one model[​](#running-only-one-model "Direct link to Running only one model") [Models](/speech-to-text/models.md) (previously called Operating Points) represent different levels of model complexity. To save GPU memory for throughput, you can run the server with only one model loaded. To do this, pass the `SM_MODEL` environment variable to the container and set it to either `standard` or `enhanced`. `SM_MODEL` replaces the older `SM_OPERATING_POINT` environment variable. `SM_OPERATING_POINT` is deprecated but still works and accepts the same `standard` and `enhanced` values; use `SM_MODEL` going forward. When running the all language standard model GPU inference server you must set the `SM_MODEL` environment variable to `standard` ### Loading only selected languages[​](#loading-only-selected-languages "Direct link to Loading only selected languages") Images that contain more than one language model load every language at startup. To save GPU memory for throughput, pass the `SM_LANGUAGES` environment variable to the container and set it to a comma-separated list of language codes. Only those languages are loaded: ``` # load only English, German, and French -e SM_LANGUAGES=en,de,fr ``` Codes must be languages the image contains. Any code the image does not contain causes startup to fail with an error. When `SM_LANGUAGES` is unset, all of the image's languages are loaded. `SM_LANGUAGES` is available from container version 15.18.0. `SM_LANGUAGES` and `SM_MODEL` can be used together. For example, `SM_MODEL=standard` with `SM_LANGUAGES=en,de` loads the Standard model for English and German only. ### Monitoring the server[​](#monitoring-the-server "Direct link to Monitoring the server") The inference server is based on [Nvidia's Triton architecture](https://developer.nvidia.com/nvidia-triton-inference-server) and as such can be monitored using Triton's inbuilt Prometheus metrics, or the GRPC/HTTP APIs. To expose these, configure an external mapping for port 8002(Prometheus) or 8000(HTTP). ### Models in GPU inference[​](#models-in-gpu-inference "Direct link to Models in GPU inference") When inference is outsourced to a GPU server, alternative GPU-specific models are used, so you should not expect to see identical results compared to CPU-based inference. For convenience, the GPU models are also designated as 'standard' and 'enhanced'. ## Docker compose example[​](#docker-compose-example "Direct link to Docker compose example") This Docker Compose file will create a Speechmatics GPU Inference Server: (assumes your `license.json` file is in the current working directory) docker-compose.yml ``` version: "3.8" networks: transcriber: driver: bridge services: triton: image: speechmaticspublic.azurecr.io/sm-gpu-inference-server-en:15.19.0 deploy: resources: reservations: devices: - driver: nvidia ### Limit to N GPUs # count: 1 ### Pick specific GPUs by device ID # device_ids: # - 0 # - 3 capabilities: - gpu container_name: triton networks: - transcriber expose: - 8000/tcp - 8001/tcp - 8002/tcp environment: - NVIDIA_DRIVER_CAPABILITIES=all - NVIDIA_VISIBLE_DEVICES=all - CUDA_VISIBLE_DEVICES=0 volumes: - $PWD/license.json:/license.json:ro ``` --- # GPU Speech to text container (Melia 1) Learn about the Speechmatics Melia 1 GPU container system Melia 1 is a multilingual, GPU-only model, currently available in Batch. For an overview of the model, see [Models](/speech-to-text/models.md#melia-1). ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [A license file or a license token](/deployments/container/licensing.md) * There is no specific license for the GPU Inference Container, it will run using an existing Speechmatics license for the Batch Container * [Access to our Docker repository](/deployments/container/accessing-images.md) ## System requirements[​](#system-requirements "Direct link to System requirements") The system must have: * Nvidia GPU(s) with at least 16GB of GPU memory * Nvidia drivers (see below for supported versions) * CUDA [compute capability](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#compute-capabilities) of 7.5-12.1 inclusive, which corresponds to the Turing, Ampere, Lovelace, Hopper, Blackwell architectures. Cards with the Volta architecture or below are not able to run the models * 24 GB RAM * The [nvidia-container-toolkit](https://github.com/NVIDIA/nvidia-docker) installed * Docker version > 19.03 See [Performance and cost](/deployments/container/performance-and-cost.md) for more information on the performance and cost of the container. ### Nvidia drivers[​](#nvidia-drivers "Direct link to Nvidia drivers") * The GPU Inference Container is based on CUDA 13.0.1, which requires NVIDIA Driver release 580 or later. Driver installation can be validated by running `nvidia-smi`. This command should return the Nvidia driver version and show additional information about the GPU(s). #### Cloud instances[​](#cloud-instances "Direct link to Cloud instances") The GPU node can be provisioned in the cloud. ## Running the image[​](#running-the-image "Direct link to Running the image") Currently, each GPU Inference Container can only run on a single GPU. If a system has more than one GPU, the device must be specified using `CUDA_VISIBLE_DEVICES` or selecting the device using the `--gpus` argument. See [Nvidia/CUDA documentation for details](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#cuda-environment-variables). Pull the Melia 1 inference server image as described in [Accessing images](/deployments/container/accessing-images.md#melia-1-model), then run it: ``` docker run --rm -it \ -v $PWD/license.json:/license.json \ --gpus '"device=0"' \ -e CUDA_VISIBLE_DEVICES \ -p 8001:8001 \ speechmaticspublic.azurecr.io/sm-asr-inference-server-melia-1:1.3.0 ``` When the Container starts you should see output similar to this, indicating that the server has started and is ready to serve requests. ``` I0705 08:12:55.419608 1 server.cc:709] +--------------+---------+--------+ | Model | Version | Status | +--------------+---------+--------+ | aed | 1 | READY | | body | 1 | READY | | detokenizer | 1 | READY | | dz | 1 | READY | | ensemble | 1 | READY | | preprocessor | 1 | READY | +--------------+---------+--------+ ... I0705 08:12:55.515473 1 grpc_server.cc:2579] "Started GRPCInferenceService at 0.0.0.0:8001" I0705 08:12:55.515940 1 http_server.cc:4961] "Started HTTPService at 0.0.0.0:8000" I0705 08:12:55.598672 1 http_server.cc:400] "Started Metrics Service at 0.0.0.0:8002" ``` ### Batch inference[​](#batch-inference "Direct link to Batch inference") The Melia 1 inference server currently runs in *batch* mode only, processing whole files and returning the transcript at the end. Once the inference server is running, run the Melia 1 transcriber with the `SM_INFERENCE_ENDPOINT` environment variable set to the gRPC endpoint of the inference server. The transcriber offloads inference to the server, so it does not require a GPU: ``` docker run --rm -it \ -e SM_INFERENCE_ENDPOINT=: \ -v $PWD/license.json:/license.json \ -v $PWD/example.wav:/input.audio \ ``` To accept multiple jobs over an HTTP API without restarting the container between jobs, run the transcriber as a [batch persistent worker](/deployments/container/batch-persistent-worker.md). ### Monitoring the server[​](#monitoring-the-server "Direct link to Monitoring the server") The inference server is based on [Nvidia's Triton architecture](https://developer.nvidia.com/nvidia-triton-inference-server) and as such can be monitored using Triton's inbuilt Prometheus metrics, or the GRPC/HTTP APIs. To expose these, configure an external mapping for port 8002(Prometheus) or 8000(HTTP). ## Docker compose example[​](#docker-compose-example "Direct link to Docker compose example") This Docker Compose file will create a Speechmatics Melia 1 GPU Inference Server: (assumes your `license.json` file is in the current working directory) docker-compose.yml ``` version: "3.8" networks: transcriber: driver: bridge services: triton: image: speechmaticspublic.azurecr.io/sm-asr-inference-server-melia-1:1.3.0 deploy: resources: reservations: devices: - driver: nvidia ### Limit to N GPUs # count: 1 ### Pick specific GPUs by device ID # device_ids: # - 0 # - 3 capabilities: - gpu container_name: triton networks: - transcriber expose: - 8000/tcp - 8001/tcp - 8002/tcp environment: - NVIDIA_DRIVER_CAPABILITIES=all - NVIDIA_VISIBLE_DEVICES=all - CUDA_VISIBLE_DEVICES=0 volumes: - $PWD/license.json:/license.json:ro ``` --- # Translation GPU inference container Learn about the Speechmatics Translation GPU container system ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [A license file or a license token](/deployments/container/licensing.md) * The Inference Container itself does not need a license, but its client (a transcriber Container) must have a valid license with Translation enabled * [Access to our Docker repository](/deployments/container/accessing-images.md) ## System requirements[​](#system-requirements "Direct link to System requirements") Note: System requirements for the Translation inference server are the same as for the [GPU Inference Container](/deployments/container/gpu-speech-to-text.md) for transcription, except for RAM and CPU requirements which are lower. The two servers cannot use the same GPU. The system must have: * Nvidia GPU(s) with at least 16GB of GPU memory * Nvidia drivers (see below for supported versions) * CUDA [compute capability](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#compute-capabilities) of 7.5-12.1 inclusive, which corresponds to the Turing, Ampere, Lovelace, Hopper, Blackwell architecture. Cards with the Volta architecture or below are not able to run the models * 5GB RAM * 4 vCPUs * The [nvidia-container-toolkit](https://github.com/NVIDIA/nvidia-docker) installed * Docker version > 19.03 The raw Docker image size of the Translation Container is around 10GB. ### Nvidia drivers[​](#nvidia-drivers "Direct link to Nvidia drivers") * 15.0.0 container version or higher: The GPU Inference Container is based on CUDA 13.0.1, which requires NVIDIA Driver release 580 or later. * 14.13.0 container version or lower: The GPU Inference Container is based on CUDA 12.4.1, which requires NVIDIA Driver release 525 or later. Driver installation can be validated by running `nvidia-smi`. This command should return the Nvidia driver version and show additional information about the GPU(s). #### Cloud instances[​](#cloud-instances "Direct link to Cloud instances") The GPU node can be provisioned in the cloud. ## Running the image[​](#running-the-image "Direct link to Running the image") Currently, each Translation Container can only run on a single GPU. If a system has more than one GPU, the device must be specified using `CUDA_VISIBLE_DEVICES` or selecting the device using the `--gpus` argument. See [Nvidia/CUDA documentation for details](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#cuda-environment-variables). ``` docker run --rm -it \ --gpus '"device=0"' \ -e CUDA_VISIBLE_DEVICES \ -p 8001:8001 \ # the grpc endpoint uses port 8001, can be mapped to any host port speechmaticspublic.azurecr.io/sm-translation-inference-server:15.19.0 ``` On startup you will see logs detailing available GPU memory. As set out in the [requirements](#system-requirements) section, the system must have a minimum of 16GB of GPU memory, though extra GPU memory may be used if available. ``` Total GPU memory: 40960MiB Approx. size models: 5GB Available GPU memory after models loaded: 35GB ``` ## Sending requests[​](#sending-requests "Direct link to Sending requests") Batch and Realtime (RT) transcribers handle sending requests to the Translation Inference Server. To run a transcription job with Translation, follow the [instructions for running the CPU Container](/deployments/container/cpu-speech-to-text.md) and additionally: * Set the environment variable `SM_TRANSLATION_ENDPOINT` in the transcriber to the GRPC endpoint of the running Translation Inference Server, in the form `:` where the port is the one bound to port 8001 of the Translation Docker Container (see [running the image](#running-the-image)) * Include a `translation_config` inside of your job config. [More details](/speech-to-text/features/translation.md) * Use a transcriber version 10.3.0 or newer * Ensure you use a license which allows Translation ### Translation language pairs[​](#translation-language-pairs "Direct link to Translation language pairs") The Translation Inference Container is not language specific, meaning that all 69 translation language pairs supported can run on a single Inference Container. The source language is defined by the language of the transcriber sending requests. By default, a maximum of 5 target languages can be requested at once. This behaviour can be changed by setting the environment variable `SM_TRANSLATION_MAX_TARGET_LANGUAGES` in the transcriber. Setting this to 0 will disable the limit. ### Example of running translation[​](#example-of-running-translation "Direct link to Example of running translation") Assuming the following config file: ``` { "type": "transcription", "transcription_config": { "model": "enhanced", "language": "en" }, "translation_config": { "target_languages": ["es", "de"] // Set languages here to enable translation } } ``` You can run Batch Transcription and Translation with: ``` cat ~/$AUDIO_FILE | docker run -i \ -v ~/$CONFIG_FILE:/config.json \ -e LICENSE_TOKEN=eyJhbGciOiJ... \ -e SM_TRANSLATION_ENDPOINT=: \ batch-asr-transcriber-en:15.19.0 ``` Or start a Translation enabled Realtime Container with: ``` docker run -p 9000:9000 -e LICENSE_TOKEN=eyJhbGciOiJ... \ -e SM_TRANSLATION_ENDPOINT=: \ -e SM_TRANSLATION_MAX_TARGET_LANGUAGES=10 \ # raise the allowed number of target languages rt-asr-transcriber-en:15.19.0 ``` ### Monitoring the server[​](#monitoring-the-server "Direct link to Monitoring the server") The Inference Server is based on [Nvidia's Triton architecture](https://developer.nvidia.com/nvidia-triton-inference-server) and as such can be monitored using Triton's inbuilt Prometheus metrics, or the GRPC/HTTP APIs. To expose these, configure an external mapping for port 8002(Prometheus) or 8000(HTTP). ## Docker compose example[​](#docker-compose-example "Direct link to Docker compose example") This docker-compose file will create a Speechmatics GPU translation server: docker-compose.yml ``` --- version: '3.8' networks: transcriber: driver: bridge services: triton: image: speechmaticspublic.azurecr.io/sm-translation-inference-server:15.19.0 deploy: resources: reservations: devices: - driver: nvidia ### Limit to N GPUs # count: 1 ### Pick specific GPUs by device ID # device_ids: # - 0 # - 3 capabilities: - gpu container_name: triton networks: - transcriber expose: - 8000/tcp - 8001/tcp - 8002/tcp environment: - NVIDIA_DRIVER_CAPABILITIES=all - NVIDIA_VISIBLE_DEVICES=all - CUDA_VISIBLE_DEVICES=0 ``` ## Error handling[​](#error-handling "Direct link to Error handling") ### Unsupported target language (Batch)[​](#unsupported-target-language-batch "Direct link to Unsupported target language (Batch)") If one or more of the target languages are not supported for the source language, an error message will be included in the final JSON output. No translations will be returned for that language pair. ``` { "job": { ... }, "metadata": { "created_at": "2023-05-26T15:01:48.412714Z", "type": "transcription", "transcription_config": {...}, "translation_config": { "target_languages": [ "es", "zz" ] }, "translation_errors": [ {"type": "unsupported_translation_pair", "message": "Translation from en to zz currently not supported"} ], ... }, "results": [...] } ``` Please note, this behaviour is different when using our SaaS Deployment. For all other errors, please see our documentation. --- # Language ID container Learn about the Speechmatics language ID Container This guide will walk you through the steps needed to deploy the Speechmatics Batch Language Identification Container. Looking for how to use this in cloud SaaS? See the documentation [here](/speech-to-text/batch/language-identification.md). This Container will allow you to predict the most likely, predominant language spoken in a media file. You can use the predicted language to select the correct transcriber when the language spoken in your file is unknown. The following steps are required to use this in your environment: * Check system requirements * Pull the Docker Image into your local Docker Registry * Run the Container ## Prerequisites[​](#prerequisites "Direct link to Prerequisites") * [A license file or a license token with Language ID enabled](/deployments/container/licensing.md) * [Access to our Docker repository](/deployments/container/accessing-images.md) * Audio file (we recommend having at least 60 seconds of speech for high accuracy) If you do not have a license or access to the Docker repository, please contact reach out to [Support](https://support.speechmatics.com). ## System requirements[​](#system-requirements "Direct link to System requirements") Speechmatics Containerized deployments are built on the Docker platform. A single Docker image can be used to create and run multiple Containers concurrently, for each running Container the following resources are required: * 1 vCPU * 1 GB RAM * The host machine should have an AMD or Intel CPU with modern AVX instructions as this generally improves transcription processing speed. Exact impact on transcription speed varies on brand, generation, which instruction sets are available and resource allocation * If you are using a hypervisor, you should pass through to the VM all AVX related instruction sets The raw image size of the Language Identification Container is around 2.1GB. ## Workflow[​](#workflow "Direct link to Workflow") * Run the Language ID Docker Container with an audio file * Receive the output JSON with the predicted language code * Use that language code to run transcription with any of the Speechmatics deployments ## Licensing[​](#licensing "Direct link to Licensing") You should have received a confidential license file from Speechmatics containing a token to use to license your Container. The contents of the file received should look similar to this: ``` { "contractid": 1, "creationdate": "2022-06-01 09:04:11", "customer": "Speechmatics", "id": "c18a4eb990b143agadeb384cbj7b04c3", "metadata": { "key_pair_id": 1, "request": { "customer": "Speechmatics", "features": ["MAPBA", "ALID"], "notValidAfter": "2023-01-01", "validFrom": "2022-01-01" } }, "signedclaimstoken": "example" } ``` There are two ways to apply the license to the Container. * As a volume-mapped file The license file should be mapped to the path `/license.json` within the Container. For example: ``` docker run ... -v /my_license.json:/license.json:ro speechmaticspublic.azurecr.io/langid:2.2.1 ``` * As an environment variable Setting an environment variable named `LICENSE_TOKEN` is also a valid way to license the Container. The contents of this variable should be set to the value of the `signedclaimstoken` from within the license file. For example, copy the `signedclaimstoken` from the license file (without the quotation marks) and set the environment variable as below. The token example is not a full example: ``` docker run ... -e LICENSE_TOKEN=eyJhbGciOiJ... speechmaticspublic.azurecr.io/langid:2.2.1 ``` There should be no reason to do this, but if both a volume-mapped file and an environment variable are provided simultaneously then the volume-mapped file will be ignored. ## Using the container[​](#using-the-container "Direct link to Using the container") To reliably identify the predominant language, the file should contain at least 60 seconds of speech in that language. Once the Docker image has been pulled into a local environment, it can be started using the Docker run command. More details about operating and managing the Container are available in the [Docker API](https://docs.docker.com/engine/api/latest) documentation. There are two different methods for passing a media file into a Container: * STDIN: Streams media file into the Container through the standard command line entry point * File Location: Pulls media file from a file location Here are some examples below to demonstrate these modes of operating the Container. Example 1: passing a file using the cat command to the Container ``` cat ~/$AUDIO_FILE | docker run -i -e LICENSE_TOKEN=eyJhbGciOiJ... speechmaticspublic.azurecr.io/langid:2.2.1 ``` Example 2: pulling a media file from a mapped directory into the Container ``` docker run -v $AUDIO_FILE:/input.audio -e LICENSE_TOKEN=eyJhbGciOiJ... speechmaticspublic.azurecr.io/langid:2.2.1 ``` The media file must be volume-mapped into the Container path `/input.audio` Both the methods will produce the same identification result. STDOUT is used to provide the result in JSON format. Here's an example of the returned JSON: ``` { "format": "1.1", "metadata": { "created_at": "2023-08-30T10:45:27+0000", "type": "language_identification", "language_identification_config": {}, "duration": 60.029388, "processed_duration": 60 }, "results": [ { "alternatives": [ { "language": "cs", "confidence": 0.94 }, { "language": "sk", "confidence": 0.02 }, { "language": "uk", "confidence": 0.02 }, { "language": "pl", "confidence": 0.01 }, { "language": "en", "confidence": 0 }, { "language": "sl", "confidence": 0 }, { "language": "el", "confidence": 0 }, { "language": "bg", "confidence": 0 }, { "language": "be", "confidence": 0 }, { "language": "ru", "confidence": 0 } ], "start_time": 0, "end_time": 60.03 } ], "predicted_language": "cs" } ``` In the regular case the predicted language code will be in the `predicted_language` field. The `alternatives` in `results` contains the top 10 predicted languages based on the confidence score. A list of possible [Language Codes can be found here](/speech-to-text/languages.md). The following languages are not supported for Language Identification: Interlingua (ia), Esperanto (eo), Uyghur (ug), Cantonese (yue). In case the language can't be identified, the `error` field contains one of the following reasons: * `LOW_CONFIDENCE`: The language can't be determined with sufficient confidence * `UNEXPECTED_LANGUAGE`: The language identified is not among the `expected_languages` list * `NO_SPEECH`: The audio file does not contain any speech * `FILE_UNREADABLE`: Failure to read the file. E.g. due to an unsupported audio format, Container exits with exit code 1 * `OTHER`: Generic error with details provided in `message` field, Container exits with exit code 1 For example the response for input with no speech looks like: ``` { "format": "1.1", "metadata": { "created_at": "2023-09-01T12:49:27+0000", "type": "language_identification", "language_identification_config": {}, "duration": 183.913313, "processed_duration": 90 }, "results": [], "error": "NO_SPEECH", "message": "No speech found for language identification" } ``` ### Setting expected languages[​](#setting-expected-languages "Direct link to Setting expected languages") If you expect the audio to be one of a restricted set of languages, you can provide this information through the `expected_languages` config. You can either specify them as comma-separated string as CLI argument: ``` docker run -v $AUDIO_FILE:/input.audio -e LICENSE_TOKEN=eyJhbGciOiJ... speechmaticspublic.azurecr.io/langid:2.2.1 --expected-languages cs,sk,en ``` or provide the list in a JSON config file: ``` { "type": "language_identification", "language_identification_config": { "expected_languages": ["cs", "sk", "en"] } } ``` The config needs to be volume-mapped into `/config.json` to apply the configuration to the identification: ``` docker run -v $(pwd)/config.json:/config.json -v $AUDIO_FILE:/input.audio -e LICENSE_TOKEN=eyJhbGciOiJ... speechmaticspublic.azurecr.io/langid:2.2.1 ``` ## Ability to run a container with multiple cores[​](#ability-to-run-a-container-with-multiple-cores "Direct link to Ability to run a container with multiple cores") For customers who are looking to improve job turnaround time and who are able to assign sufficient resources, it is possible to pass a parallel parameter to the Container to take advantage of multiple CPUs. The parameter is called `parallel` and the following example shows how it can be used. In this case to use 2 cores to process the audio you would run the Container like this: ``` docker run -v $AUDIO_FILE:/input.audio -e LICENSE_TOKEN=eyJhbGciOiJ... speechmaticspublic.azurecr.io/langid:2.2.1 --parallel=2 ``` Depending on your hardware, you may need to experiment to find the optimum performance. We've noticed an improvement in turnaround time for jobs by using this approach. If you limit or are limited on the number of CPUs you can use (for example your platform places restrictions on the number of cores you can use, or you use the --cpu flag in your docker run command), then you should ensure that you do not set the parallel value to be more than the number of available cores. If you attempt to use a setting in excess of your free resources, then the Container will only use the available cores. If you are running the Container on a shared resource, you may experience different results depending on what other processes are running at the same time. ## Determining success[​](#determining-success "Direct link to Determining success") The exit code of the Container will determine if the identification was successful. There are two exit code possibilities: * `Exit Code == 0` : The identification was a success; the output will contain a JSON output defining the identification result * `Exit Code != 0` : the output will contain useful information why the job failed. This output should be used in any communication with [Speechmatics Support](https://support.speechmatics.com) to aid understanding and resolution of any problems that may occur ## Limitations[​](#limitations "Direct link to Limitations") * It's not possible to predict the language of each channel independently in a multichannel media file; any multichannel files are converted to mono before identifying the language * Inverted multichannel audio is not supported, this is where the second channel is the inverse of first * The Container uses CPU and doesn't run on a GPU ## Enable logging[​](#enable-logging "Direct link to Enable logging") If you are seeing problems then we recommend that you [reach out to Support](https://support.speechmatics.com). Please include the logging output from the Container if you do open a ticket, and ideally enable verbose logging. Verbose logging is enabled by running the Container with the argument `-vv`. All logs are written to STDERR. ``` docker run ... speechmaticspublic.azurecr.io/langid:2.2.1 -vv ``` --- # Licensing Learn about the licensing for Speechmatics containers **To** run our containers you need a license. To get a license you need to speak to [Support](https://support.speechmatics.com). The license is easy to pass into the container deployment either as a file or a environment variable. Licensing does not require network access and works in offline deployments. ## License structure[​](#license-structure "Direct link to License structure") The contents of the file received should look similar to this: ``` { "contractid": 1, "creationdate": "2020-03-24 17:43:35", "customer": "Speechmatics", "id": "c18a4eb990b143agadeb384cbj7b04c3", "is_trial": true, "metadata": { "key_pair_id": 1, "request": { "customer": "Speechmatics", "features": ["MAPBA", "LANY"], "isTrial": true, "notValidAfter": "2021-01-01", "validFrom": "2020-01-01" } }, "signedclaimstoken": "exampleClaimsToken" } ``` The `validFrom` and `notValidAfter` keys in the license file specify the start and end dates for the validity of your license. The license is valid from 00:00 UTC on the start date to 00:00 UTC on the expiry date. After the expiry date, the Container will continue to run but will not transcribe audio. You should apply for a new license before this happens. ## Volume mapping in a license file[​](#volume-mapping-in-a-license-file "Direct link to Volume mapping in a license file") The license file should be mapped to the path `/license.json` within the Container. For example: ``` docker run -v ./my_license.json:/license.json batch-asr-transcriber-en:15.19.0 ``` ## Setting the license with an environment variable[​](#setting-the-license-with-an-environment-variable "Direct link to Setting the license with an environment variable") For example, copy the `signedclaimstoken` value from the license file and set the `LICENSE_TOKEN` environment variable: ``` docker run -e LICENSE_TOKEN=exampleClaimsToken batch-asr-transcriber-en:15.19.0 ``` There should be no reason to do this, but if both a volume-mapped file and an environment variable are provided simultaneously then the volume-mapped file will be ignored. --- # Performance and cost Get an overview of the performance and cost of Speechmatics container deployments ## Speech to text containers[​](#speech-to-text-containers "Direct link to Speech to text containers") This is a comparison of the performance and estimated running costs of transcription executing on standard Azure VMs. The comparison highlights the maximum number of concurrent real-time sessions (session density) and the maximum throughput for batch jobs on a single instance. ### Batch transcription[​](#batch-transcription "Direct link to Batch transcription") | Models | [CPU Standard](/deployments/container/cpu-speech-to-text.md) | [CPU Enhanced](/deployments/container/cpu-speech-to-text.md) | [GPU Standard](/deployments/container/gpu-speech-to-text.md) | [GPU Enhanced](/deployments/container/gpu-speech-to-text.md) | [GPU Melia 1](/deployments/container/gpu-speech-to-text-melia-1.md) | | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------------- | | Lowest Processing Cost (US ¢ per hour) | 1.7 | 3.8 | 0.34 | 1.67 | 0.19 | | Cost vs CPU Standard (%) | - | 224% | 20% | 98% | 11% | | Cost vs CPU Enhanced (%) | 45% | - | 9% | 44% | 5% | | Cost vs GPU Standard (%) | 500% | 1118% | - | 491% | 56% | | Cost vs GPU Enhanced (%) | 102% | 228% | 20% | - | 11% | | Maximum Throughput[1](#user-content-fn-1) | 53.2 | 23.7 | 350 | 45 | 3600 | | Representative Real-Time Factor (RTF)[2](#user-content-fn-2) | 0.085 | 0.2 | 0.034 | 0.120 | 0.018 | | Job Concurrency | 20 | 20 | 15[3](#user-content-fn-3) | 6[3](#user-content-fn-3) | 70[3](#user-content-fn-3) | The benchmark uses the following configuration: | Benchmark details | | | ----------------- | -------------------------------------------- | | Version | 15.7.0 (Standard, Enhanced), 1.3.0 (Melia-1) | | Language | English only | | CPU | D16ds\_v5 | | GPU Standard | Standard\_NC16as\_T4\_v3 | | GPU Enhanced | Standard\_NC8as\_T4\_v3 | | GPU Melia-1 | Standard\_NC40ads\_H100\_v5 | | Price Basis | Azure PAYG East US, Linux, Standard | For GPU Models, transcribers and inference servers were all run on a single VM node. The Batch transcription results above were measured with diarization disabled. ### Realtime transcription[​](#realtime-transcription "Direct link to Realtime transcription") | Models | [CPU Standard](/deployments/container/cpu-speech-to-text.md#realtime-transcription) | [CPU Enhanced](/deployments/container/cpu-speech-to-text.md#realtime-transcription) | [GPU Standard](/deployments/container/gpu-speech-to-text.md#batch-and-realtime-inference) | [GPU Enhanced](/deployments/container/gpu-speech-to-text.md#batch-and-realtime-inference) | | -------------------------------------- | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | | Lowest Processing Cost (US ¢ per hour) | 1.97 | 2.95 | 0.86 | 2.51 | | Cost vs. CPU Standard (%) | - | 150% | 44% | 127% | | Cost vs. CPU Enhanced (%) | 67% | - | 29% | 85% | | Session Density | 40 | 24 | 140[3](#user-content-fn-3) | 30[3](#user-content-fn-3) | This benchmark uses the following configuration[4](#user-content-fn-4): | Benchmark details | Value | | ----------------- | ----------------------------------- | | Version | 13.4.0 | | Language | English only | | CPU | D16ds\_v5 | | GPU Standard | Standard\_NC16as\_T4\_v3 | | GPU Enhanced | Standard\_NC8as\_T4\_v3 | | Price Basis | Azure PAYG East US, Linux, Standard | For GPU Models, the transcribers and inference servers were run on a single VM node. Each first session, transcriber requires 0.25 cores for both OPs, with 1.2 GB memory (Standard OP) or 3 GB memory (Enhanced OP). Every additional session consumes 0.1 cores and 100 MB of memory. The RT transcription results above were measured with diarization disabled. ## Translation (GPU)[​](#translation-gpu "Direct link to Translation (GPU)") [Translation](/deployments/container/gpu-translation.md) running on a 4-core T4 has an RTF of roughly 0.008. It can handle up to 125 hours of batch audio per hour, or 125 Realtime Transcription streams. However, each translation target language is counted as a stream, meaning that a single Realtime Transcription stream which requests 5 target languages adds the same load on the Translation Inference Server as 5 transcription streams each requesting a single target language. ## Footnotes[​](#footnote-label "Direct link to Footnotes") 1. Throughput is measured as hours of audio per hour of runtime. A throughput of 50 would mean that in one hour, the system as a whole can transcribe 50 hours of audio. [↩](#user-content-fnref-1) 2. An RTF of 1 would mean that a one hour file would take one hour to transcribe. An RTF of 0.1 would mean that a one hour file would take six minutes to transcribe. Benchmark RTFs are representative for processing audio files over 20 minutes in duration using `parallel=4`. [↩](#user-content-fnref-2) 3. Multiple jobs (batch mode) or sessions (Realtime mode) are handled by a single worker configured with the required concurrency. [↩](#user-content-fnref-3) [↩2](#user-content-fnref-3-2) [↩3](#user-content-fnref-3-3) [↩4](#user-content-fnref-3-4) [↩5](#user-content-fnref-3-5) 4. Benchmark results reflect performance on a fully loaded inference server operating at the session density recommended for the respective GPU platform. [↩](#user-content-fnref-4) --- # Speaker identification secrets Prepare and manage Speaker Identification secrets for Speechmatics deployments [Speaker identification](/speech-to-text/features/speaker-identification.md) requires the Batch or Realtime Transcriber to access one or more cryptographically strong secrets. These secrets are used to generate and validate speaker identifiers. This page explains how to prepare these secret files and mount them into the Transcriber Container in both Batch and Realtime deployments. Batch speaker identification requires GPU inference. It is not supported by the [CPU speech-to-text container](/deployments/container/cpu-speech-to-text.md) in batch mode. For batch processing, use the [GPU speech-to-text container](/deployments/container/gpu-speech-to-text.md). Realtime speaker identification is supported on both CPU and GPU containers. ## Preparing the secret directory[​](#preparing-the-secret-directory "Direct link to Preparing the secret directory") Speaker ID secrets must be stored in a dedicated directory on the host machine. Each secret is contained in a separate file whose name follows the pattern: `/s.`. For example: ``` /speaker_id_secrets/s.1 /speaker_id_secrets/s.2 ``` Secret files may contain either binary data or plain text (for example, Base64-encoded values). The Transcriber will read all files matching the `s.` pattern in the directory. Multiple secret files allow operators to rotate secrets without disrupting active workloads. When more than one secret is present, the Transcriber will continue to accept identifiers encrypted with older secrets while using the most recent secret for generating new identifiers. The most recent secret is determined by the highest-numbered secret file. In most on-prem deployments, operators typically only need to maintain a single secret file unless secret rotation is required for their environment. ## Mounting the secret directory[​](#mounting-the-secret-directory "Direct link to Mounting the secret directory") The secret directory must be mounted into the Transcriber Container as a read-only volume. The location inside the Container must then be provided to the Transcriber using the `SM_SPEAKER_ID_SECRETS_DIR` environment variable. Below is an example for Docker: ``` docker run --rm -i -v :/speaker_id_secrets:ro -e SM_SPEAKER_ID_SECRETS_DIR=/speaker_id_secrets -e LICENSE_TOKEN=$TOKEN_VALUE $IMAGE_NAME ``` The Transcriber will automatically load all secret files in the specified directory when it starts. --- # Troubleshooting Troubleshooting for Speechmatics containers ## Batch troubleshooting[​](#batch-troubleshooting "Direct link to Batch troubleshooting") ### Enabling logging[​](#enabling-logging "Direct link to Enabling logging") If you are seeing problems then we recommend that you enable logging and [reach out to Support](https://support.speechmatics.com). The following example shows how to enable logging, using the `-stderr` argument to output the logs to `stderr`: ``` docker run --rm -e SM_JOB_ID=123 -e SM_LOG_DIR=/logs \ -v ~/$AUDIO_FILE:/input.audio \ -e LICENSE_TOKEN=f787b0051e2768b1f619d75faab97f23ee9b7931890c05f97e9f550702 \ batch-asr-transcriber-en:15.19.0 \ -stderr ``` To store the output of logs, add two environment variables: * `SM_JOB_ID`: - a job id, for example: 1 * `SM_LOG_DIR`: - the directory inside the container where to write the logs, for example: `/logs` When raising a Support Ticket it is normally easier to write the log output to a specific file. You can do this by creating a volume mount where the logs will be accessible from after the Container has finished. Before running the Container you need to create a directory for the log file and ensure it has the correct permissions. In this example we use a local logs directory to store the output of the log for a job with ID 124: ``` mkdir -p logs/124 / sudo chown -R nobody:nogroup logs/ sudo chmod -R a+rwx logs/ ``` then ``` docker run --rm -v ${PWD}/logs:/logs -e SM_JOB_ID=124 -e SM_LOG_DIR=/logs \ -v ~/sm_audio.wav:/input.audio \ -e LICENSE_TOKEN=f787b0051e2768b1f619d75faab97f23ee9b7931890c05f97e9f550702 \ batch-asr-transcriber-en:15.19.0 tail logs/124/sigurd.log ``` ### Common problems[​](#common-problems "Direct link to Common problems") There are occasions where the transcription container will fail to transcribe the media file provided and will exit without error code 0 (success). Speechmatics heavily advise enabling logging (see instruction above). The logs will show the reasons for the failed job especially when multiple errors can cause the same error code. Below are some errors with suggestions and how they can be revolved. | Error Code | Error | Resolution | | ---------- | ------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 1 | “err: signal: illegal instruction” | This means that the models couldn't be loaded within the Container. Please ensure that the host that's running the Docker engine has an AVX compatible CPU.

The following can also be done inside the Container to check that AVX is listed in the CPU flags.

`$ docker run -it --entrypoint /bin/bash batch-asr-transcriber-en:{smVariables.latestContainerVersion}`

`$ cat /proc/cpuinfo \| grep flags` | | 1 | “Unable to set up logging” | This can occur when a directory is volume mapped into the Containers and a log file cannot be created into that directory.

Example command to map in a tmp directory inside the container to /xxx path:

`$ docker run --rm -e SM_LOG_DIR=/xxx -e SM_JOB_ID=1 -v $PWD/tmp:/xxx batch-asr-transcriber-en:{smVariables.latestContainerVersion}` | | 1 | “/input.audio is not valid” | If volume mapping the file into the Container, ensure that a valid audio file is being mapped in. | | 1 | “failed to get sample rate” | The sample rate from the audio file that was passed for recognition did not have a sample rate. Check the audio file is valid and that a sample rate can be read.

The following ffmpeg can be used to identify if there is a valid sample rate:

`$ ffmpeg -i /home/user/example.wav` | | 1 | “exit status 1” | If the container is memory (RAM) starved it can quit during the transcription process. Verify the minimum resource (CPU and RAM) requirements are being assigned to a Transcription Container.

The inspect command in Docker can be useful to identify if the lack of memory shutdown the container. Look out for the “OOMKilled” value. Here is an example.

`$ docker inspect --format='{{json.state}}' $containerID` | | 1 | "License Error: illegal base64 data at input byte $NUMBER" | The license token value has been truncated or otherwise altered from the initial value generated. Please ensure that you have copied token value correctly or that the license file is not corrupt | | 1 | "ERROR sentryserver could not load license: stat /license.json: no such file or directory" | The license file or license token has not been passed when attempting to run the Container. Please ensure that the license file or license token value is passed as documented | | 2 | --parallel/-parallel: invalid check\_parallel value: '0' | If using the parallel option to speed up the processing time on files more than 5 minutes in length the -–parallel switch needs to have an integer at least 1. A non-zero value must be provided if the parallel command is to be used.

The example below shows a valid command:

`$ docker run -i -v /home/user/config.json:/config.json -v /home/user/example.wav:/input.audio -e LICENSE_TOKEN=$TOKEN_VALUE batch-asr-transcriber-en:{smVariables.latestContainerVersion} --parallel 2` | If you still continue to face issues, please [reach out to Support](https://support.speechmatics.com). ## Realtime troubleshooting[​](#realtime-troubleshooting "Direct link to Realtime troubleshooting") ### Enabling logging[​](#enabling-logging-1 "Direct link to Enabling logging") If you are seeing problems then we recommend that you [reach out to Support](https://support.speechmatics.com). Please include the logging output from the Container if you do open a ticket, and ideally enable verbose logging. Verbose logging is enabled by running the Container with the environment variable `DEBUG` set to `true`. e.g. ``` docker run -e DEBUG=true rt-asr-transcriber-en:15.19.0 ``` ### Licensing[​](#licensing "Direct link to Licensing") The best way to identify licensing errors with the Container is to look at the container logs. See for more information about doing this. If licensing is successful then the logs upon startup should look similar to this: ``` INFO:__main__:Starting health service INFO:orchestrator.health:Health check server starting... INFO:__main__:Health service started. INFO:orchestrator.license:Starting sentry server... time="2020-03-27T11:50:18.9774596Z" level=info msg="Listening to port 52000, secure mode = false" time="2020-03-27T11:50:18.9776369Z" level=info msg="Reading license from /license.json" time="2020-03-27T11:50:18.9866595Z" level=info msg="Read token eyJkbGciOjJS..." INFO:orchestrator.license:Sentry server started time="2020-03-27T11:50:18.990334Z" level=info msg="License : licensed=true, customer=Speechmatics, contract_id=0, expires_at=2021-03-16 00:00:00 +0000 UTC, trial=false, features=MAPRT,MAPBA,AMCC,APD,APR,ASS" time="2020-03-27T11:50:18.9904803Z" level=info msg="Starting server 3.0.0 [master]" time="2020-03-27T11:50:18.9918058Z" level=info msg="Monitoring parent pid 1" 2020-03-27 11:50:19,005 orchestrator.transport.ws.common INFO Waiting for the model to be ready - checking /model/manifest.json 2020-03-27 11:50:20,673 orchestrator.transport.ws.common INFO Loading model en 2020-03-27 11:50:26,107 orchestrator.transport.ws.ws INFO transport websocket listening at ws://0.0.0.0:9000 2020-03-27 11:50:26,107 orchestrator.transport.ws.health_update INFO Transport marked as started for health updates. ``` If your Container is not licensed, or has an invalid license then it will exit upon startup with an error message similar to this: ``` RuntimeError: Failed to launch sentry server licensing process on port 52000 ``` Please ensure that you have correctly followed the instructions in the quick start guide for setting up licensing, and that have you a license file which has not expired (the `metadata` section in the file tells you when the license is valid until). There can be several reasons for a licensing error: * No license has been provided If you see the following message in the container logs then the most likely cause is that no license file has been provided: ``` level=error msg="could not load license file data: stat /license.json: no such file or directory" ``` Please review the quick start guide and ensure that the license has been provided properly, either as a volume-mapped file or as an environment variable. * The license has expired ``` level=info msg="License : licensed=false, customer=Speechmatics, contract_id=99, expires_at=2020-03-26 00:00:00 +0000 UTC, trial=false, features=" level=error msg="Error in license : token is expired by 36h6m37s" ``` This message indicates that your license has expired. Please request a new license from Speechmatics Support. * You are attempting to use a feature for which you are not licensed Not all licenses are valid for all features of our product. If you are not licensed for a feature which you attempt to use for transcription, then transcription will not be performed. Please [get in touch with Speechmatics Support](https://support.speechmatics.com) if you are interested in using a feature which you are not licensed for. If this error case happens you should see a log message similar to this one: ``` 2020-03-27 12:11:04,230 orchestrator.transport.ws.protocol WARNING Sending an error to client: not_allowed - Unable to use provided configuration: No license for requested language - LEN; session ID de1ec62d-a22d-47a3-8f03-def025a52f60 ``` * An improperly formatted license file has been provided *Only relevant if using a volume-mapped file to license the container* ``` level=error msg="could not load license file data: unexpected end of JSON input" ``` or ``` level=error msg="could not load license file data: No valid signedclaimstoken field found in license (too short)" ``` Please ensure that you are using the license file which has been provided to you by the Speechmatics support team, and that no changes have been made to the file accidentally. The license file should be a valid JSON file and should contain a key named `signedclaimstoken` which is your license token. ### Common problems[​](#common-problems-1 "Direct link to Common problems") You should ensure, when using the config object in the `StartRecognition` message, that the JSON is correctly formatted. --- # Kubernetes Learn about the Kubernetes deployment options for Speechmatics Speechmatics applications can be deployed on Kubernetes using Helm charts, providing a flexible and consistent deployment option for cloud-native environments. The Helm charts package all required Kubernetes resources and configuration options, making it easier to install, configure, and manage Speechmatics applications. Using Helm, customers can customize deployments through configurable values, manage upgrades and rollbacks, and integrate Speechmatics applications into existing Kubernetes workflows. ## Quickstart[​](#quickstart "Direct link to Quickstart") #### [Realtime Kubernetes](/deployments/kubernetes/realtime.md) [Kubernetes deployment options for Realtime](/deployments/kubernetes/realtime.md) ## Supported applications[​](#supported-applications "Direct link to Supported applications") Speechmatics Kubernetes deployment supports the following applications: * [Realtime](/speech-to-text/realtime/quickstart.md): Stream audio from an input device or file and receive real-time transcription updates as audio is processed. --- # Prerequisites Prerequisites for deploying Realtime on Kubernetes ## Access to Docker and Helm registry[​](#access-to-docker-and-helm-registry "Direct link to Access to Docker and Helm registry") It is important to create and reference resources within the same namespace if you do not have full control of your Kubernetes cluster or if it is a shared cluster The Speechmatics Docker images and Helm chart are obtained from the `speechmaticspublic.azurecr.io` registry. If you do not have credentials for the registry account or have lost your details, please reach out to [Support](https://support.speechmatics.com) for help. Add a `docker-registry` secret to the cluster so it can successfully pull Speechmatics docker images. ``` # Add the speechmatics registry credentials for image pulling kubectl create secret docker-registry speechmatics-registry \ --docker-server=speechmaticspublic.azurecr.io \ --docker-username= \ --docker-password= ``` Using the same credentials, authenticate Helm with the `speechmaticspublic` registry to install the Helm chart: ``` # Authenticate to the Speechmatics Helm repository helm registry login speechmaticspublic.azurecr.io \ --username \ --password ``` ## Speechmatics license[​](#speechmatics-license "Direct link to Speechmatics license") Please speak to `support@speechmatics.com` if you do not already have a valid Speechmatics license. The Helm chart requires you to have a valid Speechmatics license stored in a secret called `speechmatics-license`, on the Kubernetes cluster. You can add a secret to the cluster with these commands: ``` kubectl create secret generic speechmatics-license \ --from-literal=license.json="$(cat $LICENSE_FILE)" ``` Alternatively, you can configure the chart to create the secret for your Speechmatics license secret for you using the following values: ``` global: licensing: createSecret: true license: $B64_ENCODED_LICENSE ``` ## GPU drivers[​](#gpu-drivers "Direct link to GPU drivers") The Speechmatics inference server runs Nvidia Triton Server, which requires an Nvidia GPU. When running GPU nodes in Kubernetes, you will require the Nvidia device plugin which allows containers on the cluster to access the GPUs. Refer to the [Nvidia driver requirements](/deployments/container/gpu-speech-to-text.md#nvidia-drivers) for the supported driver releases corresponding to different Speechmatics container versions. Below is a list of the common cloud providers and their recommended way of deploying the Nvidia device plugin on a cluster: * AWS - [Amazon EKS Setup](https://docs.aws.amazon.com/deep-learning-containers/latest/devguide/deep-learning-containers-eks-setup.html#deep-learning-containers-eks-setup-gpu-clusters) * Azure - [GPU Workloads on AKS](https://learn.microsoft.com/en-us/azure/aks/gpu-cluster?tabs=add-ubuntu-gpu-node-pool#nvidia-device-plugin-installation) * GCP - [Deploy GPU workloads in GKE](https://cloud.google.com/kubernetes-engine/docs/how-to/gpus#installing_drivers) Alternatively, see the [Nvidia Device Plugin](https://github.com/NVIDIA/k8s-device-plugin#deployment-via-helm) docs. You can validate a node has allocatable GPU resources with: ``` kubectl get nodes -o yaml | yq .[].[].status.allocatable | grep nvidia ``` ## Nginx ingress controller[​](#nginx-ingress-controller "Direct link to Nginx ingress controller") When setting up Speechmatics via an ingress controller, it is recommended to use the `ingress-nginx` ingress controller with snippet annotations enabled. You can confirm if your cluster supports Nginx with snippet annotations enabled using the following command: ``` # The default is false kubectl get cm -o yaml -l app.kubernetes.io/instance=nginx | grep allow-snippet-annotations ``` If you are not already running `ingress-nginx`, follow the below steps: 1. Create an `nginx.values.yaml` file: ``` controller: service: # This is needed to preserve the source IP of the client # See: https://kubernetes.io/docs/tasks/access-application-cluster/create-external-load-balancer/#preserving-the-client-source-ip externalTrafficPolicy: Local # This should be set to the IP address used to access services on your cluster loadBalancerIP: $CLUSTER_INGRESS_IP config: # This prevents nginx worker process from shutting down for 24h in case of active sessions worker-shutdown-timeout: 86400s extraArgs: # Needed to support annotations added by the chart ingresses annotations-prefix: nginx.ingress.kubernetes.io # This prevents nginx main process from shutting down for 24h in case of active sessions shutdown-grace-period: 86400 # Used to prevent nginx pods being terminated for 24h while there are active sessions terminationGracePeriodSeconds: 86400 # Needed to allow ingresses to add snippet annotations to add necessary headers allowSnippetAnnotations: true ``` 2. Install the Nginx chart with: ``` helm repo add nginx https://kubernetes.github.io/ingress-nginx helm install nginx nginx/ingress-nginx --version 4.11.4 -f nginx.values.yaml ``` ### Using another ingress controller[​](#using-another-ingress-controller "Direct link to Using another ingress controller") If you are running another ingress controller, when enabling ingress on the chart, you need to ensure that a `Request-Id` header is passed through. This is used to manage session usage. In Nginx, it looks like this: ``` proxy: ingress: annotations: # Add headers to all requests coming through this ingress nginx.ingress.kubernetes.io/configuration-snippet: |+ more_set_headers "Request-Id: $req_id"; ``` --- # Realtime Learn about the Kubernetes deployment options for Realtime ## Quickstart[​](#quickstart "Direct link to Quickstart") ### Install[​](#install "Direct link to Install") Providing the [Prerequisites](/deployments/kubernetes/prerequisites.md) have been met for the Speechmatics Helm chart, use the command below to install: ``` # Install the sm-realtime chart helm upgrade --install speechmatics-realtime \ oci://speechmaticspublic.azurecr.io/sm-charts/sm-realtime \ --version 1.4.0 \ --set proxy.ingress.hostname="speechmatics.example.com" ``` `proxy.ingress.url` Helm value is deprecated in favour of `proxy.ingress.hostname` for configuring the Ingress hostname. Existing deployments remain compatible: if `proxy.ingress.hostname` is not set, the chart continues to honour `proxy.ingress.url` as a fallback. ### Validate[​](#validate "Direct link to Validate") #### Capacity check[​](#capacity-check "Direct link to Capacity check") You can confirm whether the transcribers and inference servers are available using: ``` kubectl get sessiongroups ``` If the transcribers and inference servers are available, it will show `CAPACITY` meaning that they have successfully registered. ``` NAME REPLICAS CAPACITY USAGE VERSION SPEC HASH inference-server-enhanced-recipe1 1 480 0 1 b5784af49332f9948481195451eab6ca rt-transcriber-en 1 2 0 1 83929f2b9b2448cdc818d0e46e37600b ``` #### Run a session[​](#run-a-session "Direct link to Run a session") ``` speechmatics rt transcribe \ --url wss://speechmatics.example.com/v2 \ --lang en \ --operating-point enhanced \ --ssl-mode insecure \ ``` ## Hardware recommendations[​](#hardware-recommendations "Direct link to Hardware recommendations") Below are the recommended Azure node sizes for running Realtime on Kubernetes: | Service | Node Size | | ------------------ | ----------------------- | | Inference Server | Standard\_NC4as\_T4\_v3 | | Transcriber | Standard\_E16s\_v5 | | All Other Services | Standard\_D\*s\_v5 | ## Configuration[​](#configuration "Direct link to Configuration") For detailed configuration options, refer to sm-realtime Helm chart README.md See the examples below on how to configure the Helm chart for different deployment scenarios. All LanguagesAll LanguagesEnglish Standard + EnhancedEnglish Standard + EnhancedAuto-ScalingAuto-Scaling ``` global: transcriber: languages: ["ar", "ba", "be", "bg", "bn", "ca", "cmn", "cmn_en", "cmn_en_ms_ta", "cs", "cy", "da", "de", "el", "en", "en_ms", "en_ta", "eo", "es", "es-bilingual-en", "et", "eu", "fa", "fi", "fr", "ga", "gl", "he", "hi", "hr", "hu", "ia", "id", "it", "ja", "ko", "lt", "lv", "mn", "mr", "ms", "mt", "nl", "no", "pl", "pt", "ro", "ru", "sk", "sl", "sv", "sw", "ta", "th", "tl", "tr", "ug", "uk", "ur", "vi", "yue"] # Enable all enhanced and standard inference server recipes inferenceServerEnhancedRecipe1: enabled: true inferenceServerEnhancedRecipe2: enabled: true inferenceServerEnhancedRecipe3: enabled: true inferenceServerEnhancedRecipe4: enabled: true inferenceServerStandardAll: enabled: true ``` ``` # Disable default enhanced inference server recipes inferenceServerEnhancedRecipe1: enabled: false # Enable custom inference server deployment with just en models inferenceServerCustom: enabled: true fullnameOverride: inference-server-en tritonServer: image: # Repository for the en-only inference server triton container repository: sm-gpu-inference-server-en inferenceSidecar: enabled: true # Configuration for custom model deployments registerFeatures: capacity: 600 languages: ["en"] operatingPoint: ["standard", "enhanced"] ``` ``` global: # Enable scaling for all sessiongroups resources sessionGroups: scaling: enabled: true inferenceServerEnhancedRecipe1: sessionGroups: scaling: # Scale up inference server pods when there are 300 inference tokens remaining scaleOnCapacityLeft: 300 transcribers: sessionGroups: scaling: # Scale up transcriber pods when there is only capacity for 1 more session scaleOnCapacityLeft: 1 ``` ## Uninstall[​](#uninstall "Direct link to Uninstall") Run the following command to uninstall Realtime from the cluster: ``` helm uninstall speechmatics-realtime ``` Depending on the configuration setup, you may also need to remove PVCs created from the redis deployment: ``` # Delete any left-over PVCs with `kubectl delete pvc` kubectl get pvc | grep redis-data ``` --- # Usage reporting Learn about the usage reporting for on-prem deployments For on-prem deployments, usage reporting is required for accurate billing. There are two types of usage reporting we offer: **automatic** and **offline**. #### [Automatic](/deployments/usage-reporting/automatic.md) [Learn more about automatic usage reporting](/deployments/usage-reporting/automatic.md) #### [Offline](/deployments/usage-reporting/offline.md) [Learn how to set up offline usage reporting using a usage container](/deployments/usage-reporting/offline.md) ## How the reporting mode is configured[​](#how-the-reporting-mode-is-configured "Direct link to How the reporting mode is configured") Both modes are controlled by the same two settings. `SM_EATS_URL` sets the destination for usage events. `SM_ENABLE_USAGE_REPORTING` sets whether the transcriber restricts events to the subset needed for billing (`true`, the default) or sends the full activity event stream (`false`). If `SM_EATS_URL` is not set, the destination depends on `SM_ENABLE_USAGE_REPORTING`: at its default of `true`, the transcriber falls back to `usage.speechmatics.com`; set to `false`, there is no destination and reporting is disabled entirely. Automatic reporting relies on the default destination, with `SM_ENABLE_USAGE_REPORTING` left at its default of `true`. Offline reporting sets `SM_EATS_URL` to your own [Usage Container](/deployments/usage-reporting/offline.md#container) instead, while `SM_ENABLE_USAGE_REPORTING` remains at its default — the same billing events are sent to your container rather than to Speechmatics. ## What data do we record?[​](#what-data-do-we-record "Direct link to What data do we record?") We will **never** send customer audio data over the network. We only record metadata about what our transcriber is doing. We aim to be completely transparent about the kind of data we record. In future releases, we may improve our service by recording additional technical data in relation to the operation of the transcriber. This will not include sensitive customer data, such as audio files or transcripts. ### Data we record:[​](#data-we-record "Direct link to Data we record:") * Audio duration information to accurately calculate billing * Customer information including: Contract ID, Customer name and License ID * Summary of job configuration information, excluding any potentially sensitive information such as Custom Dictionary content * Audio information such as codec and format type * System information, including certain error events We record events about what the transcriber is doing as small JSON objects. Here is an example of the kind of data we send: ``` { "container_id": "aee112010359", "contract_id": "0", "customer_id": "Speechmatics", "engine_instance_id": "930ab901-c408-457d-9f6d-dee21ecaac7f", "event": "TRANSCRIBER_DONE", "event_specific_data": { "bit_rate": "512000", "build_id": "", "channels": 2, "codec_name": "pcm_s16le", "file_metadata": {}, "format_name": "wav", "inference": "cpu", "job_config": { "transcription_config": { "language": "en" }, "type": "transcription" }, "license_id": "b6616911c59c40cbbb32ebf442696a14", "rtf": 1.975401759147644, "sample_rate": "16000", "total_received_bytes": 149562, "total_received_duration": 2, "total_speech_duration": 2 }, "session_id": "a29bd398-5745-43ca-9d16-07cdc6eabdd0", "source": "uniasr-batch", "time_stamp": "2022-09-14T13:01:13.239Z" } ``` --- # Automatic usage reporting Learn about automatic usage reporting for on-prem deployments Automatic usage reporting (a.k.a. online usage reporting) automatically sends usage data from the transcriber to Speechmatics via `usage.speechmatics.com`. Speechmatics will use this data to appropriately bill users for their usage. Depending on the on-prem solution you require, the setup process is slightly different see the sections below on [containers](#container) and [appliances](#virtual-appliance) for details. We will **never** send customer audio data over the network. See [What Data Do We Record](/deployments/usage-reporting/.md#what-data-do-we-record) for a full description of what information will be recorded. ## Technical details[​](#technical-details "Direct link to Technical details") The Batch transcriber will report one `TRANSCRIBER_DONE` event at the event of transcription. The Realtime transcriber will report one `SESSION_ENDED` event at the end of each session. During a session, the Realtime transcriber also sends `SESSION_STATUS` every few minutes. The payload size is only several KB, so it won’t have a meaningful impact on the duration of transcription or your bandwidth costs. If usage reporting is successful then at the end of the session the following message will be visible in the transcriber logs: ``` 2024-04-15 13:17:38,799 INFO sentryserver Transcribed 36 seconds of speech 2024-04-15 13:17:38,974 INFO prod.orchestrator.eats.api Eats client process finished ``` In verbose debug mode, the contents of the usage report can be seen. ## Network failure[​](#network-failure "Direct link to Network failure") In the event of a network failure (for example, if your Internet connection is down or our usage server has a temporary outage) the transcriber will attempt to reconnect to our usage server several times. ``` 2024-04-15 13:21:15,556 WARNING prod.orchestrator.eats.api Exception from EATS API call [TRANSCRIBER_DONE] HTTPSConnectionPool(host='usage.speechmatics.com', port=443): Max retries exceeded with url: /v1/log (Caused by NameResolutionError(": Failed to resolve 'usage.speechmatics.com' ([Errno -3] Temporary failure in name resolution)")) ``` If, after this retry period (which takes up to 3 seconds), the transcriber is still unable to contact our usage server then it will output some `WARNING` log messages then cease attempting to send usage information. If this happens then the transcriber will exit normally with an exit code of 0. For Batch transcribers, the transcriber will exit immediately after this. For Realtime transcribers, usage reporting will be disabled for a fixed time period (currently 60 seconds). This is to minimize the impact on the duration of transcription jobs. This retry mechanism will cause a small hit to the speed of transcription, so in the event of a network outage, you may wish to temporarily disable usage reporting by setting the `SM_ENABLE_USAGE_REPORTING` variable to false when running the container. We ask that you inform our Finance Team about the duration and timing of any such outage. ## Deployment type[​](#deployment-type "Direct link to Deployment type") ### Container[​](#container "Direct link to Container") For container deployments, automatic usage reporting is turned ON by default and is currently opt out. It is turned off by setting the environment variable `SM_ENABLE_USAGE_REPORTING=false` (`false`, `no` or `0` are equally valid) when running the transcriber. For example: ``` docker run -i -v ~/$AUDIO_FILE:/input.audio \ -e LICENSE_TOKEN=eyJhbGciOiJ... \ -e SM_ENABLE_USAGE_REPORTING=false \ batch-asr-transcriber-en:15.19.0 ``` To enable automatic usage reporting, you must be running one of the following ASR Container versions: * Batch Container `10.1.0` onwards * Realtime Container `10.1.0` onwards Automatic usage reporting is turned ON by default, starting from version `10.6.0`. `SM_EATS_URL` overrides the destination for these reports. Set it to redirect automatic reporting to your own [Usage Container](/deployments/usage-reporting/offline.md#container) instead of `usage.speechmatics.com`, without changing which events are sent. See [Offline usage reporting](/deployments/usage-reporting/offline.md) for details. ### Virtual appliance[​](#virtual-appliance "Direct link to Virtual appliance") For virtual appliances, automatic usage reporting is turned ON by default and is currently opt out. The usage mode can be set via the Management API ``` curl -L -u admin:admin -X 'POST' \ "http://${APPLIANCE_HOST}/v2/management/usagereporting" \ -d '{"mode": "online"}' ``` where mode is one of `offline` | `online` --- # Offline usage reporting Learn about offline usage reporting for on-prem deployments To support on-prem solutions which require access to the internet to be blocked, usage data can be collected locally using a usage container. This data is then exported and sent to Speechmatics via email for review. Depending on the on-prem solution you require, the setup process is slightly different. See the sections below on [Containers](#container) or [Appliances](#virtual-appliance) for details. The Usage Container only collates data that is required for Speechmatics to calculate accurate financial billing and measure product usage and system performance. This data is made up of a series of events that correspond to the various stages of a Speechmatics Batch or Realtime Container as it processes a media file. No personal customer data, transcripts or media data is captured or stored at any point. See [What Data Do We Record](/deployments/usage-reporting/.md#what-data-do-we-record) for a full description of what information will be recorded. The customer is responsible for assigning storage to the Usage Container and or Batch Appliance in order to capture all usage information, and sending data to Speechmatics at regular intervals. ## Reporting cadence[​](#reporting-cadence "Direct link to Reporting cadence") Speechmatics requires customers to send all usage data by the last working day of each calendar month. You should send data for each Usage Container and/or batch appliance you have running in your environments. For customers with very large transcription volumes, more regular reporting may be recommended. Large transcription volumes can mean: * Large number of jobs * This means any Usage Container that will store data from more than 10,000 Batch jobs in a calendar month, or 1250 Realtime jobs of more than an hour * Many jobs of long duration (>60 minutes), especially when using the Realtime Container in a 'streaming' mode where it persists between sessions ## Sending data to Speechmatics[​](#sending-data-to-speechmatics "Direct link to Sending data to Speechmatics") The exported data must not be modified in any way before sending to Speechmatics. Speechmatics will request a new unmodified data export if it is found that data has been altered. Data is retained in the Usage Container for **90 days**, after which point it is purged. After exporting, Speechmatics requires data to be sent via email to . Speechmatics recommends file sizes to not exceed 25MB. This is the default limit for sending emails for many popular providers like Microsoft Office 365. Files in excess of this size may trigger an error when sending by your email provider. You will receive a confirmation email within 15 minutes if the report(s) get accepted by our billing system. If the "Reply-To" header on the email you send contains multiple email addresses, we will send a reply email to only the first address in the list. Any attachment sent to Speechmatics must have the correct file name extension: `.json.gz`. For details on how to export usage data for a given on-prem solution see here [Container](/deployments/usage-reporting/offline.md#container) and [Appliance](/deployments/usage-reporting/offline.md#virtual-appliance) for more details. ## Deployment type[​](#deployment-type "Direct link to Deployment type") ### Virtual appliance[​](#virtual-appliance "Direct link to Virtual appliance") The usage mode can be set via the Management API ``` curl -L -u admin:admin -X 'POST' \ "http://${APPLIANCE_HOST}/v2/management/usagereporting" \ -d '{"mode": "offline"}' ``` where mode is either `offline` or `online`. When the usage mode is set to `offline`, usage will be collected via a Container inside the appliance, the data collected by this Container will need to be sent to Speechmatics via email at . #### Workflow[​](#workflow "Direct link to Workflow") The following workflow is recommended: * The user downloads and runs one or more of the Virtual Appliances * Before running any jobs, the user sets the usage mode to `offline` (see above). * At intervals of no more than a calendar month, the user will extract usage data processed in that interval from each running Appliance via the Management API, see [below](#exporting-usage-data) * The user will then send this data to a designated Speechmatics email address at . #### Exporting usage data[​](#exporting-usage-data "Direct link to Exporting usage data") The exported data must not be modified in any way before sending to Speechmatics. Speechmatics will request a new unmodified data export if it is found that data has been altered. Data is retained in the appliance for **90 days**, after which point it is purged. Exported data needs to be sent to via email to . A compressed archive of the usage data can be retrieved via the Management API In `realtime` mode you may see usage for contract id `-77777777777777`, this is a prewarming job that runs during the transcriber first startup, and will not be included in your billed usage. ``` curl -X 'GET' \ "http://${APPLIANCE_HOST}/v1/export?since={start_time}&until={end_time}" \ -H 'accept: application/gzip' ``` Where `start_time` and `end_time` are inclusive and are timestamps in the [ISO-8601 format](https://www.iso.org/iso-8601-date-and-time-format.html) (YYYY-MM-DDTHH:MM:SSZ). To remain under the 25MB email attachment limit, we recommend changing `start_time` and `end_time` to chunk exports into 25MB files (usually around 10,000 batch jobs or 1250 real-time sessions of one hour). Data is exported in compressed `json.gz` format. All files must be sent in this format to Speechmatics. The name of the file does not matter. You can send multiple attachments per email, or each email as a separate attachment, so long as you are under email provider limits for sending files. ### Container[​](#container "Direct link to Container") #### Terminology[​](#terminology "Direct link to Terminology") Throughout this section there are references to different types of containers: * ASR Containers - Speechmatics containers that transcribe media or audio files into a transcript. Two types are available - those can process media in batch, and those that can process media in real-time. When these are specifically referred to they are called the Batch or Realtime Containers * Usage Containers - a new container that stores event-specific data from ASR Containers #### Getting started[​](#getting-started "Direct link to Getting started") The ASR Usage Container can be retrieved from Speechmatics Docker Registry as a Docker Image. To access the Usage Container, you should use the same credentials that you use to access Speechmatics' ASR Containers from its Docker Registry. This information should already be provided to you by [Support](https://support.speechmatics.com) when you are onboarded. You will also need to know the following information: * Docker Registry URL, e.g. `https://speechmaticspublic.azurecr.io` * Image name, e.g. `asr-usage` * Image tag, e.g. `0.3.0` The image can be downloaded by using the standard Docker workflow: ``` # Login docker login https://speechmaticspublic.azurecr.io ### Download image docker pull speechmaticspublic.azurecr.io/asr-usage:0.3.0 ``` Speechmatics require all customers to cache a copy of the Docker images within their own environment. Please do not pull directly from the Speechmatics docker registry for each deployment. #### System requirements[​](#system-requirements "Direct link to System requirements") The ASR Usage Container requires the following resources: * 1 vCPU * 1 GB memory * At least 1 GB of persistent storage per Usage Container deployed. Every 25 MB can store data for up to 13,000 batch jobs or up to 1250 (60 minute) Realtime sessions. Persisting storage to temporary locations (e.g. `tmpfs`) is supported where this is necessary as part of a user's workflow, but is not recommended. If you are required to use `tmpfs` or other such directories as a storage solution, Speechmatics recommends increasing the frequency of how often usage reports are sent to avoid any potential data loss #### Configuration[​](#configuration "Direct link to Configuration") The following section will show you how to set up an environment where you have a running ASR Usage Container that can accept data from one or multiple ASR Containers. It will show in order: * How to set up and run an ASR Usage Container * How to ensure an ASR Batch or Realtime Container can send all required data to an ASR Usage Container during transcription. You must set up a Usage Container before running Speechmatics' Batch or Realtime ASR Containers in order to ensure that all usage data is captured. A Usage Container is persistent, which means it does not shut down after receiving transcription data. #### Prerequisites[​](#prerequisites "Direct link to Prerequisites") When setting up an environment with one or multiple Speechmatics ASR Container(s) and one or multiple Usage Container(s) please ensure: * That all Batch or Realtime Containers you require to send data to the Usage Container can exchange communication with each other in their environment * That all communication between Docker containers is via HTTPS * That when running Usage Containers, you enable the required ports when necessary to send and extract data. More detail is below #### Compatibility[​](#compatibility "Direct link to Compatibility") To use the Usage Container, you must be running the following ASR Container versions: * Batch Container 8.2.0 onwards * Realtime Container 1.4.1 onwards The ASR Usage Container has been tested using Docker Version 20. Compatibility with previous versions of Docker has not been tested. #### Early access[​](#early-access "Direct link to Early access") The Usage Container has been released as an early access product that any customer using either Speechmatics' Batch or Realtime ASR Containers is entitled to use. Speechmatics encourages customers to try this solution, in order to simplify their usage logging and reporting processes. Speechmatics encourages feedback on the Usage Container, and the raising of any bugs or usability issues. These will be subject to our normal bug triage process, and should be submitted to [Support](https://support.speechmatics.com). #### Workflow[​](#workflow-1 "Direct link to Workflow") The following workflow is recommended: * The user downloads the Usage Container from Speechmatics' Docker Registry using their existing credentials * The user must cache a copy of each Container they download within their own environment * The user will run one or multiple Usage Container(s) depending on their requirements * Any Usage Container must be assigned its own persistent data volume. The user is responsible for allocating and backing up this persistent volume * Any Usage Container must also have all relevant ports opened to allow data exchange and export where necessary * When requesting transcription from a Batch or Realtime Container, the user must specify the hostname or IP address of the Usage Container via a new environment variable * Data will then be stored in the Usage Container for up to 90 days * At intervals of no more than a calendar month, the user will extract usage data processed in that interval from the ASR Usage Container via the RESTful API * The user will then send this data to a designated Speechmatics [email address](mailto:billing-reporting@speechmatics.com) #### ASR usage container[​](#asr-usage-container "Direct link to ASR usage container") The ASR Usage Container always requires a persistent storage volume to store the data. This volume must be mounted inside the container at `/data`. #### Endpoints[​](#endpoints "Direct link to Endpoints") The ASR Usage Container has 2 endpoints: | Endpoint | Use | Port | How to Set | | ----------- | ----------------------------------------------------------------------------------------------------- | ---- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `v1/log` | Receives transcription event data from Batch and Realtime Containers | 9090 | use the `SM_EATS_URL` environment variable | | `v1/export` | All event data, or time-specific event data, can be extracted from this endpoint as a compressed file | 8000 | use the `docker -p $PORT:$PORT` command. If you need to change the default port use `-e PANDAS_PORT` environment variable as well as `docker -p` with your required port | By default, all Docker Containers do not expose any ports. You must specifically request these ports to be open to ensure transcription events are captured, or that data can be extracted. The example below starts a Usage Container with * A persistent volume mounted to `/data` * Port 9090 open via the `EATS_PORT` environment variable to allow the Container to accept transcription event data * Port 8000 open via the `docker -p` command to allow data to be exported from the Usage Container ``` # Create volume docker volume create volume-1 # Mount volume docker run -it \ -v volume-1:/data \ -e EATS_PORT=9090 \ -p 8000:8000 \ speechmaticspublic.azurecr.io/asr-usage:0.3.0 ``` Further documentation on using persistent storage volumes on popular container orchestration engines: #### [Kubernetes](https://kubernetes.io/docs/concepts/storage/persistent-volumes) #### [Nomad](https://www.nomadproject.io/docs/job-specification/volume) Speechmatics recommends setting up backup policies for the persistent volume. The ASR Usage Container cannot perform recovery by itself if the data file or volume is corrupted. The ASR Usage Container accepts the following configuration option, which can be set via environment variables. | Key | Default | Type | Description | | ----------- | ------- | ---- | -------------------------------------------------------------------------------------------------------------------------- | | `EATS_PORT` | `9090` | int | Listening port for incoming data from transcribers. Must be set to accept usage data from Batch or Realtime ASR Containers | #### Sending transcription data from ASR container to ASR usage container[​](#sending-transcription-data-from-asr-container-to-asr-usage-container "Direct link to Sending transcription data from ASR container to ASR usage container") An ASR Container must be explicitly configured to send data to the ASR Usage Container when starting. By default, this is via HTTPS. The following configuration options **must** be specified when running the ASR Container to send usage data: | Key | Default | Type | Description | | ------------- | ------- | ------ | ------------------------------------------------------------------------------ | | `SM_EATS_URL` | none | string | Address and listening port of the ASR Usage Container you wish to send data to | To correctly configure the transcriber, set `SM_EATS_URL` environment variable to point to ASR Usage Container. e.g., `SM_EATS_URL=asr-usage.example.net:9090` or `SM_EATS_URL=10.244.8.32:9090`, where `asr-usage.example.net` and `10.244.8.32` correspond to the relevant ASR Usage Container instance. The port `9090` is the default listening port for incoming data from transcribers. The port number is alterable by using the `EATS_PORT` environment variable. `SM_EATS_URL` can be set alongside `SM_ENABLE_USAGE_REPORTING` — leave `SM_ENABLE_USAGE_REPORTING` at its default of `true` so the transcriber sends the same billing events described in [Automatic usage reporting](/deployments/usage-reporting/automatic.md#technical-details) to your Usage Container instead of to Speechmatics. Below is a working example of running an ASR Batch Container that will then send transcription event data to a running ASR Usage Container: ``` docker run -i -v $AUDIO_FILE:/input.audio \ -v $CONFIG_FILE:/config.json:ro \ -e LICENSE_TOKEN=$TOKEN_VALUE \ -e SM_EATS_URL=-asr-usage.example.net:9090 batch-asr-transcriber-en:8.2.0 ``` Below is a similar example of a Realtime Container that will send transcription event data to a running ASR Usage Container: ``` docker run -p 9000:9000 \ -e LICENSE_TOKEN=$TOKEN_VALUE \ -e SM_EATS_URL=asr-usage.example.net:9090 \ rt-asr-transcriber-en:1.4.0 ``` #### Logging[​](#logging "Direct link to Logging") The Usage Container will log event data sent by an ASR Batch or Realtime Container: * during transcription * when transcription has finished * For the Batch Container this is when transcription finishes as it is not a persistent container. * For the Realtime Container this is both after a endOfTranscription websocket message. The Realtime Container will send a `SESSION_ENDED` message to the Usage Container * when the Container is shut down or terminated by the user or due to system error during transcription itself (e.g. SIGTERM) #### Example logs - Success[​](#example-logs---success "Direct link to Example logs - Success") The following is an example of a log from by a Batch or Realtime ASR Container when they successfully send data to the ASR Usage Container: ``` 2021-07-19 11:24:31.314 INFO sentryserver Transcription usage registered with EATS ``` The following is an example of a log from the Usage Container when it successfully receives data from a Batch or Realtime ASR Container: ``` [2021-08-25T10:45:37Z INFO actix_web::middleware::logger] 172.19.0.3:39068 "POST /v1/log HTTP/1.1" 201 0 "-" "Go-http-client/1.1" 0.009459 ``` The following is an example of a log from the Usage Container when a customer successfully exports data: ``` [2021-09-03T14:54:00Z INFO actix_web::middleware::logger] 172.19.0.1:55820 "GET /v1/export HTTP/1.1" 200 12912 "-" "curl/7.64.1" 0.006313 ``` #### Example logs - Failure[​](#example-logs---failure "Direct link to Example logs - Failure") If data cannot be sent from the ASR Container to the ASR Usage Container, the following error message is shown in the ASR Container: ``` 2021-07-19 11:27:43.158 ERROR sentryserver Error 'Post "https://asr-usage.net:9090/v1/log": dial tcp 172.25.0.2:909: connect: connection refused' occurred when logging EATS data: retrying ``` #### Example logs - Failure upon container termination[​](#example-logs---failure-upon-container-termination "Direct link to Example logs - Failure upon container termination") If a Container is shut down or terminated, both the Batch and the Realtime Container will attempt retries for up to 1 minute after receiving `SIGTERM`. For Batch, the Container will attempt to send data when transcription finishes. For the Realtime Container, this is when Container termination is requested. After this point, any unsent data is lost with following message. ``` 2021-07-19 11:28:55.288 WARNING sentryserver Some activity events could not be sent to EATS: count: 4 ``` #### Orchestrating multiple ASR usage containers[​](#orchestrating-multiple-asr-usage-containers "Direct link to Orchestrating multiple ASR usage containers") It is up to the customer's level of risk tolerance and their internal topology and orchestration how many ASR Usage Containers they need to deploy in ratio to their number of ASR Containers. Speechmatics recommends that each environment in which Batch or Realtime Containers are deployed requires at least one Usage Container. Customers can implement multiple Usage Containers in each environment for redundancy and to reduce the risk of failure. If a customer has ASR containers in multiple availability zones or clusters, assigning Usage Containers per environment or cluster reduces latency and the requirement to send messages between clusters. Orchestrating multiple ASR Usage Containers allows redundancy in the event of network or storage failure. It is possible to deploy multiple ASR Usage Containers in a single environment and have usage data distributed to those Containers. A basic scenario example is below, The `docker-compose` example below illustrate this scenario with: * Two Speechmatics ASR Usage Containers * One proxy Container, to route telemetry data * One Speechmatics ASR Batch Container docker-compose.yml ``` # Example docker-compose file using multiple telemetry Containers. # version: "3.4" # Common setup for ASR Usage Containers x-usage-template: &usage-template image: asr-usage:x.y.z labels: - "traefik.enable=true" - "traefik.tcp.routers.usage.rule=HostSNI(`*`)" - "traefik.tcp.routers.usage.entrypoints=custom" - "traefik.tcp.routers.usage.tls=true" - "traefik.tcp.routers.usage.tls.passthrough=true" - "traefik.tcp.routers.usage.service=telemeter" - "traefik.tcp.services.usage.loadbalancer.server.port=9090" depends_on: - proxy services: # Traefik reverse proxy, to route telemetry events to multiple ASR Usage Container # containers proxy: image: traefik:v2.4 command: --providers.docker --providers.docker.exposedByDefault=false --entrypoints.custom.address=:9090 volumes: - /var/run/docker.sock:/var/run/docker.sock usagecontainer1: <<: *usage-template ports: - "8001:8000" usagecontainer2: <<: *usage-template ports: - "8002:8000" transcriber-batch: image: batch-asr-transcriber-en:x.y.z environment: SM_EATS_URL: proxy:9090 volumes: - ./input/10_sec_news.wav:/input.audio - ./input/license.json:/license.json depends_on: - proxy - usagecontainer1 - usagecontainer2 ``` The example configures the ASR Batch Container to send data, using `SM_EATS_URL`, to the proxy container instead of a specific ASR Usage Container. When receiving usage data, the proxy will forward it to one ASR Usage Container, using round robin balancing. Each ASR Usage Container will need its own persistent storage volume to store usage data. This means that when generating reports to send to Speechmatics, an export request must be made for each ASR Usage Container the user has in operation. There will be as many reports as there are ASR Usage Containers deployed. #### Exporting usage data[​](#exporting-usage-data-1 "Direct link to Exporting usage data") The exported data must not be modified in any way before sending to Speechmatics. Speechmatics will request a new unmodified data export if it is found that data has been altered. Data is retained in the Usage Container for **90 days**, after which point it is purged. Exported data needs to be sent to via email to . The data must be exported from each ASR Usage Container you have used, and then sent to Speechmatics for calculation. The ASR Usage Container has a REST API to export transcription data. You will need to send at least as many reports as you have from ASR Usage Containers. Based on heavy transcription usage, you may have to provide multiple reports per single ASR Usage Container. To remain under the 25MB email attachment limit, we recommend changing `start_time` and `end_time` to chunk exports into 25MB files (usually around 10,000 batch jobs or 1250 real-time sessions of one hour). Data is exported in compressed `json.gz` format. All files must be sent in this format to Speechmatics. The name of the file does not matter. You can send multiple attachments per email, or each email as a separate attachment, so long as you are under email provider limits for sending files. The complete API reference for extracting usage data can be found in the [API Reference](/api-ref/batch/get-usage-statistics.md) section. ``` # To export all data curl 'asr-usage.net:8000/v1/export' > ExportExampleFile.json.gz # To export data within a date window, e.g. 1-Jan-2020 to 1-Feb-2020 curl 'asr-usage.net:8000/v1/export?since=2020-01-01T00:00:00.000000Z&until=2020-02-01T00:00:00Z' > ExportExampleFile-01-01_2020-02-01.json.gz ``` If the number of jobs extracted is too large a 4XX response may be returned. Generally this has been shown in testing to be circa. 25,000 Batch jobs or 5,000 Realtime jobs of an hour long. In such cases, please select a smaller time window with `since` and `until` parameters. It is fine to have overlapping reports with duplicate data. Transcriptions will always be billed once; the billing cycle will be determined by their time of completion. The following example script exports reports by each week for the whole month: export-usage-data.sh ``` #!/bin/bash # Use ISO-8601 format START="2021-11-01T00:00:00Z" END="2021-12-01T00:00:00Z" CHUNK="7 day" d=$(date -d "$START" -I) while [ $(date -d "$d" +%s) -le $(date -d "$END" +"%s") ]; do SINCE=$(date -d "$d" +"%Y-%m-%dT%H:%M:%SZ") d=$(date -I -d "$d + $CHUNK") UNTIL=$(date -d "$d" +"%Y-%m-%dT%H:%M:%SZ") curl "asr-usage.net:8000/v1/export?since=${SINCE}&until=${UNTIL}" > exported_$(date -d "${SINCE}" -I)_$(date -d "${UNTIL}" -I).json.gz done ``` The exported usage data is a compressed JSON file; it is possible to inspect the contents by unpacking it and opening the text file. The following example uses the [jq](https://stedolan.github.io/jq/) JSON parser. ``` $ cat exported_2020-01-01_2020-02-01.json.gz | gunzip | jq . { "header": { "alg": "HS512" }, "payload": { "events": [ { ... ``` --- # Virtual Appliance Deploy Speechmatics to your own hardware. Virtual appliances are pre-configured virtual machines that can be deployed to your own hardware. They are designed to be easy to use and manage, and are a great way to get started with Speechmatics. --- # Adding Languages Add languages to a Virtual Appliance deployment As of 08 October 2025 we have moved our public containers from artifactory to Azure Container Registry (ACR), please [reach out to Support](https://support.speechmatics.com) to update your credentials. Adding languages is currently only supported for `batch` mode, if you need extra realtime languages please [reach out to Support](https://support.speechmatics.com). If you have access to the Speechmatics Container Repository, you can install additional CPU-only languages on the Appliance. Log into the appliance using SSH with the username `smadmin`, as described in [Remote Access](/deployments/virtual-appliance/administration/remote-access.md). At the commandline, run the `configure_languages.sh` script. You will be prompted for credentials for the Speechmatics Container Repository. You will need to know which version of the images are compatible with your appliance, this will be documented here, but until then [reach out to Support](https://support.speechmatics.com). Language codes are the standard two-letter ISO codes used in configuring transcription, for example: "de", "fr". Appliance versions up to and including `6.2.1` will pull from our artifactory repository by default, as of 08 October 2025 we will deprecate public access to this repository (see above). If running an appliance version `6.2.1` or earlier, you will need to configure your appliance to pull from the ACR repository instead. ``` export SM_DOCKER_PUBLIC=speechmaticspublic.azurecr.io ``` Access to the Speechmatics repository is needed to download new languages, these credentials can be set in the environment but need to be base64 encoded, alternately if unset the user will be prompted for the username and password after running the configure languages script. If you need access to the Speechmatics container repository please [reach out to Support](https://support.speechmatics.com). ``` export CRICTL_AUTH="$(echo -n : | base64)" ``` ### Add a language:[​](#add-a-language "Direct link to Add a language:") ``` sudo CRICTL_AUTH="$(echo -n : | base64)" SM_DOCKER_PUBLIC=speechmaticspublic.azurecr.io BUILD_MODE=batch configure_languages.sh [ ...] ``` ### Remove a language:[​](#remove-a-language "Direct link to Remove a language:") ``` sudo CRICTL_AUTH="$(echo -n : | base64)" SM_DOCKER_PUBLIC=speechmaticspublic.azurecr.io BUILD_MODE=batch configure_languages.sh -r [ ...] ``` Running the configure language script requires superuser access. After the script has completed, the new language will be available to use. Adding a transcription language with this script will also make it available for `auto` Language Identification. ## Reverting[​](#reverting "Direct link to Reverting") If something appears to have gone wrong after running the script, the system configuration can be restored by running ``` sudo kubectl apply -f /root/.siab_deployment/k8s/bjapi.yaml ``` To assist Speechmatics Support, running this command will show the currently configured languages: ``` kubectl get -o json deployment/api | \ jq -r '.spec.template.spec.containers[0].env[]|select(.name="SM_FEATURES_VALIDATION").value' ``` This command will print a configuration file, under `batch->transcription->languages` there should be a section like: ``` { ... "batch": { "transcription": [ { "version": "latest", "languages": [ "en" ], ... } ] } ... } ``` --- # Language Identification Configure Language Identification on a Virtual Appliance deployment By default, Language Identification is disabled, and is currently only available in `Batch` mode. You can enable it using the Management API with ``` curl -sSL -u admin:admin -X 'POST' \ "http://${APPLIANCE_HOST}/v2/management/host/langid" \ -H 'Content-Type: application/json' \ -d '{"max_thread_count": 1, "memory_per_thread": "3Gi"}' ``` * `max_thread_count` is the maximum number of threads the Language Identification will use. When set to 0, Language Identification is disabled. Each Language Identification task will use up to one thread, so to allow 2 parallel identification tasks, you should set this value to 2. * `memory_per_thread` is the memory reserved for each configured thread. If you set `max_thread_count` to 2 and `memory_per_thread` to 3Gi, 6Gi will be reserved for the Language Identification component. `memory_per_thread` is optional and if not set will default to 3Gi. When submitting large audio files, you may need to increase the memory assigned to each thread. You can learn more about Language Identification in the [SaaS documentation](/speech-to-text/batch/language-identification.md). --- # Logcli Help Usage manual for the \`logcli\` command-line tool. ``` usage: logcli [] [; ...] A command-line for loki. ``` ## Flags[​](#flags "Direct link to Flags") ``` --help Show context-sensitive help (also try --help-long and --help-man). --version Show application version. -q, --quiet Suppress query metadata --stats Show query statistics -o, --output=default Specify output mode [default, raw, jsonl]. raw suppresses log labels and timestamp. -z, --timezone=Local Specify the timezone to use when formatting output timestamps [Local, UTC] --cpuprofile="" Specify the location for writing a CPU profile. --memprofile="" Specify the location for writing a memory profile. --stdin Take input logs from stdin --addr="http://localhost:3100" Server address. Can also be set using LOKI_ADDR env var. --username="" Username for HTTP basic auth. Can also be set using LOKI_USERNAME env var. --password="" Password for HTTP basic auth. Can also be set using LOKI_PASSWORD env var. --ca-cert="" Path to the server Certificate Authority. Can also be set using LOKI_CA_CERT_PATH env var. --tls-skip-verify Server certificate TLS skip verify. Can also be set using LOKI_TLS_SKIP_VERIFY env var. --cert="" Path to the client certificate. Can also be set using LOKI_CLIENT_CERT_PATH env var. --key="" Path to the client certificate key. Can also be set using LOKI_CLIENT_KEY_PATH env var. --org-id="" adds X-Scope-OrgID to API requests for representing tenant ID. Useful for requesting tenant data when bypassing an auth gateway. Can also be set using LOKI_ORG_ID env var. --query-tags="" adds X-Query-Tags http header to API requests. This header value will be part of `metrics.go` statistics. Useful for tracking the query. Can also be set using LOKI_QUERY_TAGS env var. --bearer-token="" adds the Authorization header to API requests for authentication purposes. Can also be set using LOKI_BEARER_TOKEN env var. --bearer-token-file="" adds the Authorization header to API requests for authentication purposes. Can also be set using LOKI_BEARER_TOKEN_FILE env var. --retries=0 How many times to retry each query when getting an error response from Loki. Can also be set using LOKI_CLIENT_RETRIES env var. --min-backoff=0 Minimum backoff time between retries. Can also be set using LOKI_CLIENT_MIN_BACKOFF env var. --max-backoff=0 Maximum backoff time between retries. Can also be set using LOKI_CLIENT_MAX_BACKOFF env var. --auth-header="Authorization" The authorization header used. Can also be set using LOKI_AUTH_HEADER env var. --proxy-url="" The http or https proxy to use when making requests. Can also be set using LOKI_HTTP_PROXY_URL env var. ``` ## Commands[​](#commands "Direct link to Commands") ### Help \[\...][​](#help-command "Direct link to Help \[...]") Show help. ### Query \[\] \[​](#query-flags-query "Direct link to Query \[] ") Run a LogQL query. The "query" command is useful for querying for logs. Logs can be returned in a few output modes: ``` raw: log line default: log timestamp + log labels + log line jsonl: JSON response from Loki API of log line ``` The output of the log can be specified with the "-o" flag, for example, "-o raw" for the raw output format. The "query" command will output extra information about the query and its results, such as the API URL, set of common labels, and set of excluded labels. This extra information can be suppressed with the --quiet flag. By default, we look over the last hour of data; use --since to modify or provide specific start and end times with --from and --to respectively. Notice that when using --from and --to then ensure to use RFC3339Nano time format, but without timezone at the end. The local timezone will be added automatically or if using --timezone flag. Example: ``` logcli query --timezone=UTC --from="2021-01-19T10:00:00Z" --to="2021-01-19T20:00:00Z" --output=jsonl 'my-query' ``` The output is limited to 30 entries by default; use --limit to increase. While "query" does support metrics queries, its output contains multiple data points between the start and end query time. This output is used to build graphs, similar to what is seen in the Grafana Explore graph view. If you are querying metrics and just want the most recent data point (like what is seen in the Grafana Explore table view), then you should use the "instant-query" command instead. ### Parallelization[​](#parallelization "Direct link to Parallelization") You can download an unlimited number of logs in parallel, there are a few flags which control this behaviour: ``` --parallel-duration --parallel-max-workers --part-path-prefix --overwrite-completed-parts --merge-parts --keep-parts ``` Refer to the help for each flag for details about what each of them do. Example: ``` logcli query --timezone=UTC --from="2021-01-19T10:00:00Z" --to="2021-01-19T20:00:00Z" --output=jsonl --parallel-duration="15m" --parallel-max-workers="4" --part-path-prefix="/tmp/my_query" --merge-parts 'my-query' ``` This example will create a queue of jobs to execute, each being 15 minutes in duration. In this case, that means, for the 10-hour total duration, there will be forty 15-minute jobs. The --limit flag is ignored. It will start four workers, and they will each take a job to work on from the queue until all the jobs have been completed. Each job will save a "part" file to the location specified by the --part-path-prefix. Different prefixes can be used to run multiple queries at the same time. The timestamp of the start and end of the part is in the file name. While the part is being downloaded, the filename will end in ".part", when it is complete, the file will be renamed to remove this ".part" extension. By default, if a completed part file is found, that part will not be downloaded again. This can be overridden with the --overwrite-completed-parts flag. Part file example using the previous command, adding --keep-parts so they are not deleted: Since we don't have the --forward flag, the parts will be downloaded in reverse. Two of the workers have finished their jobs (last two files), and have picked up the next jobs in the queue. Running ls, this is what we should expect to see. ``` ls -1 /tmp/my_query* /tmp/my_query_20210119T183000_20210119T184500.part.tmp /tmp/my_query_20210119T184500_20210119T190000.part.tmp /tmp/my_query_20210119T190000_20210119T191500.part.tmp /tmp/my_query_20210119T191500_20210119T193000.part.tmp /tmp/my_query_20210119T193000_20210119T194500.part /tmp/my_query_20210119T194500_20210119T200000.part ``` If you do not specify the `--merge-parts` flag, the part files will be downloaded, and logcli will exit, and you can process the files as you wish. With the flag specified, the part files will be read in order, and the output printed to the terminal. The lines will be printed as soon as the next part is complete, you don't have to wait for all the parts to download before getting output. The `--merge-parts` flag will remove the part files when it is done reading each of them. To change this, you can use the `--keep-parts` flag, and the part files will not be removed. ### Instant-Query \[\] \[​](#instant-query-flags-query "Direct link to Instant-Query \[] ") Run an instant LogQL query. The "instant-query" command is useful for evaluating a metric query for a single point in time. This is equivalent to the Grafana Explore table view; if you want a metrics query that is used to build a Grafana graph, you should use the "query" command instead. This command does not produce useful output when querying for log lines; you should always use the "query" command when you are running log queries. For more information about log queries and metric queries, refer to the LogQL documentation: ### Labels \[\] \[\