# GPU Speech to text container (Standard and Enhanced)

Learn about the Speechmatics Transcription GPU container system

## Prerequisites[​](#prerequisites "Direct link to Prerequisites")

* [A license file or a license token](/deployments/container/licensing.md)
  * There is no specific license for the GPU Inference Container, it will run using an existing Speechmatics license for the Realtime or Batch Container
* [Access to our Docker repository](/deployments/container/accessing-images.md)

## System requirements[​](#system-requirements "Direct link to System requirements")

The system must have:

* Nvidia GPU(s) with at least 16GB of GPU memory
* Nvidia drivers (see below for supported versions)
* CUDA [compute capability](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#compute-capabilities) of 7.5-12.1 inclusive, which corresponds to the Turing, Ampere, Lovelace, Hopper, Blackwell architectures. Cards with the Volta architecture or below are not able to run the models
* 24 GB RAM
* The [nvidia-container-toolkit](https://github.com/NVIDIA/nvidia-docker) installed
* Docker version > 19.03

The raw image size of the GPU Inference Container is around 15GB.

See [Performance and cost](/deployments/container/performance-and-cost.md) for more information on the performance and cost of the container.

### Nvidia drivers[​](#nvidia-drivers "Direct link to Nvidia drivers")

* 15.0.0 container version or higher: The GPU Inference Container is based on CUDA 13.0.1, which requires NVIDIA Driver release 580 or later.
* 14.13.0 container version or lower: The GPU Inference Container is based on CUDA 12.4.1, which requires NVIDIA Driver release 525 or later.

Driver installation can be validated by running `nvidia-smi`. This command should return the Nvidia driver version and show additional information about the GPU(s).

#### Cloud instances[​](#cloud-instances "Direct link to Cloud instances")

The GPU node can be provisioned in the cloud.

## Running the image[​](#running-the-image "Direct link to Running the image")

Currently, each GPU Inference Container can only run on a single GPU. If a system has more than one GPU, the device must be specified using `CUDA_VISIBLE_DEVICES` or selecting the device using the `--gpus` argument. See [Nvidia/CUDA documentation for details](https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html#cuda-environment-variables).

```
docker run --rm -it \
  -v $PWD/license.json:/license.json \
  --gpus '"device=0"' \
  -e CUDA_VISIBLE_DEVICES \
  -p 8001:8001 \
  speechmaticspublic.azurecr.io/sm-gpu-inference-server-en:15.19.0
```

When the Container starts you should see output similar to this, indicating that the server has started and is ready to serve requests.

```
I1215 11:43:57.300390 1 server.cc:633]
+----------------------+---------+--------+
| Model                | Version | Status |
+----------------------+---------+--------+
| am_en_enhanced       | 1       | READY  |
| am_en_standard       | 1       | READY  |
| body_enhanced        | 1       | READY  |
| body_standard        | 1       | READY  |
| diar_enhanced        | 1       | READY  |
| diar_standard        | 1       | READY  |
| ensemble_en_enhanced | 1       | READY  |
| ensemble_en_standard | 1       | READY  |
| lm_en_enhanced       | 1       | READY  |
+----------------------+---------+--------+
...
I1215 11:43:57.375233 1 grpc_server.cc:4819] Started GRPCInferenceService at 0.0.0.0:8001
I1215 11:43:57.375473 1 http_server.cc:3477] Started HTTPService at 0.0.0.0:8000
I1215 11:43:57.417749 1 http_server.cc:184] Started Metrics Service at 0.0.0.0:8002
```

### Batch and Realtime inference[​](#batch-and-realtime-inference "Direct link to Batch and Realtime inference")

The Inference server can run in two modes: *batch*, for processing whole files and returning the transcript at the end, and *real-time* for processing audio streams. The default mode is batch. To configure the GPU server for real-time, set the environment variable `SM_BATCH_MODE=false` by passing it into the `docker run` command.

The modes correspond to the two types of client speech Container, which are distinguished by their name:

* **rt**-asr-transcriber-en:\<version>
* **batch**-asr-transcriber-en:\<version>

The server can only support one of these modes at once.

### Linking to a GPU inference container[​](#linking-to-a-gpu-inference-container "Direct link to Linking to a GPU inference container")

Once the GPU Server is running, follow the [Instructions for Linking a CPU Container](/deployments/container/cpu-speech-to-text.md#linking-to-a-gpu-inference-container).

### Running only one model[​](#running-only-one-model "Direct link to Running only one model")

[Models](/speech-to-text/models.md) (previously called Operating Points) represent different levels of model complexity. To save GPU memory for throughput, you can run the server with only one model loaded. To do this, pass the `SM_MODEL` environment variable to the container and set it to either `standard` or `enhanced`.

`SM_MODEL` replaces the older `SM_OPERATING_POINT` environment variable. `SM_OPERATING_POINT` is deprecated but still works and accepts the same `standard` and `enhanced` values; use `SM_MODEL` going forward.

When running the all language standard model GPU inference server you must set the `SM_MODEL` environment variable to `standard`

### Loading only selected languages[​](#loading-only-selected-languages "Direct link to Loading only selected languages")

Images that contain more than one language model load every language at startup. To save GPU memory for throughput, pass the `SM_LANGUAGES` environment variable to the container and set it to a comma-separated list of language codes. Only those languages are loaded:

```
# load only English, German, and French
-e SM_LANGUAGES=en,de,fr
```

Codes must be languages the image contains. Any code the image does not contain causes startup to fail with an error.

When `SM_LANGUAGES` is unset, all of the image's languages are loaded. `SM_LANGUAGES` is available from container version 15.18.0.

`SM_LANGUAGES` and `SM_MODEL` can be used together. For example, `SM_MODEL=standard` with `SM_LANGUAGES=en,de` loads the Standard model for English and German only.

### Monitoring the server[​](#monitoring-the-server "Direct link to Monitoring the server")

The inference server is based on [Nvidia's Triton architecture](https://developer.nvidia.com/nvidia-triton-inference-server) and as such can be monitored using Triton's inbuilt Prometheus metrics, or the GRPC/HTTP APIs. To expose these, configure an external mapping for port 8002(Prometheus) or 8000(HTTP).

### Models in GPU inference[​](#models-in-gpu-inference "Direct link to Models in GPU inference")

When inference is outsourced to a GPU server, alternative GPU-specific models are used, so you should not expect to see identical results compared to CPU-based inference. For convenience, the GPU models are also designated as 'standard' and 'enhanced'.

## Docker compose example[​](#docker-compose-example "Direct link to Docker compose example")

This Docker Compose file will create a Speechmatics GPU Inference Server:

(assumes your `license.json` file is in the current working directory)

docker-compose.yml

```
version: "3.8"

networks:
  transcriber:
    driver: bridge

services:
  triton:
    image: speechmaticspublic.azurecr.io/sm-gpu-inference-server-en:15.19.0
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              ### Limit to N GPUs
              # count: 1
              ### Pick specific GPUs by device ID
              # device_ids:
              #   - 0
              #   - 3
              capabilities:
                - gpu
    container_name: triton
    networks:
      - transcriber
    expose:
      - 8000/tcp
      - 8001/tcp
      - 8002/tcp
    environment:
      - NVIDIA_DRIVER_CAPABILITIES=all
      - NVIDIA_VISIBLE_DEVICES=all
      - CUDA_VISIBLE_DEVICES=0
    volumes:
      - $PWD/license.json:/license.json:ro
```
