# Output

Learn about the supported output formats for the Speechmatics Batch API

Transcription jobs are processed asynchronously by default: you submit a job, then check its status to see whether it has completed. To block for the result in a single request instead, use [Synchronous transcription](/speech-to-text/batch/synchronous.md).

You can also configure notifications to be sent to a webhook when a job is completed. See [Notifications](/speech-to-text/batch/notifications.md) for more details.

## Check single job status[​](#check-single-job-status "Direct link to Check single job status")

If you wish to retrieve a particular job, you can do so using the job ID for up to 7 days, after which time it will be automatically deleted in accordance with our [Data Retention Policy](/speech-to-text/batch/limits.md#data-retention-limits).

You can make a GET request to check the status of a job as follows:

UnixUnixWindowsWindows

```
# $JOB_ID is from the submit command output
curl -L -X GET "https://eu1.asr.api.speechmatics.com/v2/jobs/$JOB_ID" \
-H "Authorization: Bearer $API_KEY"
```

```
curl -L -X GET "https://eu1.asr.api.speechmatics.com/v2/jobs/" \
-H "Authorization: Bearer $API_KEY"
```

This endpoint applies a short default wait, returning as soon as the job reaches a terminal state. Pass the `wait` query parameter to control the duration, or `wait=0` to return immediately. See [Synchronous transcription](/speech-to-text/batch/synchronous.md#default-wait-on-the-get-endpoints).

The response is a JSON object containing details of the job, with the `status` field showing whether the job is still processing or not. The possible values are:

* **`running`**: The job is still processing.
* **`done`**: The job has been completed. The transcript is available for download at the [`/jobs/:jobid/transcript` endpoint](/api-ref/batch/get-the-transcript-for-a-transcription-job.md).
* **`rejected`**: Transcription was not possible. This will be accompanied by an error message.

An example response looks like this:

```
HTTP/1.1 200 OK
Content-Type: application/json
{
  "job": {
    "config": {
      "notification_config": null,
      "transcription_config": {
        "additional_vocab": null,
        "channel_diarization_labels": null,
        "language": "en"
      },
      "type": "transcription"
    },
    "created_at": "2019-01-17T17:50:54.113Z",
    "data_name": "example.wav",
    "duration": 275,
    "id": "yjbmf9kqub",
    "status": "running"
  }
}
```

## Check multiple job statuses[​](#check-multiple-job-statuses "Direct link to Check multiple job statuses")

You can retrieve the status of the 100 most recent jobs submitted in the past 7 days. This is done by making a GET request without a Job ID as follows:

UnixUnixWindowsWindows

```
curl -L -X GET "https://eu1.asr.api.speechmatics.com/v2/jobs/" \
     -H "Authorization: Bearer $API_KEY"
```

```
curl -L -X GET "https://eu1.asr.api.speechmatics.com/v2/jobs/" `
     -H "Authorization: Bearer $API_KEY"
```

A deleted job is not included in the list unless you request it explicitly, as described in the [API reference](/api-ref/batch/list-all-jobs.md).

## Load and process the transcript[​](#load-and-process-the-transcript "Direct link to Load and process the transcript")

The transcript endpoint returns plain text, JSON, or SRT. Request the format using the `format` query parameter. See the [API reference](/api-ref/batch/get-the-transcript-for-a-transcription-job.md#request).

A few useful things to know about transcript formats:

* The default format is JSON.
* Use the `format=txt` query parameter to get the transcript in plain text. Useful for quick access to the transcript.
* Use the `format=srt` query parameter to get the transcript in SRT format. Useful for displaying the transcript in a subtitle file.
* To access other data, including word timestamps, translations, and speech intelligence features, use the default JSON format.

Like the job status endpoint, this endpoint applies a short default wait. Pass the `wait` query parameter alongside `format` to block for the transcript instead of polling. See [Synchronous transcription](/speech-to-text/batch/synchronous.md#wait-for-the-transcript).

### Transcript response schema[​](#transcript-response-schema "Direct link to Transcript response schema")

Below is the schema for the transcript response when using the default JSON format.

Refer to the [API reference](/api-ref/batch/get-the-transcript-for-a-transcription-job.md#responses) for further details.

**format**stringrequired

Speechmatics JSON transcript format version number.

**Example:&#x20;**`2.1`

**job** <!-- -->objectrequired

Summary information about an ASR job, to support identification and tracking.

**created\_at**date-timerequired

The UTC date time the job was created.

**Example:&#x20;**`2018-01-09T12:29:01.853047Z`

**data\_name**stringrequired

Name of data file submitted for job.

**duration**integerrequired

The data file audio duration (in seconds).

Possible values: `>= 0`

**id**stringrequired

The unique id assigned to the job.

**Example:&#x20;**`a1b2c3d4e5`

**tracking** <!-- -->object

**title**string

The title of the job.

**reference**string

External system reference.

**tags**string\[]

**details**object

Customer-defined JSON structure.

**metadata** <!-- -->objectrequired

Summary information about the output from an ASR job, comprising the job type and configuration parameters used when generating the output.

**created\_at**date-timerequired

The UTC date time the transcription output was created.

**Example:&#x20;**`2018-01-09T12:29:01.853047Z`

**type**stringrequired

Possible values: \[`transcription`]

**transcription\_config** <!-- -->object

**language**stringrequired

Language model to process the audio input, normally specified as an ISO language code

**domain**string

Request a specialized model based on 'language' but optimized for a particular field, e.g. 'finance' or 'medical'.

**output\_locale**string

Language locale to be used when generating the transcription output, normally specified as an ISO language code

**model**string

Specific model to use in transcription (previously called operating point).

Possible values: \[`standard`, `enhanced`, `melia-1`]

**operating\_point**stringdeprecated

**Deprecated**: Use `model` instead. `operating_point` is left only for backward compatibility.

Possible values: \[`standard`, `enhanced`, `melia-1`]

**additional\_vocab** <!-- -->object\[]

List of custom words or phrases that should be recognized. Alternative pronunciations can be specified to aid recognition.

**Possible values:** `<= 20000`

* Array \[

**content**stringrequired

**sounds\_like**string\[]

* ]

**punctuation\_overrides** <!-- -->object

Control punctuation settings.

**sensitivity**float

Ranges between zero and one. Higher values will produce more punctuation. The default is 0.5.

Possible values: `>= 0` and `<= 1`

**permitted\_marks**string\[]

The punctuation marks which the client is prepared to accept in transcription output, or the special value 'all' (the default). Unsupported marks are ignored. This value is used to guide the transcription process.

Possible values: Value must match regular expression `^(.|all)$`

**diarization**string

Specify whether speaker or channel labels are added to the transcript. The default is `none`.

* **none**: no speaker or channel labels are added.
* **speaker**: speaker attribution is performed based on acoustic matching; all input channels are mixed into a single stream for processing.
* **channel**: multiple input channels are processed individually and collated into a single transcript.

Possible values: \[`none`, `speaker`, `channel`]

**channel\_diarization\_labels**string\[]

Transcript labels to use when using collating separate input channels.

Possible values: Value must match regular expression `^[A-Za-z0-9._]+$`

**enable\_entities**boolean

Include additional 'entity' objects in the transcription results (e.g. dates, numbers) and their original spoken form. These entities are interleaved with other types of results. The concatenation of these words is represented as a single entity with the concatenated written form present in the 'content' field. The entities contain a 'spoken\_form' field, which can be used in place of the corresponding 'word' type results, in case a spoken form is preferred to a written form. They also contain a 'written\_form', which can be used instead of the entity, if you want a breakdown of the words without spaces. They can still contain non-breaking spaces and other special whitespace characters, as they are considered part of the word for the formatting output. In case of a written\_form, the individual word times are estimated and might not be accurate if the order of the words in the written form does not correspond to the order they were actually spoken (such as 'one hundred million dollars' and '$100 million').

**max\_delay\_mode**string

Whether or not to enable flexible endpointing and allow the entity to continue to be spoken.

Possible values: \[`fixed`, `flexible`]

**audio\_filtering\_config** <!-- -->object

Configuration for limiting the transcription of quiet audio.

**volume\_threshold**float

Controls the lower limit of audio volume at which speech and audio events will be transcribed. If the volume limit is very low, then most sound will be passed to the speech recognition engine. Higher numbers will cut out increasing amounts of sound.

Possible values: `>= 0` and `<= 100`

**transcript\_filtering\_config** <!-- -->object

Configuration for applying filtering to the transcription

**remove\_disfluencies**boolean

If true, words that are identified as disfluencies will be removed from the transcript. If false (default), they are tagged in the transcript as 'disfluency'.

**replacements** <!-- -->object\[]

An array of objects defining custom replacements. Each replacement contains a pair of strings: the text to find `from` and the text to replace it with `to`.

* Array \[

**from**stringrequired

**to**stringrequired

* ]

**speaker\_diarization\_config** <!-- -->object

Configuration for speaker diarization

**prefer\_current\_speaker**boolean

If true, the algorithm will prefer to stay with the current active speaker if it is a close enough match, even if other speakers may be closer. This is useful for cases where we can flip incorrectly between similar speakers during a single speaker section.

**speaker\_sensitivity**float

Controls how sensitive the algorithm is in terms of keeping similar speakers separate, as opposed to combining them into a single speaker. Higher values will typically lead to more speakers, as the degree of difference between speakers in order to allow them to remain distinct will be lower. A lower value for this parameter will conversely guide the algorithm towards being less sensitive in terms of retaining similar speakers, and as such may lead to fewer speakers overall. The default is 0.5.

Possible values: `>= 0` and `<= 1`

**get\_speakers**boolean

If true, speaker identifiers will be returned at the end of transcript.

**speakers** <!-- -->object\[]

Use this option to provide speaker labels linked to their speaker identifiers. When passed, the transcription system will tag spoken words in the transcript with the provided speaker labels whenever any of the specified speakers is detected in the audio. A maximum of 50 speakers identifiers across all speakers can be provided.

* Array \[

**label**stringrequired

Speaker label, which must not match the format used internally (e.g. S1, S2, etc)

Possible values: `non-empty`

**speaker\_identifiers**bytes\[]required

Possible values: `>= 1`

* ]

**translation\_errors** <!-- -->object\[]

List of errors that occurred in the translation stage.

* Array \[

**type**string

Possible values: \[`translation_failed`, `unsupported_translation_pair`]

**message**string

Human readable error message

* ]

**summarization\_errors** <!-- -->object\[]

List of errors that occurred in the summarization stage.

* Array \[

**type**string

Possible values: \[`summarization_failed`, `unsupported_language`]

**message**string

Human readable error message

* ]

**sentiment\_analysis\_errors** <!-- -->object\[]

List of errors that occurred in the sentiment analysis stage.

* Array \[

**type**string

Possible values: \[`sentiment_analysis_failed`, `unsupported_language`]

**message**string

Human readable error message

* ]

**topic\_detection\_errors** <!-- -->object\[]

List of errors that occurred in the topic detection stage.

* Array \[

**type**string

Possible values: \[`topic_detection_failed`, `unsupported_list_of_topics`, `unsupported_language`]

**message**string

Human readable error message

* ]

**auto\_chapters\_errors** <!-- -->object\[]

List of errors that occurred in the auto chapters stage.

* Array \[

**type**string

Possible values: \[`auto_chapters_failed`, `unsupported_language`]

**message**string

Human readable error message

* ]

**output\_config** <!-- -->object

**srt\_overrides** <!-- -->object

Parameters that override default values of srt conversion. max\_line\_length: sets maximum count of characters per subtitle line including white space. max\_lines: sets maximum count of lines in a subtitle section.

**max\_line\_length**integer

**max\_lines**integer

**language\_pack\_info** <!-- -->object

Properties of the language pack.

**language\_description**string

Full descriptive name of the language, e.g. 'Japanese'.

**word\_delimiter**stringrequired

The character to use to separate words.

**writing\_direction**string

The direction that words in the language should be written and read in.

Possible values: \[`left-to-right`, `right-to-left`]

**itn**boolean

Whether or not ITN (inverse text normalization) is available for the language pack.

**adapted**boolean

Whether or not language model adaptation has been applied to the language pack.

**language\_identification** <!-- -->object

Result of the language identification of the audio, configured using `language_identification_config`, or setting the transcription language to `auto`.

**results** <!-- -->object\[]

* Array \[

**alternatives** <!-- -->object\[]

* Array \[

**language**string

**confidence**number

* ]

**start\_time**number

**end\_time**number

* ]

**error**string

Possible values: \[`LOW_CONFIDENCE`, `UNEXPECTED_LANGUAGE`, `NO_SPEECH`, `FILE_UNREADABLE`, `OTHER`]

**message**string

**orchestrator\_version**string

Orchestrator version in PEP 440 Format or set to 'version\_not\_found' as default.

**results** <!-- -->object\[]required

* Array \[

**channel**string

**start\_time**floatrequired

**end\_time**floatrequired

**volume**float

An indication of the volume of audio across the time period the word was spoken.

Possible values: `>= 0` and `<= 100`

**is\_eos**boolean

Whether the punctuation mark is an end of sentence character. Only applies to punctuation marks.

**type**stringrequired

New types of items may appear without being requested; unrecognized item types can be ignored.

Possible values: \[`word`, `punctuation`, `entity`]

**written\_form** <!-- -->object\[]

* Array \[

**alternatives** <!-- -->object\[]required

* Array \[

**content**stringrequired

**confidence**floatrequired

**language**stringrequired

**display** <!-- -->object

**direction**stringrequired

Possible values: \[`ltr`, `rtl`]

**speaker**string

**tags**string\[]

* ]

**end\_time**floatrequired

**start\_time**floatrequired

**type**stringrequired

What kind of object this is. See #/Definitions/RecognitionResult for definitions of the enums.

Possible values: \[`word`]

* ]

**spoken\_form** <!-- -->object\[]

* Array \[

**alternatives** <!-- -->object\[]required

* Array \[

**content**stringrequired

**confidence**floatrequired

**language**stringrequired

**display** <!-- -->object

**direction**stringrequired

Possible values: \[`ltr`, `rtl`]

**speaker**string

**tags**string\[]

* ]

**end\_time**floatrequired

**start\_time**floatrequired

**type**stringrequired

What kind of object this is. See #/Definitions/RecognitionResult for definitions of the enums.

Possible values: \[`word`, `punctuation`]

* ]

**alternatives** <!-- -->object\[]

* Array \[

**content**stringrequired

**confidence**floatrequired

**language**stringrequired

**display** <!-- -->object

**direction**stringrequired

Possible values: \[`ltr`, `rtl`]

**speaker**string

**tags**string\[]

* ]

**attaches\_to**string

Attachment direction of the punctuation mark. This only applies to punctuation marks. This information can be used to produce a well-formed text representation by placing the `word_delimiter` from `language_pack_info` on the correct side of the punctuation mark.

Possible values: \[`previous`, `next`, `both`, `none`]

* ]

**speakers** <!-- -->object\[]

List of unique speaker identifiers detected in the transcript.

* Array \[

**label**stringrequired

Speaker label.

Possible values: `non-empty`

**speaker\_identifiers**bytes\[]required

Possible values: `>= 1`

* ]

**translations** <!-- -->object

Translations of the transcript into other languages. It is a map of ISO language codes to arrays of translated sentences. Configured using `translation_config`.

**\[property name: string]** <!-- -->object\[]

* Array \[

**start\_time**float

**end\_time**float

**content**string

**speaker**string

**channel**string

* ]

**summary** <!-- -->object

Summary of the transcript, configured using `summarization_config`.

**content**string

**sentiment\_analysis** <!-- -->object

The main object that holds sentiment analysis data.

**sentiment\_analysis** <!-- -->object

Holds the detailed sentiment analysis information.

**segments** <!-- -->object\[]

An array of objects that represent a segment of text and its associated sentiment.

* Array \[

**text**string

Represents the transcript of the analysed segment

**sentiment**string

The assigned sentiment to the segment, which can be positive, neutral or negative

**start\_time**float

The timestamp corresponding to the beginning of the transcription segment

**end\_time**float

The timestamp corresponding to the end of the transcription segment

**speaker**string

The speaker label for the segment, if speaker diarization is enabled

**channel**string

The channel label for the segment, if channel diarization is enabled

**confidence**float,

A confidence score in the range of 0-1, indicating the model's certainty in the predicted sentiment,

* ]

**summary** <!-- -->object

An object that holds overall sentiment information, and per-speaker and per-channel sentiment data.

**overall** <!-- -->object

Summary of overall sentiment data.

**positive\_count**integer

**negative\_count**integer

**neutral\_count**integer

**speakers** <!-- -->object\[]

An array of objects that represent sentiment data for a specific speaker.

* Array \[

**speaker**string

**positive\_count**integer

**negative\_count**integer

**neutral\_count**integer

* ]

**channels** <!-- -->object\[]

An array of objects that represent sentiment data for a specific channel.

* Array \[

**channel**string

**positive\_count**integer

**negative\_count**integer

**neutral\_count**integer

* ]

**topics** <!-- -->object

Main object that holds topic detection results.

**segments** <!-- -->object\[]

An array of objects that represent a segment of text and its associated topic information.

* Array \[

**text**string

**start\_time**float

**end\_time**float

**topics** <!-- -->object\[]

* Array \[

**topic**string

* ]

* ]

**summary** <!-- -->object

An object that holds overall information on the topics detected.

**overall** <!-- -->object

Summary of overall topic detection results.

**\[property name: string]**&#x69;nteger

**chapters** <!-- -->object\[]

An array of objects that represent summarized chapters of the transcript

* Array \[

**title**string

The auto-generated title for the chapter

**summary**string

An auto-generated paragraph-style, short summary of the chapter

**start\_time**number

The start time of the chapter in the audio file

**end\_time**number

* ]

**audio\_events** <!-- -->object\[]

Timestamped audio events, only set if `audio_events_config` is in the config

* Array \[

**type**string

Kind of audio event. E.g. music

**start\_time**float

Time (in seconds) at which the audio event starts

**end\_time**float

Time (in seconds) at which the audio event ends

**confidence**float

Prediction confidence associated with this event

**channel**string

Input channel this event occurred on

* ]

**audio\_event\_summary** <!-- -->object

Summary statistics per event type, keyed by `type`, e.g. music

**overall** <!-- -->object

Overall summary on all channels

**\[property name: string]** <!-- -->object

Summary statistics for this audio event type

**total\_duration**float

Total duration (in seconds) of all audio events of this type

**count**number

Number of events of this type

**channels** <!-- -->object

Summary keyed by channel, only set if channel diarization is enabled

**\[property name: string]** <!-- -->object

**\[property name: string]** <!-- -->object

Summary statistics for this audio event type

**total\_duration**float

Total duration (in seconds) of all audio events of this type

**count**number

Number of events of this type

### Example response[​](#example-response "Direct link to Example response")

The following is an example of a transcript response, which you should see as an output of the provided [example.wav](https://github.com/speechmatics/speechmatics-js-sdk/raw/7d219bfee9166736e6aa21598535a194387b84be/examples/nodejs/example.wav) file used in the [code samples in the quickstart](/speech-to-text/batch/quickstart.md).

```
{
  "format": "2.9",
  "job": {
    "created_at": "2025-06-30T11:43:54.135Z",
    "data_name": "example.wav",
    "duration": 15,
    "id": "650krlru2e"
  },
  "metadata": {
    "created_at": "2025-06-30T11:44:08.133526Z",
    "language_pack_info": {
      "adapted": false,
      "itn": true,
      "language_description": "English",
      "word_delimiter": " ",
      "writing_direction": "left-to-right"
    },
    "orchestrator_version": "2025.06.28+1eb4127132+13.4.0",
    "transcription_config": {
      "language": "en",
      "model": "enhanced"
    },
    "type": "transcription"
  },
  "results": [
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "Welcome",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 1.36,
      "start_time": 0.72,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "to",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 1.44,
      "start_time": 1.36,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 0.98,
          "content": "Speechmatics",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 2.44,
      "start_time": 1.48,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": ".",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "attaches_to": "previous",
      "end_time": 2.44,
      "is_eos": true,
      "start_time": 2.44,
      "type": "punctuation"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "We're",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 3.24,
      "start_time": 3.04,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "delighted",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 3.64,
      "start_time": 3.24,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "that",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 3.8,
      "start_time": 3.64,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "you've",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 4.04,
      "start_time": 3.8,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "decided",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 4.44,
      "start_time": 4.04,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "to",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 4.56,
      "start_time": 4.44,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "try",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 4.92,
      "start_time": 4.56,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "our",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 5.2,
      "start_time": 4.92,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "speech",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 5.48,
      "start_time": 5.2,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "to",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 5.6,
      "start_time": 5.48,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "text",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 5.92,
      "start_time": 5.6,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "software",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 6.56,
      "start_time": 5.92,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "to",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 7.16,
      "start_time": 7.04,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "get",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 7.4,
      "start_time": 7.2,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "going",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 7.88,
      "start_time": 7.4,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": ".",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "attaches_to": "previous",
      "end_time": 7.88,
      "is_eos": true,
      "start_time": 7.88,
      "type": "punctuation"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "Just",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 8.12,
      "start_time": 7.92,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "create",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 8.44,
      "start_time": 8.12,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "an",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 8.64,
      "start_time": 8.48,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "API",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 9,
      "start_time": 8.64,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "key",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 9.36,
      "start_time": 9,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "and",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 9.6,
      "start_time": 9.36,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "submit",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 9.88,
      "start_time": 9.6,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "a",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 10.04,
      "start_time": 9.92,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "transcription",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 10.6,
      "start_time": 10.04,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "request",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 11.04,
      "start_time": 10.6,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "to",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 11.12,
      "start_time": 11.04,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "our",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 11.32,
      "start_time": 11.16,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "API",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 11.88,
      "start_time": 11.32,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": ".",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "attaches_to": "previous",
      "end_time": 11.88,
      "is_eos": true,
      "start_time": 11.88,
      "type": "punctuation"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "We",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 12.36,
      "start_time": 12.24,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "hope",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 12.56,
      "start_time": 12.4,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "you'll",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 12.8,
      "start_time": 12.56,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "be",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 12.88,
      "start_time": 12.8,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "very",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 13.08,
      "start_time": 12.88,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "impressed",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 13.4,
      "start_time": 13.08,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "by",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 13.56,
      "start_time": 13.4,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "the",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 13.68,
      "start_time": 13.56,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "results",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 14.36,
      "start_time": 13.68,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": ".",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "attaches_to": "previous",
      "end_time": 14.36,
      "is_eos": true,
      "start_time": 14.36,
      "type": "punctuation"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "Thank",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 15.08,
      "start_time": 14.8,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": "you",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "end_time": 15.4,
      "start_time": 15.08,
      "type": "word"
    },
    {
      "alternatives": [
        {
          "confidence": 1,
          "content": ".",
          "language": "en",
          "speaker": "UU"
        }
      ],
      "attaches_to": "previous",
      "end_time": 15.4,
      "is_eos": true,
      "start_time": 15.4,
      "type": "punctuation"
    }
  ]
}
```

### Multilingual transcript output[​](#multilingual-transcript-output "Direct link to Multilingual transcript output")

For a Melia 1 job, the `language` property on each word reflects the language detected for that word, so it can change across the transcript. For Enhanced and Standard jobs, which transcribe one selected language, the same language is reported for every word.

The example below shows two words in different languages within one transcript:

```
{
  "results": [
    {
      "alternatives": [
        { "content": "Hello", "confidence": 0.98, "language": "en" }
      ],
      "start_time": 0.20,
      "end_time": 0.52,
      "type": "word"
    },
    {
      "alternatives": [
        { "content": "مرحبا", "confidence": 0.95, "language": "ar" }
      ],
      "start_time": 0.60,
      "end_time": 1.04,
      "type": "word"
    }
  ]
}
```

For multilingual transcripts, `language_pack_info` reports the word delimiter and writing direction per language rather than for a single language pack:

```
{
  "metadata": {
    "language_pack_info": {
      "per_language_word_delimiters": {
        "en": " ",
        "ar": " "
      },
      "per_language_writing_direction": {
        "en": "left-to-right",
        "ar": "right-to-left"
      }
    }
  }
}
```

`per_language_word_delimiters` gives the word delimiter for each language in the transcript, and `per_language_writing_direction` gives its writing direction.

## Tracking metadata[​](#tracking-metadata "Direct link to Tracking metadata")

You can optionally add tracking metadata to a job when you create it. This can be used for tracking the job through your own management workflow.

The metadata will be returned when you [check job status](#check-single-job-status), or [get the transcript](/api-ref/batch/get-the-transcript-for-a-transcription-job.md) for a transcription job.

The metadata object can contain the following properties:

### `tracking`[​](#tracking "Direct link to tracking")

**title**string

The title of the job.

**reference**string

External system reference.

**tags**string\[]

**details**object

Customer-defined JSON structure.

For example:

```
{
  "type": "transcription",
  "transcription_config": {
    "model": "enhanced",
    "language": "en"
  },
  "tracking": {
    "title": "ACME Q12018 Statement",
    "reference": "/data/clients/ACME/statements/segs/2018Q1-seg8",
    "tags": [
      "quick-review",
      "segment"
    ],
    "details": {
      "user_type": "agent",
      "client": {
        "name": "ACME Corp",
        "type": "enterprise"
        },
      "seg_start": 963.201,
      "seg_end": 1091.481
    }
  }
}
```

## App usage tracking[​](#app-usage-tracking "Direct link to App usage tracking")

First, please contact [Support](https://support.speechmatics.com) to enable this feature.

For integrations where customers use their own Speechmatics API keys, gain insights into application usage by aggregating data on unique users, processing hours, languages used, and more. Once enabled, use the `sm-app` query parameter when starting a job. For example:

```
APP_ID="YourAppID"
API_KEY="YOUR_API_KEY"
PATH_TO_FILE="example.wav"

curl -L -X POST "https://eu1.asr.api.speechmatics.com/v2/jobs/?sm-app=${APP_ID}" \
    -H "Authorization: Bearer ${API_KEY}" \
    -F data_file=@${PATH_TO_FILE} \
    -F config='{"type": "transcription","transcription_config": { "model": "enhanced","language": "en" }}'
```

## Next steps[​](#next-steps "Direct link to Next steps")

#### [Synchronous transcription](/speech-to-text/batch/synchronous.md)

[Get a transcript in one call instead of polling](/speech-to-text/batch/synchronous.md)

#### [Troubleshooting](/speech-to-text/batch/troubleshooting.md)

[Resolve common issues with the Batch API](/speech-to-text/batch/troubleshooting.md)

#### [API Reference](/api-ref/batch/create-a-new-job.md)

[Full details about the Batch API](/api-ref/batch/create-a-new-job.md)

#### [Inputs](/speech-to-text/batch/input.md)

[See which file types are supported](/speech-to-text/batch/input.md)

#### [Limits](/speech-to-text/batch/limits.md)

[See the limits for the Batch API](/speech-to-text/batch/limits.md)
