Start a Real-time STT task

Updated

Start speech-to-text transcription.

Real-Time STT API versions v5.x and v6.x are deprecated since June 2025 and will reach end-of-life on June 11, 2026. Use the v7 REST API for new projects.

Endpoint

  • Method: POST

  • Endpoint: https://api.agora.io/v1/projects/{appId}/rtsc/speech-to-text/tasks

After you acquire a builderToken, call this method within 5 minutes to start speech-to-text conversion.

Request

Path parameters

appId Type: string Required

The App ID of the project

Query parameters

builderToken Type: string Required

The tokenName value you obtained in the response body of the acquire method. To stop a task, use the same builderToken you used to start the task.

Request body

APPLICATION/JSON

BODY

languages Type: array[string] Required

The transcription languages to recognize. You can specify a maximum of 2 languages. Refer to Supported Languages for details.

maxIdleTime Type: integer Optional Default: 30 Values: 5 to 2592000

Maximum channel idle time, in seconds. When the specified time is exceeded, the transcription task ends automatically.

rtcConfig Type: object Required

channelName Type: string Required

The name of the channel to transcribe.

subBotUid Type: string Required

The ID of the bot that subscribes to the audio stream. All UIDs within a channel must be unique. Ensure no other user or service bot is using this UID in the same channel.

subBotToken Type: string Optional

The token used by the subscribing bot for channel authentication. Required only when your project has App Certificate enabled. Generate this token on your token server. For details, see Token authentication.

pubBotUid Type: string Required

The ID of the bot that pushes subtitle information to the channel. All UIDs within a channel must be unique. Ensure no other user or service bot is using this UID in the same channel.

pubBotToken Type: string Optional

The token used by the subtitle-pushing bot for channel authentication. Required only when your project has App Certificate enabled. Generate this token on your token server. For details, see Token authentication.

subscribeAudioUids Type: array[string] Optional

The user IDs of the audio streams you want to subscribe to. Specify this parameter only if you need to subscribe to specific users. Maximum array length: 3.

cryptionMode Type: integer Optional Values: 0 to 8

The encryption and decryption mode. When enabled, this mode is used for both decrypting incoming streams and encrypting outgoing subtitles.

  • 0: No encryption
  • 1: AES_128_XTS 128-bit AES encryption, XTS mode
  • 2: AES_128_ECB 128-bit AES encryption, ECB mode
  • 3: AES_256_XTS 256-bit AES encryption, XTS mode
  • 4: SM4_128_ECB 128-bit SM4 encryption, ECB mode
  • 5: AES_128_GCM 128-bit AES encryption, GCM mode
  • 6: AES_256_GCM 256-bit AES encryption, GCM mode
  • 7: AES_128_GCM2 128-bit AES encryption, GCM mode, requires key and salt
  • 8: AES_256_GCM2 256-bit AES encryption, GCM mode, requires key and salt The decryption method must match the encryption method set for the channel.

secret Type: string Optional

The encryption/decryption key. Required when cryptionMode is not 0.

salt Type: string Optional

A Base64-encoded, 32-byte encryption/decryption salt. Required only when cryptionMode is 7 or 8.

enableJsonProtocol Type: boolean Optional Default: false

Set the encoding format of the subtitle data pushed to the channel.

  • true: Use JSON to push subtitles and compress data with gzip. Uses less bandwidth, but requires decoding.
  • false: Use Protobuf to push subtitles (default). The data volume is smaller. Suitable for scenarios with high transmission efficiency requirements.

transferMode Type: array[string] Optional

Select the channel to push, RTM or RTC

captionConfig Type: object Optional

sliceDuration Type: integer Optional Default: 60 Values: 5 to 28800

The slice size of the recorded subtitle file, in seconds.

storage Type: object Required

accesskey Type: string Required

The access key of the third-party cloud storage.

secretkey Type: string Required

The secret key of the third-party cloud storage.

bucket Type: string Required

The bucket name of the third-party cloud storage.

vendor Type: integer Required Values: 1,5,6

The third-party cloud storage platform:

  • 1: Amazon S3
  • 5: Microsoft Azure
  • 6: Google Cloud

region Type: integer Required

The region information for the third-party cloud storage. To ensure successful and real-time uploading of recorded files, the cloud storage region must match the region of the application server where you initiate the request. For example, if your App server is in East US, set the cloud storage region to East US as well. See third-party storage regions for details.

fileNamePrefix Type: array[string] Optional

The storage location of the recorded file in the third-party cloud storage. The prefix length (including slashes) must not exceed 128 characters. The following characters are supported:

  • Lowercase English letters (a-z)
  • Uppercase English letters (A-Z)
  • Numbers (0-9)
    Symbols like slashes, underscores, and brackets must not appear in the string.

translateConfig Type: object Optional

forceTranslateInterval Type: integer Optional Default: 3 Values: 2 to 5

The time interval for forced translation, in seconds.

languages Type: array Required

The translation language array. You can specify a maximum of 2 different source languages. The source language and target language must be different, otherwise an error is reported.
Each array item is an object with: source Type: string Required

The source language for translation. Refer to Supported Languages for details.

target Type: array[string] Required

The target languages for translation. You can specify a maximum of 5 target languages for each source language. Refer to Supported Languages for details.

Response

  • If the returned status code is 200, the request was successful. The response body contains the result of the request.

    OK

taskId Type: string Required

The unique identifier of this transcription task.

createTs Type: integer Required

The Unix timestamp (in seconds) when the transcription task was created.

status Type: string Required

The current status of the transcription task:

  • IDLE: Task not initialized

  • PREPARING: Task has received an initialization request

  • PREPARED: Task initialization completed

  • STARTING: Task is beginning to start

  • CREATED: Task startup partially completed

  • STARTED: Task startup fully completed

  • IN_PROGRESS: Task is currently running

  • STOPPING: Task is in the process of being paused

  • STOPPED: Task has been terminated

  • FAILURE_STOP: Task termination failed

  • If the returned status code is not 200, the request failed. Refer to the message field to understand the possible reasons for failure.

Non-200

message Type: string

The reason why the request failed.

Authorization

This endpoint requires Basic Auth.

Request example

      curl --request POST \
        --url 'https://api.agora.io/v1/projects/:appId/rtsc/speech-to-text/tasks?builderToken=your_builder_token' \
        --header 'Authorization: Basic <credentials>'   

Response example

200

    {
      "taskId": "The unique identifier of this transcription task",
      "createTs": null,
      "status": "The status of the task"
    }

Non-200

    {
      "message": "The reason why the request failed.",
    }