Start a Real-time STT task
Updated
Start speech-to-text transcription.
Real-Time STT API versions v5.x and v6.x are deprecated since June 2025 and will reach end-of-life on June 11, 2026. Use the v7 REST API for new projects.
Endpoint
-
Method:
POST -
Endpoint:
https://api.agora.io/v1/projects/{appId}/rtsc/speech-to-text/tasks
After you acquire a builderToken, call this method within 5 minutes to start speech-to-text conversion.
Request
Path parameters
appId Type: string Required
The App ID of the project
Query parameters
builderToken Type: string Required
The tokenName value you obtained in the response body of the acquire method. To stop a task, use the same builderToken you used to start the task.
Request body
APPLICATION/JSON
BODY
languages Type: array[string] Required
The transcription languages to recognize. You can specify a maximum of 2 languages. Refer to Supported Languages for details.
maxIdleTime Type: integer Optional Default: 30 Values: 5 to 2592000
Maximum channel idle time, in seconds. When the specified time is exceeded, the transcription task ends automatically.
rtcConfig Type: object Required
channelName Type: string Required
The name of the channel to transcribe.
subBotUid Type: string Required
The ID of the bot that subscribes to the audio stream. All UIDs within a channel must be unique. Ensure no other user or service bot is using this UID in the same channel.
subBotToken Type: string Optional
The token used by the subscribing bot for channel authentication. Required only when your project has App Certificate enabled. Generate this token on your token server. For details, see Token authentication.
pubBotUid Type: string Required
The ID of the bot that pushes subtitle information to the channel. All UIDs within a channel must be unique. Ensure no other user or service bot is using this UID in the same channel.
pubBotToken Type: string Optional
The token used by the subtitle-pushing bot for channel authentication. Required only when your project has App Certificate enabled. Generate this token on your token server. For details, see Token authentication.
subscribeAudioUids Type: array[string] Optional
The user IDs of the audio streams you want to subscribe to. Specify this parameter only if you need to subscribe to specific users. Maximum array length: 3.
cryptionMode Type: integer Optional Values: 0 to 8
The encryption and decryption mode. When enabled, this mode is used for both decrypting incoming streams and encrypting outgoing subtitles.
0: No encryption1:AES_128_XTS128-bit AES encryption, XTS mode2:AES_128_ECB128-bit AES encryption, ECB mode3:AES_256_XTS256-bit AES encryption, XTS mode4:SM4_128_ECB128-bit SM4 encryption, ECB mode5:AES_128_GCM128-bit AES encryption, GCM mode6:AES_256_GCM256-bit AES encryption, GCM mode7:AES_128_GCM2128-bit AES encryption, GCM mode, requires key and salt8:AES_256_GCM2256-bit AES encryption, GCM mode, requires key and salt The decryption method must match the encryption method set for the channel.
secret Type: string Optional
The encryption/decryption key. Required when cryptionMode is not 0.
salt Type: string Optional
A Base64-encoded, 32-byte encryption/decryption salt. Required only when cryptionMode is 7 or 8.
enableJsonProtocol Type: boolean Optional Default: false
Set the encoding format of the subtitle data pushed to the channel.
true: Use JSON to push subtitles and compress data with gzip. Uses less bandwidth, but requires decoding.false: Use Protobuf to push subtitles (default). The data volume is smaller. Suitable for scenarios with high transmission efficiency requirements.
transferMode Type: array[string] Optional
Select the channel to push, RTM or RTC
captionConfig Type: object Optional
sliceDuration Type: integer Optional Default: 60 Values: 5 to 28800
The slice size of the recorded subtitle file, in seconds.
storage Type: object Required
accesskey Type: string Required
The access key of the third-party cloud storage.
secretkey Type: string Required
The secret key of the third-party cloud storage.
bucket Type: string Required
The bucket name of the third-party cloud storage.
vendor Type: integer Required Values: 1,5,6
The third-party cloud storage platform:
1: Amazon S35: Microsoft Azure6: Google Cloud
region Type: integer Required
The region information for the third-party cloud storage. To ensure successful and real-time uploading of recorded files, the cloud storage region must match the region of the application server where you initiate the request. For example, if your App server is in East US, set the cloud storage region to East US as well. See third-party storage regions for details.
fileNamePrefix Type: array[string] Optional
The storage location of the recorded file in the third-party cloud storage. The prefix length (including slashes) must not exceed 128 characters. The following characters are supported:
- Lowercase English letters (a-z)
- Uppercase English letters (A-Z)
- Numbers (0-9)
Symbols like slashes, underscores, and brackets must not appear in the string.
translateConfig Type: object Optional
forceTranslateInterval Type: integer Optional Default: 3 Values: 2 to 5
The time interval for forced translation, in seconds.
languages Type: array Required
The translation language array. You can specify a maximum of 2 different source languages. The source language and target language must be different, otherwise an error is reported.
Each array item is an object with:
source Type: string Required
The source language for translation. Refer to Supported Languages for details.
target Type: array[string] Required
The target languages for translation. You can specify a maximum of 5 target languages for each source language. Refer to Supported Languages for details.
Response
-
If the returned status code is
200, the request was successful. The response body contains the result of the request.OK
taskId Type: string Required
The unique identifier of this transcription task.
createTs Type: integer Required
The Unix timestamp (in seconds) when the transcription task was created.
status Type: string Required
The current status of the transcription task:
-
IDLE: Task not initialized -
PREPARING: Task has received an initialization request -
PREPARED: Task initialization completed -
STARTING: Task is beginning to start -
CREATED: Task startup partially completed -
STARTED: Task startup fully completed -
IN_PROGRESS: Task is currently running -
STOPPING: Task is in the process of being paused -
STOPPED: Task has been terminated -
FAILURE_STOP: Task termination failed -
If the returned status code is not
200, the request failed. Refer to themessagefield to understand the possible reasons for failure.
Non-200
message Type: string
The reason why the request failed.
Authorization
This endpoint requires Basic Auth.
Request example
curl --request POST \
--url 'https://api.agora.io/v1/projects/:appId/rtsc/speech-to-text/tasks?builderToken=your_builder_token' \
--header 'Authorization: Basic <credentials>' import requests
url = "https://api.agora.io/v1/projects/:appId/rtsc/speech-to-text/tasks?builderToken=your_builder_token"
headers = {"Authorization": "Basic <credentials>"}
response = requests.request("POST", url, headers=headers)
print(response.text) const url = 'https://api.agora.io/v1/projects/:appId/rtsc/speech-to-text/tasks?builderToken=your_builder_token';
const options = {method: 'POST', headers: {Authorization: 'Basic <credentials>'}};
fetch(url, options)
.then(res => res.json())
.then(json => console.log(json))
.catch(err => console.error(err));Response example
200
{
"taskId": "The unique identifier of this transcription task",
"createTs": null,
"status": "The status of the task"
}Non-200
{
"message": "The reason why the request failed.",
}