Start a Real-time STT agent
Updated
Starts subtitle recording and translation.
https://api.agora.io/api/speech-to-text/v1/projects/{appid}/joinUse this method to start subtitle recording and subtitle translation.
Path Parameters
The App ID of the project.
Request Body
application/json
The transcription languages you want to recognize. You can specify up to four languages. For a complete list, see Supported languages. Choosing multiple transcription languages can affect both quality and cost. For best practices, see Optimize transcription quality and cost.
4Configure the transcription language for the specified user ID. Supports up to 5 configuration items.
5Configure the transcription language for the specified user ID. Supports up to 5 configuration items.
The ID of the user to be transcribed. You may configure a maximum of 5 uids for language recognition at the uid level.
The transcription languages to recognize. Each uid can support a maximum of 4 languages. Refer to Supported Languages for details.
4Maximum channel idle time, in seconds. Value range: [0,259200]. Set maxIdleTime to 0 to disable automatic termination due to idle time. When the specified time is exceeded, the task ends automatically. Idle time means that there is no host in a live broadcast channel, or there is no user in a communication channel.
Independent of maxIdleTime, every task also has a maximum lifetime of 72 hours (259200 seconds). Once a task reaches this limit, Agora terminates it, even if maxIdleTime is 0.
30[0, 259200]Real-time subtitle configuration. After a user's voice is converted to text, the information is pushed to the channel as subtitles to match the UI real-time display.
The name of the channel to transcribe.
The ID of the bot that subscribes to the audio stream. This is always identical to the value of the pubBotUid.
The token used by the subscribing bot for channel authentication. Required only when your project has App Certificate enabled. Generate this token on your token server. For details, see Token authentication.
The ID of the bot that pushes subtitle information to the channel. All UIDs within a channel must be unique. Ensure no other user or service bot is using this UID in the same channel.
The token used by the subtitle-pushing bot for channel authentication. Required only when your project has App Certificate enabled. Generate this token on your token server. For details, see Token authentication.
The user IDs for the audio streams you want to subscribe. Set this parameter if you need to subscribe to the audio stream of certain users. Maximum array length: 32. You can set either subscribeAudioUids or unSubscribeAudioUids.
32The user IDs for the audio streams you do not want to subscribe. Set this parameter if you don't need to subscribe to the audio stream of certain users. Maximum array length: 5. You can set either subscribeAudioUids or unSubscribeAudioUids.
5The encryption and decryption mode. When enabled, this mode is used for both decrypting incoming streams and encrypting outgoing subtitles.
0: No encryption1:AES_128_XTS128-bit AES encryption, XTS mode2:AES_128_ECB128-bit AES encryption, ECB mode3:AES_256_XTS256-bit AES encryption, XTS mode4:SM4_128_ECB128-bit SM4 encryption, ECB mode5:AES_128_GCM128-bit AES encryption, GCM mode6:AES_256_GCM256-bit AES encryption, GCM mode7:AES_128_GCM2128-bit AES encryption, GCM mode, Compared withAES_128_GCMencryption mode, this encryption mode is more secure and requires setting a key and salt.8:AES_256_GCM2256-bit AES encryption, GCM mode, Compared withAES_256_GCMencryption mode, this encryption mode is more secure and requires setting a key and salt. The decryption method must match the encryption method set for the channel.
0 | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8The encryption/decryption key. Required when cryptionMode is not 0.
A Base64-encoded, 32-byte encryption/decryption salt. Required only when cryptionMode is 7 or 8.
Set the encoding format of the subtitle data pushed to the channel.
true: Use JSON to push subtitles and compress data with gzip. Uses less bandwidth, but requires decoding.false: Use Protobuf to push subtitles (default). The data volume is smaller. Suitable for scenarios with high transmission efficiency requirements.
A single update can contain both a stabilized prefix segment and a segment that may still change. Both JSON and Protobuf formats can carry multiple segments in a single message. For details, see Parse transcription data.
Subtitle translation configuration.
The translation languages array. You can specify a maximum of 4 different source languages.
The translation language array. You can specify a maximum of 4 different source languages.
Each array item is an object with:
4Translation language pair configuration.
The source language for translation. Refer to Supported Languages for details.
The target languages for translation. You can configure up to 10 target languages for each source language. Refer to Supported Languages for details.
- Single-language input: If you set the source language to a single language, the target language must be different, otherwise an error is returned. For example, if you set the source language to English, you cannot set the target language to English.
- Mixed-language input: If you set the source language to mixed-language input, you can set the target language to one of the source languages. For example, if you set the source languages to Chinese and English, setting the target language to English translates both into English.
10Subtitle recording configuration.
The slice size of the recorded subtitle file, in seconds.
60[5, 28800]The configuration for third-party cloud storage.
The access key of the third-party cloud storage.
The secret key of the third-party cloud storage.
The bucket name of the third-party cloud storage.
The third-party cloud storage platform:
1: Amazon S32: Alibaba Cloud3: Tencent Cloud5: Microsoft Azure6: Google Cloud7: Huawei Cloud8: Baidu Smart Cloud11: Other S3-compatible object storage systems, such as MinIO and self-hosted cloud storage systems
1 | 2 | 3 | 5 | 6 | 7 | 8 | 11The region information for the third-party cloud storage. To ensure successful and real-time uploading of recorded files, the cloud storage region must match the region of the application server where you initiate the request. For example, if your App server is in East US, set the cloud storage region to East US as well. See third-party storage regions for details.
The storage location of the recorded file in the third-party cloud storage. The prefix length (including slashes) must not exceed 128 characters. The following characters are supported:
- Lowercase English letters (a-z)
- Uppercase English letters (A-Z)
- Numbers (0-9) Symbols like slashes, underscores, and brackets must not appear in the string.
128Optional third-party cloud storage extension configuration. When storage.vendor is set to 11, use this field to specify access information for standard S3-compatible object storage.
The access URL for the S3-compatible service, including the scheme. For example, http://host:9002. Required when storage.vendor is 11.
The storage provider name. For example, Minio for MinIO.
The rclone S3 backend region. If provided, this overrides the default region inferred from storage.region.
A base string for the object tag. Only effective when storage.vendor is Tencent Cloud, Alibaba Cloud, or Amazon S3.
Appends key=value to the tag according to the filename rules. When both the tag and the filename exist, the rule takes precedence. Only applicable to Tencent Cloud, Alibaba Cloud, and Amazon S3.
Server-side encryption method. aes256 for AES-256 encryption, kms for AWS KMS. Only available for Amazon S3.
aes256 | kmsMaps the target object name using the uploaded file extension, used to override DstFileName.
Keyword list. Use it to improve the recognition accuracy of specific words during transcription. Supports up to 500 words.
500Unique ID of the agent. Maximum length is 64 characters. You cannot use the same ID repeatedly.
64Response
- If the returned status code is
200, the request was successful. The response body contains the result of the request.
Response Body
application/json
application/json
Response schema
200OK
The ID of the agent.
The Unix timestamp (in seconds) when the agent was created.
The current status of the agent:
IDLE: The agent is not initializedSTARTING: The agent is startingRUNNING: The agent is runningSTOPPING: The agent is exitingSTOPPED: The agent exited successfullyRECOVERING: The agent is recoveringFAILED: Agent exit failed
IDLE | STARTING | RUNNING | STOPPING | STOPPED | RECOVERING | FAILEDResponse
Refer to the detail and reason fields to understand the possible reasons for failure.
Request examples
curl --request POST \
--url https://api.agora.io/api/speech-to-text/v1/projects/:appid/join \
--header 'Authorization: Basic <credentials>' \
--data '{
"languages": [
"en-US"
],
"keywords": [
"Agora",
"STT"
],
"name": "agora-test",
"maxIdleTime": 50,
"rtcConfig": {
"channelName": "agora-test",
"pubBotUid": "88222"
},
"translateConfig": {
"languages": [
{
"source": "en-US",
"target": [
"ar-SA",
"id-ID",
"fr-FR",
"ja-JP"
]
}
]
},
"captionConfig": {
"sliceDuration": 60,
"storage": {
"accessKey": "test-oss",
"secretKey": "test-oss",
"bucket": "test-oss",
"vendor": 2,
"region": 3
}
}
}'
Response example
{ "agent_id": "Agent ID.", "create_ts": null, "status": "RUNNING"}