TypeScript
Updated
Full API reference for the Agora Agent TypeScript SDK — AgoraClient, Agent, AgentSession, and vendor classes.
Full API reference for the Agora Conversational AI TypeScript SDK.
AgoraClient
AgoraClient extends the Fern-generated base client with domain pool support for regional URL cycling and three authentication modes. Pass appId and appCertificate only for the recommended app-credentials mode. The SDK mints fresh REST tokens per request and generates RTC join tokens at session start.
import { AgoraClient, Area } from 'agora-agents';Constructor
const client = new AgoraClient(options: AgoraClient.Options);The authentication mode is resolved automatically from the options you provide.
| Option | Type | Required | Description |
|---|---|---|---|
area | Area | Yes | Region for API routing (Area.US, Area.EU, Area.AP, Area.CN) |
appId | string | Yes | Agora App ID |
appCertificate | string | Yes | Agora App Certificate. Keep this secret and never expose it client-side |
customerId | string | No | Customer ID for Basic Auth |
customerSecret | string | No | Customer Secret for Basic Auth |
authToken | string | No | Raw Agora token. The SDK sends Authorization: agora token=<authToken> |
timeoutInSeconds | number | No | Default request timeout in seconds |
maxRetries | number | No | Maximum retry attempts |
fetch | typeof fetch | No | Custom fetch implementation for unsupported runtimes |
Authentication mode is resolved from the options you provide:
| Options provided | Resolved authMode |
|---|---|
customerId + customerSecret | "basic" |
authToken | "token" |
| Neither | "app-credentials" |
Passing customerId without customerSecret throws an error.
See Authentication for details on each mode.
Properties
The following read-only properties are available on any AgoraClient instance.
| Property | Type | Description |
|---|---|---|
appId | string | The Agora App ID |
appCertificate | string | The Agora App Certificate |
authMode | AgoraAuthMode | The resolved authentication mode |
area | Area | The configured region |
pool | Pool | The underlying domain pool instance used for regional routing |
Methods
The following methods are available in addition to the Fern-generated sub-client methods.
nextRegion()
Cycles to the next region prefix in the domain pool. Call this after a request failure to try a different regional endpoint.
client.nextRegion();selectBestDomain(signal?)
Runs a DNS check to select the best domain suffix. Calls made within 30 seconds of the last successful selection are no-ops.
await client.selectBestDomain();| Parameter | Type | Description |
|---|---|---|
signal | AbortSignal | Optional abort signal to cancel the DNS check |
stopAgent(agentId)
Stops an agent by ID without an AgentSession reference, for example, from an end-call handler. The method treats a 404 response as success.
stopAgent(agentId: string): Promise<void>getCurrentURL()
Returns the full API URL currently in use as a string.
const url = client.getCurrentURL();
// Example: 'https://api-us-west-1.agora.io/api/conversational-ai-agent'Sub-clients
AgoraClient exposes Fern-generated sub-clients for direct REST API access. You typically do not need these when using the agentkit layer.
| Property | Description |
|---|---|
client.agents | Start, stop, update, speak, interrupt, get history, list agents |
client.agentManagement | Inject instructions into a running agent (think) |
client.telephony | Telephony operations |
client.phoneNumbers | Phone number management |
For full method signatures and request parameters, see the REST API reference.
Agent
Agent is an immutable configuration object. Each builder method returns a new Agent instance — the original is never modified. Define one Agent at startup and call createSession() on it for each user conversation.
import { Agent } from 'agora-agents';Constructor
new Agent(options: AgentOptions)client is required. All other options are optional. Use the builder methods to set vendor configuration after construction.
| Option | Type | Default | Description |
|---|---|---|---|
client | AgoraClient | — | Required. Agora client used by createSession() |
pipelineId | string | undefined | Published AI Studio pipeline ID used as the base configuration. SessionOptions.pipelineId overrides it |
instructions | string | undefined | Deprecated. Set systemMessages on the LLM vendor instead |
greeting | string | undefined | Deprecated. Set greetingMessage on the LLM or MLLM vendor instead |
failureMessage | string | undefined | Deprecated. Set failureMessage on the LLM or MLLM vendor instead |
maxHistory | number | undefined | Deprecated. Set maxHistory on the LLM vendor instead |
greetingConfigs | LlmGreetingConfigs | undefined | Deprecated. Set greetingConfigs on the LLM vendor instead |
turnDetection | TurnDetectionConfig | undefined | Voice activity detection settings |
interruption | InterruptionConfig | undefined | Unified interruption control settings |
sal | SalConfig | undefined | Selective Attention Locking configuration |
avatar | AvatarConfig | undefined | Avatar configuration |
advancedFeatures | AdvancedFeatures | undefined | Enable MLLM mode, AI-VAD, and other advanced features |
parameters | SessionParamsInput | undefined | Session parameters including silence and farewell config |
geofence | GeofenceConfig | undefined | Regional access restriction |
labels | Labels | undefined | Custom key-value labels returned in notification callbacks |
rtc | RtcConfig | undefined | RTC media encryption |
fillerWords | FillerWordsConfig | undefined | Filler words played while waiting for the LLM response |
Builder methods
All builder methods return a new Agent instance. The original is never modified.
withLlm(vendor)
Sets the LLM vendor. Pass an instance of OpenAI, AzureOpenAI, Anthropic, Gemini, or any other LLM vendor.
withLlm(vendor: LlmVendor): Agent<TTSSampleRate, TArea>withTts(vendor)
Sets the TTS vendor. The sample rate type is captured and tracked for avatar compatibility.
withTts<SR extends number>(vendor: TtsVendor<SR>): Agent<SR, TArea>withStt(vendor)
Sets the STT vendor. Pass an instance of any STT vendor class.
withStt(vendor: SttVendor): Agent<TTSSampleRate, TArea>withMllm(vendor)
Sets the MLLM vendor for multimodal mode. Pass OpenAIRealtime, AzureOpenAIRealtime, GeminiLive, VertexAI, XaiGrok, or OpenAIGPTLive. Calling withMllm() automatically sets mllm.enable = true. MLLM mode does not require withTts() / withLlm() / withStt().
Avatars are only supported with the cascading ASR + LLM + TTS pipeline. If you combine withMllm() with withAvatar(), the SDK throws an error when toProperties() or session.start() is called.
withMllm(vendor: GlobalMllmVendor | CNMllmVendor): Agent<TTSSampleRate, TArea>withAvatar(vendor)
Sets the avatar vendor. The this constraint enforces at compile time that the agent's TTS sample rate matches the avatar's required rate.
Requires the cascading ASR + LLM + TTS pipeline. If you combine withAvatar() with withMllm(), the SDK throws an error when toProperties() or session.start() is called.
withAvatar<RequiredSR extends number>(
this: Agent<RequiredSR>,
vendor: AvatarVendor<RequiredSR>
): Agent<RequiredSR, TArea>withTurnDetection(config)
Configures cascading-flow turn detection. Pass { config: { start_of_speech, end_of_speech } } for SOS/EOS detection. Use withInterruption() for interruption behavior and MLLM vendor turnDetection for MLLM turn detection.
withTurnDetection(config: TurnDetectionConfig): Agent<TTSSampleRate, TArea>withInterruption(config)
Configures unified interruption behavior using the top-level interruption object. Use this for start_of_speech and keywords interruption modes.
withInterruption(config: InterruptionConfig): Agent<TTSSampleRate, TArea>withInstructions(text)
Overrides the LLM system prompt on a new Agent instance. Deprecated. Set systemMessages on the LLM vendor instead.
withInstructions(instructions: string): Agent<TTSSampleRate, TArea>withGreeting(text)
Overrides the greeting message on a new Agent instance. Deprecated. Set greetingMessage on the LLM or MLLM vendor instead.
withGreeting(greeting: string): Agent<TTSSampleRate, TArea>Other builder methods
The following methods follow the same pattern — each returns a new Agent instance with the updated configuration.
| Method | Parameter type | Description |
|---|---|---|
withSal(config) | SalConfig | Set Selective Attention Locking configuration |
withAdvancedFeatures(features) | AdvancedFeatures | Set advanced features |
withTools(enabled = true) | boolean | Enable or disable MCP tool and custom tool invocation |
withParameters(parameters) | SessionParamsInput | Set session parameters |
withAudioScenario(audioScenario) | ParametersAudioScenario | Set parameters.audio_scenario. Use the exported AudioScenario constants for discoverability, for example agent.withAudioScenario(AudioScenario.Aiserver) |
withFailureMessage(message) | string | Deprecated. Set failureMessage on the LLM or MLLM vendor instead |
withMaxHistory(n) | number | Deprecated. Set maxHistory on the LLM vendor instead. For Azure OpenAI Realtime MLLM, set maxHistory on the MLLM vendor |
withGreetingConfigs(configs) | LlmGreetingConfigs | Deprecated. Set greetingConfigs on the LLM vendor instead |
withGeofence(geofence) | GeofenceConfig | Set geofence configuration |
withLabels(labels) | Labels | Set custom labels |
withRtc(rtc) | RtcConfig | Set RTC configuration |
withFillerWords(fillerWords) | FillerWordsConfig | Set filler words configuration |
createSession(options)
Creates an AgentSession for a channel using the agent's client. Doesn't start the agent. Call session.start() to join the channel.
createSession(options: SessionOptions): AgentSessionSessionOptions fields:
| Option | Type | Required | Description |
|---|---|---|---|
channel | string | Yes | Channel name to join |
agentUid | string | Yes | The agent's RTC UID. Must be a numeric string when token is omitted |
remoteUids | string[] | Yes | Remote user UIDs the agent listens and responds to |
name | string | No | Agent instance name sent to the Agora API. Defaults to agent-{timestamp} |
token | string | No | Pre-built RTC+RTM token. Omit to auto-generate from app credentials |
expiresIn | number | No | Token lifetime in seconds. Only applies when the token is auto-generated. Valid range: 1–86400. Use ExpiresIn helpers for clarity |
idleTimeout | number | No | Seconds before the agent auto-exits when no audio is detected. 0 disables the timeout |
enableStringUid | boolean | No | Use string UIDs instead of numeric UIDs |
preset | PresetInput | No | Advanced project-specific presets. Use only when Agora provides a specific preset ID for your project. Most applications should not set this field |
pipelineId | string | No | Published AI Studio pipeline ID to use as the base configuration |
debug | boolean | No | Log API requests to the console |
warn | (message: string) => void | No | Custom warning logger; pass a no-op to silence warnings |
preset is session-scoped because the underlying Agora start/join API applies presets per session, not per reusable Agent definition.
When you omit credentials for supported reseller-backed vendor models, AgentKit infers the matching session preset automatically:
- Deepgram STT:
nova-2,nova-3 - OpenAI LLM:
gpt-4o-mini,gpt-4.1-mini,gpt-5-nano,gpt-5-mini - OpenAI TTS:
tts-1 - MiniMax TTS:
speech-2.6-turbo,speech-2.8-turbo
If you provide your own vendor API key for those same models, AgentKit keeps the request in BYOK mode and does not infer a preset.
Properties
Read-only properties available on any Agent instance.
| Property | Type | Description |
|---|---|---|
pipelineId | string | undefined | AI Studio pipeline ID |
greetingConfigs | LlmGreetingConfigs | undefined | Greeting playback configuration |
instructions | string | undefined | LLM system prompt |
greeting | string | undefined | Greeting message |
failureMessage | string | undefined | Message spoken when LLM fails |
maxHistory | number | undefined | Maximum conversation history length |
llm | LlmConfig | undefined | LLM configuration |
tts | TtsConfig | undefined | TTS configuration |
stt | SttConfig | undefined | STT configuration |
mllm | MllmConfig | undefined | MLLM configuration |
avatar | AvatarConfig | undefined | Avatar configuration |
turnDetection | TurnDetectionConfig | undefined | Turn detection configuration |
interruption | InterruptionConfig | undefined | Interruption configuration |
sal | SalConfig | undefined | SAL configuration |
advancedFeatures | AdvancedFeatures | undefined | Advanced features |
parameters | SessionParams | undefined | Session parameters |
geofence | GeofenceConfig | undefined | Geofence configuration |
labels | Labels | undefined | Custom labels |
rtc | RtcConfig | undefined | RTC configuration |
fillerWords | FillerWordsConfig | undefined | Filler words configuration |
config | AgentOptions | Full read-only configuration snapshot |
toProperties(opts)
Low-level method to convert the agent configuration to the Fern request format. Used internally by AgentSession.start(). You typically do not need to call this directly unless building custom request bodies.
toProperties(opts): StartAgentsRequest.PropertiesIn cascading mode, throws if TTS or LLM isn't set. If you don't call withStt(), ASR defaults to ARES (fengming for Area.CN). turnDetection.language defaults to en-US and is also sent as asr.language.
Type aliases
Public aliases over Fern-generated types include LlmConfig, SttConfig, AsrConfig (= SttConfig), MllmConfig, AvatarConfig, session/conversation types, and think types (ThinkOnListeningAction, etc.).
Think value constants: ThinkOnListeningActionInject, ThinkOnListeningActionInterrupt, ThinkOnListeningActionIgnore, ThinkOnThinkingActionInterrupt, ThinkOnThinkingActionIgnore, ThinkOnSpeakingActionInterrupt, ThinkOnSpeakingActionIgnore.
The think action options also accept the string append, which has no convenience constant. With append, the instruction doesn't interrupt the current interaction. The agent waits for the current user turn, LLM inference, or TTS playback to finish, then appends the instruction to the context as a separate user message and starts a new turn.
FillerWordsConfig supports static and generated filler words. Generated mode requires static fallback phrases, while generated_config and prompt are optional. Agora hosts the generation service, so you can't configure its model endpoint, API key, or model parameters. For the field structure, see Configure generated filler words.
FillerWordsContentGeneratedConfig sets the conversation context for generated filler words. Use context_message_limit (1 to 6, defaults to 1) to set how many of the most recent conversation messages to use, counting the current turn's message, and history_character_limit (0 to 10000, defaults to 1000) to cap the combined characters of the earlier messages. For details, see Talking while waiting.
Public aliases also include FillerWordsContentGeneratedConfig and the custom tool types LlmTool, LlmToolFunction, LlmToolServer, and LlmToolExecution. llm.tools is independent of MCP server configuration, but both require withTools(true) to enable tool calling.
AgentSession
AgentSession manages the full lifecycle of a running agent. Create sessions using agent.createSession(); do not call the constructor directly.
import { AgentSession } from 'agora-agents';State machine
A session progresses through the following states:
idle ──► starting ──► running ──► stopping ──► stopped
│ │
▼ ▼
error error| Transition | Trigger |
|---|---|
idle → starting | start() called |
starting → running | API responds with agent ID |
starting → error | API request fails |
running → stopping | stop() called |
stopping → stopped | API confirms agent stopped |
stopping → error | Stop request fails with a non-404 error |
start() can also be called from stopped or error state to restart the session. say(), interrupt(), update(), and think() reject on failure but don't change the session state.
Methods
The following methods are available on an AgentSession instance.
start()
Starts the agent session. Generates tokens if not provided, sends the start request, and returns the agent ID. Resolves explicit preset values and also infers reseller presets from supported vendor configs when credentials are omitted.
start(): Promise<string>- Transitions:
idle/stopped/error→starting→running - Throws if called in
starting,running, orstoppingstate - Throws if avatar config is invalid (wrong TTS sample rate)
- Throws if no
tokenis provided and the client has noappCertificate - Throws if
agentUid(or avataragoraUid) isn't numeric when tokens are auto-generated - Throws if
appIdorappCertificateisn't exactly 32 characters when tokens are auto-generated - Throws if MLLM is enabled together with an enabled avatar — avatars are only supported with the cascading ASR + LLM + TTS pipeline
- Applies explicit
presetvalues when provided and sends Agora-managed configuration when supported vendor credentials are omitted - Fills generic avatar
agora_appidandagora_channelfrom the session when omitted - Generates avatar
agora_tokenforHeyGenAvatar,LiveAvatarAvatar,GenericAvatar,Tavus,Protoface, andLemonSlicewhenagoraTokenis omitted and the client has anappCertificate. Other vendors (AkoolAvatar,AnamAvatar) never receive an auto-generated token.
stop()
Stops the agent session and removes the agent from the channel. If the agent has already stopped — for example due to idle timeout — resolves silently rather than throwing a 404 error.
stop(): Promise<void>- Transitions:
running→stopping→stopped - Throws if called outside
runningstate
say(text, options?)
Instructs the agent to speak the given text.
say(text: string, options?: SayOptions): Promise<void>| Parameter | Type | Required | Description |
|---|---|---|---|
text | string | Yes | The text for the agent to speak |
options.priority | SpeakPriority | No | Message priority |
options.interruptable | boolean | No | Whether this message can be interrupted by the user |
- Only valid in
runningstate
interrupt()
Interrupts the agent's current speech.
interrupt(): Promise<void>- Only valid in
runningstate
think(text, options?)
Injects a custom text instruction into the running agent.
think(text: string, options?: ThinkOptions): Promise<ThinkResponse>| Parameter | Type | Required | Description |
|---|---|---|---|
text | string | Yes | The instruction text |
options.on_listening_action | ThinkOnListeningAction | No | Action when the agent is listening: inject, interrupt, ignore, or append |
options.on_thinking_action | ThinkOnThinkingAction | No | Action when the agent is thinking: interrupt, ignore, or append |
options.on_speaking_action | ThinkOnSpeakingAction | No | Action when the agent is speaking: interrupt, ignore, or append |
options.interruptable | boolean | No | Whether the user can interrupt the resulting response |
options.metadata | Record<string, string> | No | Custom key-value metadata |
- Only valid in
runningstate
update(config)
Updates the agent configuration mid-session without restarting. Accepts a partial configuration object in REST API format.
update(config: AgentConfigUpdate): Promise<void>- Only valid in
runningstate
getHistory()
Fetches the conversation history for this session. Requires a valid agent ID — start() must have been called successfully.
getHistory(): Promise<ConversationHistory>getTurns(options?)
Fetches turn-by-turn analytics for this session, including start/end events and latency metrics. Requires a valid agent ID and start() must have been called successfully.
getTurns(options?: GetTurnsOptions): Promise<ConversationTurns>options.page_index: page number, starting from1options.page_size: number of turns per page
getAllTurns(options?)
Fetches all turn analytics pages and merges the turns array. Requires a valid agent ID and start() must have been called successfully.
getAllTurns(options?: Omit<GetTurnsOptions, "page_index">): Promise<ConversationTurns>- Requires a valid
agentId - For very long sessions, prefer processing pages with
getTurns()to avoid holding all turns in memory
getInfo()
Fetches current agent metadata from the API. Requires a valid agent ID.
getInfo(): Promise<SessionInfo>on(event, handler)
Subscribes to a session event. Register handlers before calling start() to avoid missing the started event.
on<T>(event: AgentSessionEvent, handler: AgentSessionEventHandler<T>): voidoff(event, handler)
Unsubscribes a previously registered event handler.
off<T>(event: AgentSessionEvent, handler: AgentSessionEventHandler<T>): voidEvents
The session emits the following events. See AgentSessionEvent and AgentSessionEventHandler for type details.
| Event | Payload type | Description |
|---|---|---|
"started" | { agentId: string } | Agent successfully joined the channel |
"stopped" | { agentId: string } | Agent left the channel |
"error" | unknown | The error thrown during start() or stop() |
Properties
The following read-only properties are available on any AgentSession instance.
| Property | Type | Description |
|---|---|---|
status | "idle" | "starting" | "running" | "stopping" | "stopped" | "error" | Current session state. One of "idle", "starting", "running", "stopping", "stopped", "error" |
id | string | null | Agent ID, populated after start() resolves |
agent | Agent | The agent configuration this session was created from |
appId | string | The Agora App ID for this session |
raw | AgentsClient | Direct access to the Fern-generated AgentsClient for advanced operations |
rawAgentManagement | AgentManagementClient | Direct access to the Fern-generated agent management client |
Using session.raw
Access the generated REST client to call endpoints not yet wrapped:
await session.raw.getTurns({
appid: session.appId,
agentId: session.id!,
});You must pass appid and agentId manually when using raw methods.
Presets and BYOK
Prefer configuring vendors on the Agent builder. When you omit credentials for supported Agora-managed models, AgentKit sends the matching Agora-managed configuration at session start.
preset is an advanced session option for project-specific settings, not for selecting Agora-managed models. Most applications should use the builder instead.
- Omit vendor credentials on the builder for supported Agora-managed models.
- Provide vendor API keys when you want BYOK.
- Pass
presetonagent.createSession(...)only when you need to access specific project-specific settings.
Supported Agora-managed models:
- Deepgram STT:
nova-2,nova-3 - OpenAI LLM:
gpt-4o-mini,gpt-4.1-mini,gpt-5-nano,gpt-5-mini - OpenAI TTS:
tts-1 - MiniMax TTS:
speech-2.6-turbo,speech-2.8-turbo
Vendors
All vendor classes are imported from agora-agents. Pass vendor instances to the Agent builder methods.
LLM vendors
Use with withLlm().
OpenAI
new OpenAI(options: OpenAIOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Usually | OpenAI API key. Optional for Agora-managed preset models |
model | string | Yes | Model name, for example 'gpt-4o-mini' |
url | string | Conditional | API endpoint URL. Required when apiKey is set (BYOK); default https://api.openai.com/v1/chat/completions |
maxHistory | number | No | Maximum conversation history to cache |
temperature | number | No | Sampling temperature (0.0–2.0) |
topP | number | No | Nucleus sampling (0.0–1.0) |
maxTokens | number | No | Maximum tokens to generate |
systemMessages | Record<string, unknown>[] | No | Additional system messages |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message spoken when the LLM call fails |
inputModalities | string[] | No | Input modalities. Default: ["text"] |
outputModalities | string[] | No | Output modalities |
params | Record<string, unknown> | No | Additional LLM parameters passed to the model |
headers | Record<string, string> | No | Custom HTTP headers forwarded to the LLM provider |
vendor | string | No | Vendor override |
mcpServers | Record<string, unknown>[] | No | MCP server connections |
tools | LlmTool[] | No | Synchronous custom tool definitions the LLM can call. Requires withTools(true) |
greetingConfigs | LlmGreetingConfigs | No | Greeting playback configuration |
templateVariables | Record<string, string> | No | Template variables for messages. Custom tools can reference them with {{template_variables.<name>}} |
apiKey is optional for the following reseller preset models: gpt-4o-mini, gpt-4.1-mini, gpt-5-nano, gpt-5-mini. If apiKey is omitted for one of those models, AgentKit infers the matching session preset. This no-key branch is only available with the default OpenAI endpoint and without a custom vendor hint. If apiKey is provided, AgentKit uses standard BYOK behavior instead.
AzureOpenAI
new AzureOpenAI(options: AzureOpenAIOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Azure OpenAI API key |
model | string | Yes | Model or deployment name |
resourceName | string | Conditional | Azure resource name. Required unless endpoint is set |
endpoint | string | Conditional | Full Azure base URL. Takes precedence over resourceName; required unless resourceName is set |
deploymentName | string | Yes | Deployment name in Azure |
apiVersion | string | No | Azure API version. Default: '2024-08-01-preview' |
maxHistory | number | No | Maximum conversation history to cache |
temperature | number | No | Sampling temperature (0.0–2.0) |
topP | number | No | Nucleus sampling (0.0–1.0) |
maxTokens | number | No | Maximum tokens to generate |
systemMessages | Record<string, unknown>[] | No | Additional system messages |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message spoken when the LLM call fails |
inputModalities | string[] | No | Input modalities. Default: ["text"] |
outputModalities | string[] | No | Output modalities |
params | Record<string, unknown> | No | Additional LLM parameters |
headers | Record<string, string> | No | Custom HTTP headers forwarded to the LLM provider |
vendor | string | No | Vendor override. Defaults to azure |
mcpServers | Record<string, unknown>[] | No | MCP server connections |
tools | LlmTool[] | No | Synchronous custom tool definitions the LLM can call. Requires withTools(true) |
greetingConfigs | LlmGreetingConfigs | No | Greeting playback configuration |
templateVariables | Record<string, string> | No | Template variables for messages. Custom tools can reference them with {{template_variables.<name>}} |
Anthropic
new Anthropic(options: AnthropicOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Anthropic API key |
model | string | Yes | Model name |
url | string | Yes | Anthropic messages endpoint URL |
maxTokens | number | Yes | Maximum tokens to generate |
headers | Record<string, string> | Yes | Request headers, including anthropic-version |
maxHistory | number | No | Maximum conversation history to cache |
temperature | number | No | Sampling temperature (0.0–1.0) |
topP | number | No | Nucleus sampling (0.0–1.0) |
systemMessages | Record<string, unknown>[] | No | Additional system messages |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message spoken when the LLM call fails |
inputModalities | string[] | No | Input modalities. Default: ["text"] |
outputModalities | string[] | No | Output modalities |
params | Record<string, unknown> | No | Additional LLM parameters |
vendor | string | No | Vendor override |
mcpServers | Record<string, unknown>[] | No | MCP server connections |
tools | LlmTool[] | No | Synchronous custom tool definitions the LLM can call. Requires withTools(true) |
greetingConfigs | LlmGreetingConfigs | No | Greeting playback configuration |
templateVariables | Record<string, string> | No | Template variables for messages. Custom tools can reference them with {{template_variables.<name>}} |
Gemini
new Gemini(options: GeminiOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Google API key |
model | string | Yes | Model name, for example 'gemini-pro' |
url | string | No | API endpoint URL. Default: https://generativelanguage.googleapis.com/v1beta/models/{model}:streamGenerateContent?alt=sse&key={apiKey}, with the API key in the URL |
maxHistory | number | No | Maximum conversation history to cache |
temperature | number | No | Sampling temperature (0.0–2.0) |
topP | number | No | Nucleus sampling (0.0–1.0) |
topK | number | No | Top-k sampling |
maxOutputTokens | number | No | Maximum output tokens to generate |
systemMessages | Record<string, unknown>[] | No | Additional system messages |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message spoken when the LLM call fails |
inputModalities | string[] | No | Input modalities. Default: ["text"] |
outputModalities | string[] | No | Output modalities |
params | Record<string, unknown> | No | Additional LLM parameters |
headers | Record<string, string> | No | Custom HTTP headers forwarded to the LLM provider |
vendor | string | No | Vendor override |
mcpServers | Record<string, unknown>[] | No | MCP server connections |
tools | LlmTool[] | No | Synchronous custom tool definitions the LLM can call. Requires withTools(true) |
greetingConfigs | LlmGreetingConfigs | No | Greeting playback configuration |
templateVariables | Record<string, string> | No | Template variables for messages. Custom tools can reference them with {{template_variables.<name>}} |
Other LLM vendors
| Class | Provider | Key options |
|---|---|---|
Groq | Groq | apiKey, model, url |
VertexAILLM | Google Vertex AI | apiKey, model, projectId, location, url? |
AmazonBedrock | Amazon Bedrock | accessKey, secretKey, region, model |
Dify | Dify | apiKey, url, model, user?, conversationId? |
CustomLLM | OpenAI-compatible LLM | apiKey, model, url |
Groq and CustomLLM share OpenAI's option shape; VertexAILLM mirrors Gemini. All accept common LLM fields, such as systemMessages, greetingMessage, failureMessage, maxHistory, params, headers, and tools.
tools declares custom tools the LLM can choose to call. Each tool includes a model-visible function definition and the server configuration for a synchronous GET or POST request. withTools(true) applies to both tools and MCP server configuration. For more information, see Call custom tools.
TTS vendors
Use with withTts(). The sampleRate option determines avatar compatibility — see withAvatar().
ElevenLabsTTS
new ElevenLabsTTS<SR extends ElevenLabsSampleRate>(options: ElevenLabsTTSOptions<SR>)| Option | Type | Required | Description |
|---|---|---|---|
key | string | Yes | ElevenLabs API key |
modelId | string | Yes | Model ID, for example 'eleven_flash_v2_5' |
voiceId | string | Yes | Voice ID |
baseUrl | string | Yes | WebSocket base URL |
sampleRate | 16000 | 22050 | 24000 | 44100 | No | Audio sample rate in Hz |
optimizeStreamingLatency | number | No | Latency optimization level (0–4) |
stability | number | No | Voice stability (0.0–1.0) |
similarityBoost | number | No | Voice similarity boost (0.0–1.0) |
style | number | No | Voice style exaggeration (0.0–1.0) |
useSpeakerBoost | boolean | No | Enable speaker boost |
skipPatterns | number[] | No | Skip patterns for bracketed content |
MicrosoftTTS
new MicrosoftTTS<SR extends MicrosoftSampleRate>(options: MicrosoftTTSOptions<SR>)| Option | Type | Required | Description |
|---|---|---|---|
key | string | Yes | Azure Speech API key |
region | string | Yes | Azure region, for example 'eastus' |
voiceName | string | Yes | Voice name, for example 'en-US-JennyNeural' |
sampleRate | 16000 | 24000 | 48000 | No | Audio sample rate in Hz |
speed | number | No | Speaking rate multiplier |
volume | number | No | Audio volume |
skipPatterns | number[] | No | Skip patterns for bracketed content |
OpenAITTS
Fixed at 24,000 Hz — no configurable sample rate.
new OpenAITTS(options: OpenAITTSOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Usually | OpenAI API key |
voice | string | Yes | Voice name: 'alloy', 'echo', 'fable', 'onyx', 'nova', or 'shimmer' |
model | string | No | Model name. Required (with apiKey and baseUrl) for BYOK |
baseUrl | string | No | Endpoint URL. Required (with apiKey and model) for BYOK |
instructions | string | No | Custom voice instructions |
speed | number | No | Speech speed multiplier |
skipPatterns | number[] | No | Skip patterns for bracketed content |
apiKey is optional only for the reseller-backed tts-1 preset path. If omitted with model: 'tts-1' or no explicit model, AgentKit infers openai_tts_1. If provided, the request stays in BYOK mode.
CartesiaTTS
new CartesiaTTS<SR extends CartesiaSampleRate>(options: CartesiaTTSOptions<SR>)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Cartesia API key |
voiceId | string | Yes | Voice ID (serialized as {"mode": "id", "id": "..."}) |
modelId | string | Yes | Model ID |
baseUrl | string | No | WebSocket URL |
language | string | No | Target language |
sampleRate | 8000 | 16000 | 22050 | 24000 | 44100 | 48000 | No | Audio sample rate in Hz |
skipPatterns | number[] | No | Skip patterns for bracketed content |
Other TTS vendors
| Class | Key parameters |
|---|---|
GoogleTTS | key, voiceName, languageCode?, sampleRate? |
AmazonTTS | accessKey, secretKey, region, voiceId, engine |
DeepgramTTS | apiKey, model, baseUrl?, sampleRate?, additionalParams? |
HumeAITTS | key, voiceId, provider, configId?, baseUrl?, speed?, trailingSilence? |
RimeTTS | key, speaker, modelId, baseUrl?, credentialMode?. In managed mode (CredentialMode.Managed), only modelId and baseUrl are required |
FishAudioTTS | key, referenceId, backend |
MiniMaxTTS | key?, groupId?, model, voiceId?, url? |
MurfTTS | key, voiceId?, baseUrl?, locale?, rate?, pitch?, model?, sampleRate? |
SarvamTTS | key, speaker, targetLanguageCode, pitch?, pace?, loudness?, sampleRate? |
GradiumTTS | apiKey, url?, modelName?, voiceId?, sampleRate?, additionalParams? |
MistralTTS | apiKey, model?, voice?, additionalParams? |
TypecastTTS | apiKey, voiceId, model, additionalParams? |
XAiTTS | apiKey, language, voiceId?, sampleRate?, additionalParams? |
SmallestAITTS | apiKey, url?, model?, voiceId?, sampleRate?, speed?, language?, additionalParams? |
SmallestAITTS also accepts numberPronunciationLanguage?, mathNotation?, pronunciationDicts?, sessionId?, and requestId?. Voice identifiers are model-specific, so voiceId must belong to the model set in model. For details, see Smallest AI.
For MiniMaxTTS, key is optional only for reseller-backed models: speech-2.6-turbo, speech-2.8-turbo. If key is omitted for one of those models, AgentKit infers the matching session preset. In that preset-backed path, groupId, voiceId, and url are optional overrides rather than required fields. If key is provided, AgentKit uses BYOK, and groupId, voiceId, and url are required.
GenericTTS
Custom OpenAI-compatible HTTP TTS. url is required and must be an absolute HTTP or HTTPS address — the constructor throws if url is missing, badly formatted, or uses a non-HTTP(S) scheme such as ws: or wss:. A valid URL is serialized as tts.vendor = "generic_http".
new GenericTTS(options: GenericTTSOptions)| Option | Type | Required | Description |
|---|---|---|---|
url | string | Yes | The HTTP(S) endpoint of your custom TTS service |
headers | Record<string, string> | No | Custom HTTP headers to forward to the TTS service. Omitted from the request if not set |
apiKey | string | No | The API key used to authenticate with the TTS service |
model | string | No | The TTS model name |
voice | string | No | The voice name |
speed | number | No | The speech rate |
sampleRate | number | No | The sample rate, in Hz, of the output audio. If your TTS service doesn't support multiple sample rates, make sure the returned audio's sample rate matches this value |
responseFormat | "pcm" | No | The output audio format. Conversational AI Engine currently supports pcm |
instruction | string | No | Instructions for voice style, emotion, or other playback directives |
additionalParams | Record<string, unknown> | No | Additional parameters passed through to the TTS service. Explicit fields with the same name take precedence |
skipPatterns | number[] | No | Skip patterns for bracketed content |
STT vendors
Use with withStt().
DeepgramSTT
new DeepgramSTT(options?: DeepgramSTTOptions)Provide apiKey, or omit it and set model to nova-2 or nova-3 to use Agora-managed credentials. The constructor throws otherwise.
| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | No | Deepgram API key. Optional for nova-2 and nova-3 reseller preset usage |
model | string | Conditional | Model name, for example 'nova-2' or 'enhanced'. Required when apiKey is omitted |
language | string | No | Language code, for example 'en-US' |
keyterm | string | No | Boost specialized terms and brands |
smartFormat | boolean | No | Enable smart formatting |
punctuation | boolean | No | Enable punctuation |
additionalParams | Record<string, unknown> | No | Additional vendor parameters |
For nova-2 and nova-3, omit apiKey to use Agora-managed credentials. For all other Deepgram models, apiKey is required.
Other STT vendors
| Class | Key parameters |
|---|---|
SpeechmaticsSTT | key, language, model?, uri?. apiKey is a deprecated alias of key |
MicrosoftSTT | key, region, language |
OpenAISTT | apiKey, model?, language?, prompt?, inputAudioTranscription? |
GoogleSTT | projectId, location, adcCredentialsString, language, model? |
AmazonSTT | accessKey, secretKey, region, language |
AssemblyAISTT | apiKey, language, ws_url? |
AresSTT | keywords?, additionalParams? |
GeminiSTT | apiKey, model?, language?, languageHints?, customVocabulary?, sampleRate? (default 16000), wordTimestamp?, mode? (SMART or VERBATIM), diarization?, additionalParams? |
XAiSTT | apiKey, language?, baseUrl?, sampleRate?, additionalParams? |
SarvamSTT | apiKey, language, model? |
SmallestAISTT | apiKey, language?, url?, sampleRate?, encoding?, keywords?, additionalParams? |
SmallestAISTT also accepts the boolean options wordTimestamps?, sentenceTimestamps?, diarize?, vadEvents?, endpointing?, format?, finalizeOnWords?, punctuate?, capitalize?, itnNormalize?, fullTranscript?, redactPii?, and redactPci?, plus eouTimeoutMs? and maxWords?. Boolean options are serialized as the strings "true" or "false" for the Smallest AI wire protocol. For details, see Smallest AI.
For OpenAISTT, the serialized configuration requires a transcription prompt and language — provide them through the prompt and language options or within inputAudioTranscription. model defaults to gpt-4o-mini-transcribe.
MLLM vendors
Use with withMllm() for multimodal end-to-end audio processing without separate STT or TTS steps. Calling withMllm() automatically enables the MLLM module (sets mllm.enable = true); the older advancedFeatures.enable_mllm flag is deprecated.
OpenAIRealtime
new OpenAIRealtime(options: OpenAIRealtimeOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | OpenAI API key |
model | string | No | Model name, for example 'gpt-4o-realtime-preview' |
voice | string | No | Voice identifier |
instructions | string | No | System instructions |
inputAudioTranscription | Record<string, unknown> | No | Audio transcription settings |
url | string | No | WebSocket URL |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message played when the model call fails |
inputModalities | string[] | No | Input modalities, for example ['audio'] |
outputModalities | string[] | No | Output modalities, for example ['text', 'audio'] |
messages | Record<string, unknown>[] | No | Conversation messages for short-term memory |
params | Record<string, unknown> | No | Additional MLLM parameters |
turnDetection | MllmTurnDetectionConfig | No | MLLM turn detection configuration; overrides top-level turnDetection |
AzureOpenAIRealtime
new AzureOpenAIRealtime(options: AzureOpenAIRealtimeOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Azure OpenAI API key |
url | string | Yes | Azure OpenAI Realtime WebSocket URL |
turnDetection | MllmTurnDetectionConfig | Yes | MLLM turn detection configuration; overrides top-level turnDetection |
model | string | No | Azure OpenAI Realtime model or deployment name |
voice | string | No | Voice identifier |
instructions | string | No | System instructions |
maxHistory | number | No | Number of conversation history messages to cache |
greetingMessage | string | No | Agent greeting message |
outputModalities | string[] | No | Output modalities, for example ['text', 'audio'] |
messages | Record<string, unknown>[] | No | Conversation messages for short-term memory |
params | AzureOpenAIRealtimeParams | No | Additional Azure OpenAI parameters |
GeminiLive
new GeminiLive(options: GeminiLiveOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Google API key |
model | string | Yes | Model name, for example 'models/gemini-3.8-live-extended-thinking' |
thinkingLevel | GeminiThinkingLevel | No | Reasoning budget ('low', 'medium', or 'high'), supported only by 'models/gemini-3.8-live-extended-thinking' |
url | string | No | WebSocket URL |
instructions | string | No | System instructions for the model |
voice | string | No | Voice name, for example 'Aoede' or 'Charon' |
affectiveDialog | boolean | No | Enable affective dialog |
proactiveAudio | boolean | No | Enable proactive audio |
transcribeAgent | boolean | No | Transcribe agent speech |
transcribeUser | boolean | No | Transcribe user speech |
httpOptions | Record<string, unknown> | No | HTTP options |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message played when the model call fails |
inputModalities | string[] | No | Input modalities |
outputModalities | string[] | No | Output modalities |
messages | Record<string, unknown>[] | No | Conversation messages for short-term memory |
additionalParams | Record<string, unknown> | No | Additional parameters |
turnDetection | MllmTurnDetectionConfig | No | MLLM turn detection configuration; overrides top-level turnDetection |
VertexAI
new VertexAI(options: VertexAIOptions)| Option | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model name, for example 'gemini-live-2.5-flash-preview-native-audio-09-2025' |
projectId | string | Yes | Google Cloud project ID |
location | string | Yes | Google Cloud location or region |
adcCredentialsString | string | Yes | Application Default Credentials JSON string |
url | string | No | WebSocket URL |
instructions | string | No | System instructions for the model |
voice | string | No | Voice name, for example 'Aoede' or 'Charon' |
affectiveDialog | boolean | No | Enable affective dialog |
proactiveAudio | boolean | No | Enable proactive audio |
transcribeAgent | boolean | No | Transcribe agent speech |
transcribeUser | boolean | No | Transcribe user speech |
httpOptions | Record<string, unknown> | No | HTTP options |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message played when the model call fails |
inputModalities | string[] | No | Input modalities |
outputModalities | string[] | No | Output modalities |
messages | Record<string, unknown>[] | No | Conversation messages for short-term memory |
additionalParams | Record<string, unknown> | No | Additional parameters |
turnDetection | MllmTurnDetectionConfig | No | MLLM turn detection configuration; overrides top-level turnDetection |
XaiGrok
new XaiGrok(options: XaiGrokOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | xAI API key |
url | string | No | WebSocket URL (defaults to xAI Realtime API) |
voice | string | No | Voice identifier, for example 'eve' |
language | string | No | Language code |
sampleRate | number | No | Audio sample rate in Hz |
greetingMessage | string | No | Agent greeting message |
failureMessage | string | No | Message played when the model call fails |
inputModalities | string[] | No | Input modalities |
outputModalities | string[] | No | Output modalities |
messages | Record<string, unknown>[] | No | Conversation messages for short-term memory |
params | Record<string, unknown> | No | Additional xAI parameters |
turnDetection | MllmTurnDetectionConfig | No | MLLM turn detection configuration; overrides top-level turnDetection |
OpenAIGPTLive
OpenAI GPT-Live MLLM vendor (mllm.vendor: "openai_gpt_live"). The constructor throws if apiKey is empty. withMllm() throws if url isn't a full ws:// or wss:// endpoint, headers isn't a JSON object string, delegation isn't client or responses, or sessionParams overrides a protected field.
Early access
OpenAI GPT-Live is available in early access and isn't intended for production traffic. The SDK automatically routes sessions that use this vendor to the preview endpoint. GPT-Live doesn't support turn detection or input audio transcription.
new OpenAIGPTLive(options: OpenAIGPTLiveOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | OpenAI API key |
model | string | No | Model name. Defaults to gpt-live-1 |
voice | string | No | Output voice. Provider default: marin |
prompt | string | No | Session instructions that define the assistant's behavior |
greeting | string | No | Greeting the agent speaks when a user joins. Serialized as greeting_message |
failureMessage | string | No | Message played when the model call fails |
url | string | No | Full ws:// or wss:// endpoint. Defaults to wss://api.openai.com/v1/live/sessions |
baseUrl | string | No | Host used when url is omitted. Defaults to wss://api.openai.com |
path | string | No | WebSocket path used when url is omitted. Defaults to /v1/live/sessions |
inputModalities | string[] | No | Input modalities |
outputModalities | string[] | No | Output modalities |
messages | Record<string, unknown>[] | No | Conversation history passed to the model as context |
mcpServers | McpServersItem[] | No | MCP servers whose tools GPT-Live can call. Requires withTools(true) |
toolEnabled | boolean | No | Advertise the agent's tools to GPT-Live. Provider default: false |
delegation | 'client' | 'responses' | No | Tool delegation mode: client or responses. Provider default: responses. Can't be changed during the session |
responsesModel | string | No | Model used for delegated tool calls |
interruptOnUserTurn | boolean | No | Interrupt playback when the user speaks. Provider default: false |
outputIdleEndMs | number | No | Agent silence boundary in milliseconds. Provider default: 600. 0 disables inference |
inputIdleEndMs | number | No | User silence boundary in milliseconds. Provider default: 1500 |
outputSilencePeak | number | No | Speech amplitude threshold on the 16-bit scale. Provider default: 50 |
outputSampleRate | number | No | Output PCM sample rate in Hz. Provider default: 24000 |
outputBufferMs | number | No | Initial audio cushion in milliseconds. Provider default: 0. A negative value disables pacing |
inputBatchMs | number | No | Microphone audio batching interval in milliseconds |
alphaSelector | string | No | OpenAI-Alpha selector for preview contracts. Omitted when not set |
headers | string | No | Extra provider request headers, as a JSON object string |
sessionParams | Record<string, unknown> | No | Additional session fields. Can't override model, delegation, audio, instructions, or input |
params | Record<string, unknown> | No | Additional provider parameters. Explicit options take precedence |
instructions | string | No | Deprecated. Use prompt instead |
inputAudioTranscription and turnDetection are deprecated and accepted only for compatibility. GPT-Live doesn't support them: setting inputAudioTranscription throws, and the SDK ignores turnDetection and logs a warning.
Avatar vendors
Use with withAvatar(). Some avatar vendors require a specific TTS sample rate, enforced at compile time and runtime.
HeyGenAvatar
Deprecated. Use LiveAvatarAvatar for new integrations. HeyGenAvatar still works and serializes vendor: "heygen". Requires TTS at 24,000 Hz.
new HeyGenAvatar(options: HeyGenAvatarOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | HeyGen API key |
quality | 'low' | 'medium' | 'high' | Yes | Video quality: 360p, 480p, or 720p |
agoraUid | string | Yes | RTC UID for the avatar stream |
agoraToken | string | No | RTC token for avatar authentication |
avatarId | string | No | HeyGen avatar ID |
disableIdleTimeout | boolean | No | Disable idle timeout. Default: false |
activityIdleTimeout | number | No | Idle timeout in seconds. Default: 120 |
enable | boolean | No | Enable or disable the avatar. Default: true |
AkoolAvatar
Requires TTS at 16,000 Hz.
new AkoolAvatar(options: AkoolAvatarOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Akool API key |
avatarId | string | No | Akool avatar ID |
enable | boolean | No | Enable or disable the avatar. Default: true |
LiveAvatarAvatar
Requires TTS at 24,000 Hz.
new LiveAvatarAvatar(options: LiveAvatarAvatarOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | LiveAvatar API key |
quality | 'low' | 'medium' | 'high' | Yes | Video quality |
agoraUid | string | Yes | RTC UID for the avatar stream |
agoraToken | string | No | Avatar token override |
avatarId | string | No | Avatar ID |
disableIdleTimeout | boolean | No | Disable idle timeout |
activityIdleTimeout | number | No | Idle timeout in seconds |
enable | boolean | No | Enable or disable the avatar. Default: true |
AnamAvatar
new AnamAvatar(options: AnamAvatarOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Anam API key |
avatarId | string | No | Anam avatar ID |
enable | boolean | No | Enable or disable the avatar. Default: true |
GenericAvatar
Generic avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start().
new GenericAvatar(options: GenericAvatarOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Custom avatar provider API key |
apiBaseUrl | string | Yes | Avatar provider API base URL |
avatarId | string | Yes | Avatar ID |
agoraUid | string | Yes | RTC UID for the avatar stream |
agoraAppId | string | No | Agora App ID override |
agoraChannel | string | No | Agora channel override |
agoraToken | string | No | Avatar token override |
enable | boolean | No | Enable or disable the avatar. Default: true |
additionalParams | Record<string, unknown> | No | Additional vendor parameters. Explicit options take precedence over matching keys. |
Tavus
Tavus avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start(). Serializes vendor: "generic". For the endpoint and a full example, see Tavus.
new Tavus(options: TavusOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Tavus API key |
apiBaseUrl | string | Yes | Tavus API base URL |
avatarId | string | Yes | Tavus avatar ID |
agoraUid | string | Yes | RTC UID for the avatar stream |
agoraAppId | string | No | Agora App ID override |
agoraChannel | string | No | Agora channel override |
agoraToken | string | No | Avatar token override |
enable | boolean | No | Enable or disable the avatar. Default: true |
additionalParams | Record<string, unknown> | No | Additional vendor parameters. Explicit options take precedence over matching keys. |
Protoface
Protoface avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start(). Serializes vendor: "generic". For the endpoint and a full example, see Protoface.
new Protoface(options: ProtofaceOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | Protoface API key |
apiBaseUrl | string | Yes | Protoface API base URL |
avatarId | string | Yes | Protoface avatar ID |
agoraUid | string | Yes | RTC UID for the avatar stream |
agoraAppId | string | No | Agora App ID override |
agoraChannel | string | No | Agora channel override |
agoraToken | string | No | Avatar token override |
enable | boolean | No | Enable or disable the avatar. Default: true |
additionalParams | Record<string, unknown> | No | Additional vendor parameters. Explicit options take precedence over matching keys. |
LemonSlice
LemonSlice avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start(). Serializes vendor: "generic". For the endpoint and a full example, see LemonSlice.
new LemonSlice(options: LemonSliceOptions)| Option | Type | Required | Description |
|---|---|---|---|
apiKey | string | Yes | LemonSlice API key |
apiBaseUrl | string | Yes | LemonSlice API base URL |
avatarId | string | Yes | Always lemonslice |
agoraUid | string | Yes | RTC UID for the avatar stream |
agoraAppId | string | No | Agora App ID override |
agoraChannel | string | No | Agora channel override |
agoraToken | string | No | Avatar token override |
enable | boolean | No | Enable or disable the avatar. Default: true |
additionalParams | Record<string, unknown> | No | Additional vendor parameters. Explicit options take precedence over matching keys. |
Token utilities
Helper functions and classes for generating and managing tokens. Use these when you need control over token lifetime, or when generating tokens outside of a session.
import { generateConvoAIToken, ExpiresIn } from 'agora-agents';generateConvoAIToken(options)
Generates a Conversational AI token combining RTC and RTM privileges. This is the same token the SDK generates automatically in app-credentials mode. Use this when you need a token outside of a session, or when passing a pre-built token to SessionOptions.token.
generateConvoAIToken(options: GenerateConvoAITokenOptions): string| Option | Type | Required | Description |
|---|---|---|---|
appId | string | Yes | Agora App ID |
appCertificate | string | Yes | Agora App Certificate |
channelName | string | Yes | The channel the token grants access to |
uid | number | Yes | Numeric RTC UID the token is issued for |
tokenExpire | number | No | Token lifetime in seconds. Default: 86400 |
privilegeExpire | number | No | Privilege lifetime in seconds. Default: same as tokenExpire |
Returns: string — the generated token.
Throws: if appId or appCertificate isn't exactly 32 characters.
const token = generateConvoAIToken({
appId: 'your-app-id',
appCertificate: 'your-app-certificate',
channelName: 'support-room-123',
uid: 1,
tokenExpire: ExpiresIn.hours(12),
});ExpiresIn
Helper for specifying token lifetimes. Use with SessionOptions.expiresIn or generateConvoAIToken. Values are validated and capped at the Agora maximum of 86400 seconds (24 hours).
| Value or method | Returns | Description |
|---|---|---|
ExpiresIn.DAY | 86400 | 24 hours — the Agora maximum and default |
ExpiresIn.MAX | 86400 | The Agora maximum |
ExpiresIn.HOUR | 3600 | One hour |
ExpiresIn.seconds(n) | number | n seconds, validated and capped at 24 h |
ExpiresIn.hours(n) | number | n hours in seconds. Throws if n ≤ 0, caps at 24 h with a warning |
ExpiresIn.minutes(n) | number | n minutes in seconds. Throws if n ≤ 0, caps at 24 h with a warning |
Types and enums
Shared types and enums used across AgoraClient, Agent, AgentSession, and vendor classes.
Area
Region used for API routing. Pass to AgoraClient via the area option.
import { Area } from 'agora-agents';| Value | Region |
|---|---|
Area.US | United States |
Area.EU | Europe |
Area.AP | Asia-Pacific |
Area.CN | China mainland |
AgoraAuthMode
The resolved authentication mode on an AgoraClient instance. Read via client.authMode.
type AgoraAuthMode = "app-credentials" | "token" | "basic";| Value | Description |
|---|---|
"app-credentials" | App ID and App Certificate provided. SDK auto-generates tokens |
"token" | Pre-built authToken provided |
"basic" | customerId and customerSecret provided |
AgentSessionEvent
Union type of all valid event names for session.on() and session.off().
type AgentSessionEvent = "started" | "stopped" | "error";AgentSessionEventHandler
Generic handler type for session event callbacks.
type AgentSessionEventHandler<T = unknown> = (data: T) => void;| Event | T |
|---|---|
"started" | { agentId: string } |
"stopped" | { agentId: string } |
"error" | unknown (the thrown error) |
SpeakPriority
Controls how the agent handles a say() call relative to its current activity.
type SpeakPriority = "INTERRUPT" | "APPEND" | "IGNORE";| Value | Description |
|---|---|
"INTERRUPT" | Agent immediately stops current speech and delivers the message |
"APPEND" | Message is queued and delivered after current speech ends |
"IGNORE" | Message is discarded if the agent is currently speaking |
The SpeakPriorityInterrupt, SpeakPriorityAppend, and SpeakPriorityIgnore constants are also exported.
AgoraError
Thrown when the API returns a 4xx or 5xx response. Catch this to inspect the status code and response body.
import { AgoraError } from 'agora-agents';
try {
const agentId = await session.start();
} catch (err) {
if (err instanceof AgoraError) {
console.error('Status:', err.statusCode);
console.error('Message:', err.message);
console.error('Body:', err.body);
}
}| Property | Type | Description |
|---|---|---|
statusCode | number | undefined | HTTP status code returned by the API |
message | string | Error message, including the status code and body |
body | unknown | Raw response body from the API, if any |
rawResponse | RawResponse | undefined | The raw HTTP response |
