TypeScript

Updated

Full API reference for the Agora Agent TypeScript SDK — AgoraClient, Agent, AgentSession, and vendor classes.

Full API reference for the Agora Conversational AI TypeScript SDK.

AgoraClient

AgoraClient extends the Fern-generated base client with domain pool support for regional URL cycling and three authentication modes. Pass appId and appCertificate only for the recommended app-credentials mode. The SDK mints fresh REST tokens per request and generates RTC join tokens at session start.

import { AgoraClient, Area } from 'agora-agents';

Constructor

const client = new AgoraClient(options: AgoraClient.Options);

The authentication mode is resolved automatically from the options you provide.

OptionTypeRequiredDescription
areaAreaYesRegion for API routing (Area.US, Area.EU, Area.AP, Area.CN)
appIdstringYesAgora App ID
appCertificatestringYesAgora App Certificate. Keep this secret and never expose it client-side
customerIdstringNoCustomer ID for Basic Auth
customerSecretstringNoCustomer Secret for Basic Auth
authTokenstringNoRaw Agora token. The SDK sends Authorization: agora token=<authToken>
timeoutInSecondsnumberNoDefault request timeout in seconds
maxRetriesnumberNoMaximum retry attempts
fetchtypeof fetchNoCustom fetch implementation for unsupported runtimes

Authentication mode is resolved from the options you provide:

Options providedResolved authMode
customerId + customerSecret"basic"
authToken"token"
Neither"app-credentials"

Passing customerId without customerSecret throws an error.

See Authentication for details on each mode.

Properties

The following read-only properties are available on any AgoraClient instance.

PropertyTypeDescription
appIdstringThe Agora App ID
appCertificatestringThe Agora App Certificate
authModeAgoraAuthModeThe resolved authentication mode
areaAreaThe configured region
poolPoolThe underlying domain pool instance used for regional routing

Methods

The following methods are available in addition to the Fern-generated sub-client methods.

nextRegion()

Cycles to the next region prefix in the domain pool. Call this after a request failure to try a different regional endpoint.

client.nextRegion();

selectBestDomain(signal?)

Runs a DNS check to select the best domain suffix. Calls made within 30 seconds of the last successful selection are no-ops.

await client.selectBestDomain();
ParameterTypeDescription
signalAbortSignalOptional abort signal to cancel the DNS check

stopAgent(agentId)

Stops an agent by ID without an AgentSession reference, for example, from an end-call handler. The method treats a 404 response as success.

stopAgent(agentId: string): Promise<void>

getCurrentURL()

Returns the full API URL currently in use as a string.

const url = client.getCurrentURL();
// Example: 'https://api-us-west-1.agora.io/api/conversational-ai-agent'

Sub-clients

AgoraClient exposes Fern-generated sub-clients for direct REST API access. You typically do not need these when using the agentkit layer.

PropertyDescription
client.agentsStart, stop, update, speak, interrupt, get history, list agents
client.agentManagementInject instructions into a running agent (think)
client.telephonyTelephony operations
client.phoneNumbersPhone number management

For full method signatures and request parameters, see the REST API reference.

Agent

Agent is an immutable configuration object. Each builder method returns a new Agent instance — the original is never modified. Define one Agent at startup and call createSession() on it for each user conversation.

import { Agent } from 'agora-agents';

Constructor

new Agent(options: AgentOptions)

client is required. All other options are optional. Use the builder methods to set vendor configuration after construction.

OptionTypeDefaultDescription
clientAgoraClient—Required. Agora client used by createSession()
pipelineIdstringundefinedPublished AI Studio pipeline ID used as the base configuration. SessionOptions.pipelineId overrides it
instructionsstringundefinedDeprecated. Set systemMessages on the LLM vendor instead
greetingstringundefinedDeprecated. Set greetingMessage on the LLM or MLLM vendor instead
failureMessagestringundefinedDeprecated. Set failureMessage on the LLM or MLLM vendor instead
maxHistorynumberundefinedDeprecated. Set maxHistory on the LLM vendor instead
greetingConfigsLlmGreetingConfigsundefinedDeprecated. Set greetingConfigs on the LLM vendor instead
turnDetectionTurnDetectionConfigundefinedVoice activity detection settings
interruptionInterruptionConfigundefinedUnified interruption control settings
salSalConfigundefinedSelective Attention Locking configuration
avatarAvatarConfigundefinedAvatar configuration
advancedFeaturesAdvancedFeaturesundefinedEnable MLLM mode, AI-VAD, and other advanced features
parametersSessionParamsInputundefinedSession parameters including silence and farewell config
geofenceGeofenceConfigundefinedRegional access restriction
labelsLabelsundefinedCustom key-value labels returned in notification callbacks
rtcRtcConfigundefinedRTC media encryption
fillerWordsFillerWordsConfigundefinedFiller words played while waiting for the LLM response

Builder methods

All builder methods return a new Agent instance. The original is never modified.

withLlm(vendor)

Sets the LLM vendor. Pass an instance of OpenAI, AzureOpenAI, Anthropic, Gemini, or any other LLM vendor.

withLlm(vendor: LlmVendor): Agent<TTSSampleRate, TArea>

withTts(vendor)

Sets the TTS vendor. The sample rate type is captured and tracked for avatar compatibility.

withTts<SR extends number>(vendor: TtsVendor<SR>): Agent<SR, TArea>

withStt(vendor)

Sets the STT vendor. Pass an instance of any STT vendor class.

withStt(vendor: SttVendor): Agent<TTSSampleRate, TArea>

withMllm(vendor)

Sets the MLLM vendor for multimodal mode. Pass OpenAIRealtime, AzureOpenAIRealtime, GeminiLive, VertexAI, XaiGrok, or OpenAIGPTLive. Calling withMllm() automatically sets mllm.enable = true. MLLM mode does not require withTts() / withLlm() / withStt().

Avatars are only supported with the cascading ASR + LLM + TTS pipeline. If you combine withMllm() with withAvatar(), the SDK throws an error when toProperties() or session.start() is called.

withMllm(vendor: GlobalMllmVendor | CNMllmVendor): Agent<TTSSampleRate, TArea>

withAvatar(vendor)

Sets the avatar vendor. The this constraint enforces at compile time that the agent's TTS sample rate matches the avatar's required rate.

Requires the cascading ASR + LLM + TTS pipeline. If you combine withAvatar() with withMllm(), the SDK throws an error when toProperties() or session.start() is called.

withAvatar<RequiredSR extends number>(
  this: Agent<RequiredSR>,
  vendor: AvatarVendor<RequiredSR>
): Agent<RequiredSR, TArea>

withTurnDetection(config)

Configures cascading-flow turn detection. Pass { config: { start_of_speech, end_of_speech } } for SOS/EOS detection. Use withInterruption() for interruption behavior and MLLM vendor turnDetection for MLLM turn detection.

withTurnDetection(config: TurnDetectionConfig): Agent<TTSSampleRate, TArea>

withInterruption(config)

Configures unified interruption behavior using the top-level interruption object. Use this for start_of_speech and keywords interruption modes.

withInterruption(config: InterruptionConfig): Agent<TTSSampleRate, TArea>

withInstructions(text)

Overrides the LLM system prompt on a new Agent instance. Deprecated. Set systemMessages on the LLM vendor instead.

withInstructions(instructions: string): Agent<TTSSampleRate, TArea>

withGreeting(text)

Overrides the greeting message on a new Agent instance. Deprecated. Set greetingMessage on the LLM or MLLM vendor instead.

withGreeting(greeting: string): Agent<TTSSampleRate, TArea>

Other builder methods

The following methods follow the same pattern — each returns a new Agent instance with the updated configuration.

MethodParameter typeDescription
withSal(config)SalConfigSet Selective Attention Locking configuration
withAdvancedFeatures(features)AdvancedFeaturesSet advanced features
withTools(enabled = true)booleanEnable or disable MCP tool and custom tool invocation
withParameters(parameters)SessionParamsInputSet session parameters
withAudioScenario(audioScenario)ParametersAudioScenarioSet parameters.audio_scenario. Use the exported AudioScenario constants for discoverability, for example agent.withAudioScenario(AudioScenario.Aiserver)
withFailureMessage(message)stringDeprecated. Set failureMessage on the LLM or MLLM vendor instead
withMaxHistory(n)numberDeprecated. Set maxHistory on the LLM vendor instead. For Azure OpenAI Realtime MLLM, set maxHistory on the MLLM vendor
withGreetingConfigs(configs)LlmGreetingConfigsDeprecated. Set greetingConfigs on the LLM vendor instead
withGeofence(geofence)GeofenceConfigSet geofence configuration
withLabels(labels)LabelsSet custom labels
withRtc(rtc)RtcConfigSet RTC configuration
withFillerWords(fillerWords)FillerWordsConfigSet filler words configuration

createSession(options)

Creates an AgentSession for a channel using the agent's client. Doesn't start the agent. Call session.start() to join the channel.

createSession(options: SessionOptions): AgentSession

SessionOptions fields:

OptionTypeRequiredDescription
channelstringYesChannel name to join
agentUidstringYesThe agent's RTC UID. Must be a numeric string when token is omitted
remoteUidsstring[]YesRemote user UIDs the agent listens and responds to
namestringNoAgent instance name sent to the Agora API. Defaults to agent-{timestamp}
tokenstringNoPre-built RTC+RTM token. Omit to auto-generate from app credentials
expiresInnumberNoToken lifetime in seconds. Only applies when the token is auto-generated. Valid range: 1–86400. Use ExpiresIn helpers for clarity
idleTimeoutnumberNoSeconds before the agent auto-exits when no audio is detected. 0 disables the timeout
enableStringUidbooleanNoUse string UIDs instead of numeric UIDs
presetPresetInputNoAdvanced project-specific presets. Use only when Agora provides a specific preset ID for your project. Most applications should not set this field
pipelineIdstringNoPublished AI Studio pipeline ID to use as the base configuration
debugbooleanNoLog API requests to the console
warn(message: string) => voidNoCustom warning logger; pass a no-op to silence warnings

preset is session-scoped because the underlying Agora start/join API applies presets per session, not per reusable Agent definition.

When you omit credentials for supported reseller-backed vendor models, AgentKit infers the matching session preset automatically:

  • Deepgram STT: nova-2, nova-3
  • OpenAI LLM: gpt-4o-mini, gpt-4.1-mini, gpt-5-nano, gpt-5-mini
  • OpenAI TTS: tts-1
  • MiniMax TTS: speech-2.6-turbo, speech-2.8-turbo

If you provide your own vendor API key for those same models, AgentKit keeps the request in BYOK mode and does not infer a preset.

Properties

Read-only properties available on any Agent instance.

PropertyTypeDescription
pipelineIdstring | undefinedAI Studio pipeline ID
greetingConfigsLlmGreetingConfigs | undefinedGreeting playback configuration
instructionsstring | undefinedLLM system prompt
greetingstring | undefinedGreeting message
failureMessagestring | undefinedMessage spoken when LLM fails
maxHistorynumber | undefinedMaximum conversation history length
llmLlmConfig | undefinedLLM configuration
ttsTtsConfig | undefinedTTS configuration
sttSttConfig | undefinedSTT configuration
mllmMllmConfig | undefinedMLLM configuration
avatarAvatarConfig | undefinedAvatar configuration
turnDetectionTurnDetectionConfig | undefinedTurn detection configuration
interruptionInterruptionConfig | undefinedInterruption configuration
salSalConfig | undefinedSAL configuration
advancedFeaturesAdvancedFeatures | undefinedAdvanced features
parametersSessionParams | undefinedSession parameters
geofenceGeofenceConfig | undefinedGeofence configuration
labelsLabels | undefinedCustom labels
rtcRtcConfig | undefinedRTC configuration
fillerWordsFillerWordsConfig | undefinedFiller words configuration
configAgentOptionsFull read-only configuration snapshot

toProperties(opts)

Low-level method to convert the agent configuration to the Fern request format. Used internally by AgentSession.start(). You typically do not need to call this directly unless building custom request bodies.

toProperties(opts): StartAgentsRequest.Properties

In cascading mode, throws if TTS or LLM isn't set. If you don't call withStt(), ASR defaults to ARES (fengming for Area.CN). turnDetection.language defaults to en-US and is also sent as asr.language.

Type aliases

Public aliases over Fern-generated types include LlmConfig, SttConfig, AsrConfig (= SttConfig), MllmConfig, AvatarConfig, session/conversation types, and think types (ThinkOnListeningAction, etc.).

Think value constants: ThinkOnListeningActionInject, ThinkOnListeningActionInterrupt, ThinkOnListeningActionIgnore, ThinkOnThinkingActionInterrupt, ThinkOnThinkingActionIgnore, ThinkOnSpeakingActionInterrupt, ThinkOnSpeakingActionIgnore.

The think action options also accept the string append, which has no convenience constant. With append, the instruction doesn't interrupt the current interaction. The agent waits for the current user turn, LLM inference, or TTS playback to finish, then appends the instruction to the context as a separate user message and starts a new turn.

FillerWordsConfig supports static and generated filler words. Generated mode requires static fallback phrases, while generated_config and prompt are optional. Agora hosts the generation service, so you can't configure its model endpoint, API key, or model parameters. For the field structure, see Configure generated filler words.

FillerWordsContentGeneratedConfig sets the conversation context for generated filler words. Use context_message_limit (1 to 6, defaults to 1) to set how many of the most recent conversation messages to use, counting the current turn's message, and history_character_limit (0 to 10000, defaults to 1000) to cap the combined characters of the earlier messages. For details, see Talking while waiting.

Public aliases also include FillerWordsContentGeneratedConfig and the custom tool types LlmTool, LlmToolFunction, LlmToolServer, and LlmToolExecution. llm.tools is independent of MCP server configuration, but both require withTools(true) to enable tool calling.

AgentSession

AgentSession manages the full lifecycle of a running agent. Create sessions using agent.createSession(); do not call the constructor directly.

import { AgentSession } from 'agora-agents';

State machine

A session progresses through the following states:

idle ──► starting ──► running ──► stopping ──► stopped
            │                         │
            ▼                         ▼
          error                     error
TransitionTrigger
idle → startingstart() called
starting → runningAPI responds with agent ID
starting → errorAPI request fails
running → stoppingstop() called
stopping → stoppedAPI confirms agent stopped
stopping → errorStop request fails with a non-404 error

start() can also be called from stopped or error state to restart the session. say(), interrupt(), update(), and think() reject on failure but don't change the session state.

Methods

The following methods are available on an AgentSession instance.

start()

Starts the agent session. Generates tokens if not provided, sends the start request, and returns the agent ID. Resolves explicit preset values and also infers reseller presets from supported vendor configs when credentials are omitted.

start(): Promise<string>
  • Transitions: idle / stopped / error → starting → running
  • Throws if called in starting, running, or stopping state
  • Throws if avatar config is invalid (wrong TTS sample rate)
  • Throws if no token is provided and the client has no appCertificate
  • Throws if agentUid (or avatar agoraUid) isn't numeric when tokens are auto-generated
  • Throws if appId or appCertificate isn't exactly 32 characters when tokens are auto-generated
  • Throws if MLLM is enabled together with an enabled avatar — avatars are only supported with the cascading ASR + LLM + TTS pipeline
  • Applies explicit preset values when provided and sends Agora-managed configuration when supported vendor credentials are omitted
  • Fills generic avatar agora_appid and agora_channel from the session when omitted
  • Generates avatar agora_token for HeyGenAvatar, LiveAvatarAvatar, GenericAvatar, Tavus, Protoface, and LemonSlice when agoraToken is omitted and the client has an appCertificate. Other vendors (AkoolAvatar, AnamAvatar) never receive an auto-generated token.

stop()

Stops the agent session and removes the agent from the channel. If the agent has already stopped — for example due to idle timeout — resolves silently rather than throwing a 404 error.

stop(): Promise<void>
  • Transitions: running → stopping → stopped
  • Throws if called outside running state

say(text, options?)

Instructs the agent to speak the given text.

say(text: string, options?: SayOptions): Promise<void>
ParameterTypeRequiredDescription
textstringYesThe text for the agent to speak
options.prioritySpeakPriorityNoMessage priority
options.interruptablebooleanNoWhether this message can be interrupted by the user
  • Only valid in running state

interrupt()

Interrupts the agent's current speech.

interrupt(): Promise<void>
  • Only valid in running state

think(text, options?)

Injects a custom text instruction into the running agent.

think(text: string, options?: ThinkOptions): Promise<ThinkResponse>
ParameterTypeRequiredDescription
textstringYesThe instruction text
options.on_listening_actionThinkOnListeningActionNoAction when the agent is listening: inject, interrupt, ignore, or append
options.on_thinking_actionThinkOnThinkingActionNoAction when the agent is thinking: interrupt, ignore, or append
options.on_speaking_actionThinkOnSpeakingActionNoAction when the agent is speaking: interrupt, ignore, or append
options.interruptablebooleanNoWhether the user can interrupt the resulting response
options.metadataRecord<string, string>NoCustom key-value metadata
  • Only valid in running state

update(config)

Updates the agent configuration mid-session without restarting. Accepts a partial configuration object in REST API format.

update(config: AgentConfigUpdate): Promise<void>
  • Only valid in running state

getHistory()

Fetches the conversation history for this session. Requires a valid agent ID — start() must have been called successfully.

getHistory(): Promise<ConversationHistory>

getTurns(options?)

Fetches turn-by-turn analytics for this session, including start/end events and latency metrics. Requires a valid agent ID and start() must have been called successfully.

getTurns(options?: GetTurnsOptions): Promise<ConversationTurns>
  • options.page_index: page number, starting from 1
  • options.page_size: number of turns per page

getAllTurns(options?)

Fetches all turn analytics pages and merges the turns array. Requires a valid agent ID and start() must have been called successfully.

getAllTurns(options?: Omit<GetTurnsOptions, "page_index">): Promise<ConversationTurns>
  • Requires a valid agentId
  • For very long sessions, prefer processing pages with getTurns() to avoid holding all turns in memory

getInfo()

Fetches current agent metadata from the API. Requires a valid agent ID.

getInfo(): Promise<SessionInfo>

on(event, handler)

Subscribes to a session event. Register handlers before calling start() to avoid missing the started event.

on<T>(event: AgentSessionEvent, handler: AgentSessionEventHandler<T>): void

off(event, handler)

Unsubscribes a previously registered event handler.

off<T>(event: AgentSessionEvent, handler: AgentSessionEventHandler<T>): void

Events

The session emits the following events. See AgentSessionEvent and AgentSessionEventHandler for type details.

EventPayload typeDescription
"started"{ agentId: string }Agent successfully joined the channel
"stopped"{ agentId: string }Agent left the channel
"error"unknownThe error thrown during start() or stop()

Properties

The following read-only properties are available on any AgentSession instance.

PropertyTypeDescription
status"idle" | "starting" | "running" | "stopping" | "stopped" | "error"Current session state. One of "idle", "starting", "running", "stopping", "stopped", "error"
idstring | nullAgent ID, populated after start() resolves
agentAgentThe agent configuration this session was created from
appIdstringThe Agora App ID for this session
rawAgentsClientDirect access to the Fern-generated AgentsClient for advanced operations
rawAgentManagementAgentManagementClientDirect access to the Fern-generated agent management client

Using session.raw

Access the generated REST client to call endpoints not yet wrapped:

await session.raw.getTurns({
  appid: session.appId,
  agentId: session.id!,
});

You must pass appid and agentId manually when using raw methods.

Presets and BYOK

Prefer configuring vendors on the Agent builder. When you omit credentials for supported Agora-managed models, AgentKit sends the matching Agora-managed configuration at session start.

preset is an advanced session option for project-specific settings, not for selecting Agora-managed models. Most applications should use the builder instead.

  • Omit vendor credentials on the builder for supported Agora-managed models.
  • Provide vendor API keys when you want BYOK.
  • Pass preset on agent.createSession(...) only when you need to access specific project-specific settings.

Supported Agora-managed models:

  • Deepgram STT: nova-2, nova-3
  • OpenAI LLM: gpt-4o-mini, gpt-4.1-mini, gpt-5-nano, gpt-5-mini
  • OpenAI TTS: tts-1
  • MiniMax TTS: speech-2.6-turbo, speech-2.8-turbo

Vendors

All vendor classes are imported from agora-agents. Pass vendor instances to the Agent builder methods.

LLM vendors

Use with withLlm().

OpenAI

new OpenAI(options: OpenAIOptions)
OptionTypeRequiredDescription
apiKeystringUsuallyOpenAI API key. Optional for Agora-managed preset models
modelstringYesModel name, for example 'gpt-4o-mini'
urlstringConditionalAPI endpoint URL. Required when apiKey is set (BYOK); default https://api.openai.com/v1/chat/completions
maxHistorynumberNoMaximum conversation history to cache
temperaturenumberNoSampling temperature (0.0–2.0)
topPnumberNoNucleus sampling (0.0–1.0)
maxTokensnumberNoMaximum tokens to generate
systemMessagesRecord<string, unknown>[]NoAdditional system messages
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage spoken when the LLM call fails
inputModalitiesstring[]NoInput modalities. Default: ["text"]
outputModalitiesstring[]NoOutput modalities
paramsRecord<string, unknown>NoAdditional LLM parameters passed to the model
headersRecord<string, string>NoCustom HTTP headers forwarded to the LLM provider
vendorstringNoVendor override
mcpServersRecord<string, unknown>[]NoMCP server connections
toolsLlmTool[]NoSynchronous custom tool definitions the LLM can call. Requires withTools(true)
greetingConfigsLlmGreetingConfigsNoGreeting playback configuration
templateVariablesRecord<string, string>NoTemplate variables for messages. Custom tools can reference them with {{template_variables.<name>}}

apiKey is optional for the following reseller preset models: gpt-4o-mini, gpt-4.1-mini, gpt-5-nano, gpt-5-mini. If apiKey is omitted for one of those models, AgentKit infers the matching session preset. This no-key branch is only available with the default OpenAI endpoint and without a custom vendor hint. If apiKey is provided, AgentKit uses standard BYOK behavior instead.

AzureOpenAI

new AzureOpenAI(options: AzureOpenAIOptions)
OptionTypeRequiredDescription
apiKeystringYesAzure OpenAI API key
modelstringYesModel or deployment name
resourceNamestringConditionalAzure resource name. Required unless endpoint is set
endpointstringConditionalFull Azure base URL. Takes precedence over resourceName; required unless resourceName is set
deploymentNamestringYesDeployment name in Azure
apiVersionstringNoAzure API version. Default: '2024-08-01-preview'
maxHistorynumberNoMaximum conversation history to cache
temperaturenumberNoSampling temperature (0.0–2.0)
topPnumberNoNucleus sampling (0.0–1.0)
maxTokensnumberNoMaximum tokens to generate
systemMessagesRecord<string, unknown>[]NoAdditional system messages
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage spoken when the LLM call fails
inputModalitiesstring[]NoInput modalities. Default: ["text"]
outputModalitiesstring[]NoOutput modalities
paramsRecord<string, unknown>NoAdditional LLM parameters
headersRecord<string, string>NoCustom HTTP headers forwarded to the LLM provider
vendorstringNoVendor override. Defaults to azure
mcpServersRecord<string, unknown>[]NoMCP server connections
toolsLlmTool[]NoSynchronous custom tool definitions the LLM can call. Requires withTools(true)
greetingConfigsLlmGreetingConfigsNoGreeting playback configuration
templateVariablesRecord<string, string>NoTemplate variables for messages. Custom tools can reference them with {{template_variables.<name>}}

Anthropic

new Anthropic(options: AnthropicOptions)
OptionTypeRequiredDescription
apiKeystringYesAnthropic API key
modelstringYesModel name
urlstringYesAnthropic messages endpoint URL
maxTokensnumberYesMaximum tokens to generate
headersRecord<string, string>YesRequest headers, including anthropic-version
maxHistorynumberNoMaximum conversation history to cache
temperaturenumberNoSampling temperature (0.0–1.0)
topPnumberNoNucleus sampling (0.0–1.0)
systemMessagesRecord<string, unknown>[]NoAdditional system messages
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage spoken when the LLM call fails
inputModalitiesstring[]NoInput modalities. Default: ["text"]
outputModalitiesstring[]NoOutput modalities
paramsRecord<string, unknown>NoAdditional LLM parameters
vendorstringNoVendor override
mcpServersRecord<string, unknown>[]NoMCP server connections
toolsLlmTool[]NoSynchronous custom tool definitions the LLM can call. Requires withTools(true)
greetingConfigsLlmGreetingConfigsNoGreeting playback configuration
templateVariablesRecord<string, string>NoTemplate variables for messages. Custom tools can reference them with {{template_variables.<name>}}

Gemini

new Gemini(options: GeminiOptions)
OptionTypeRequiredDescription
apiKeystringYesGoogle API key
modelstringYesModel name, for example 'gemini-pro'
urlstringNoAPI endpoint URL. Default: https://generativelanguage.googleapis.com/v1beta/models/{model}:streamGenerateContent?alt=sse&key={apiKey}, with the API key in the URL
maxHistorynumberNoMaximum conversation history to cache
temperaturenumberNoSampling temperature (0.0–2.0)
topPnumberNoNucleus sampling (0.0–1.0)
topKnumberNoTop-k sampling
maxOutputTokensnumberNoMaximum output tokens to generate
systemMessagesRecord<string, unknown>[]NoAdditional system messages
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage spoken when the LLM call fails
inputModalitiesstring[]NoInput modalities. Default: ["text"]
outputModalitiesstring[]NoOutput modalities
paramsRecord<string, unknown>NoAdditional LLM parameters
headersRecord<string, string>NoCustom HTTP headers forwarded to the LLM provider
vendorstringNoVendor override
mcpServersRecord<string, unknown>[]NoMCP server connections
toolsLlmTool[]NoSynchronous custom tool definitions the LLM can call. Requires withTools(true)
greetingConfigsLlmGreetingConfigsNoGreeting playback configuration
templateVariablesRecord<string, string>NoTemplate variables for messages. Custom tools can reference them with {{template_variables.<name>}}

Other LLM vendors

ClassProviderKey options
GroqGroqapiKey, model, url
VertexAILLMGoogle Vertex AIapiKey, model, projectId, location, url?
AmazonBedrockAmazon BedrockaccessKey, secretKey, region, model
DifyDifyapiKey, url, model, user?, conversationId?
CustomLLMOpenAI-compatible LLMapiKey, model, url

Groq and CustomLLM share OpenAI's option shape; VertexAILLM mirrors Gemini. All accept common LLM fields, such as systemMessages, greetingMessage, failureMessage, maxHistory, params, headers, and tools.

tools declares custom tools the LLM can choose to call. Each tool includes a model-visible function definition and the server configuration for a synchronous GET or POST request. withTools(true) applies to both tools and MCP server configuration. For more information, see Call custom tools.

TTS vendors

Use with withTts(). The sampleRate option determines avatar compatibility — see withAvatar().

ElevenLabsTTS

new ElevenLabsTTS<SR extends ElevenLabsSampleRate>(options: ElevenLabsTTSOptions<SR>)
OptionTypeRequiredDescription
keystringYesElevenLabs API key
modelIdstringYesModel ID, for example 'eleven_flash_v2_5'
voiceIdstringYesVoice ID
baseUrlstringYesWebSocket base URL
sampleRate16000 | 22050 | 24000 | 44100NoAudio sample rate in Hz
optimizeStreamingLatencynumberNoLatency optimization level (0–4)
stabilitynumberNoVoice stability (0.0–1.0)
similarityBoostnumberNoVoice similarity boost (0.0–1.0)
stylenumberNoVoice style exaggeration (0.0–1.0)
useSpeakerBoostbooleanNoEnable speaker boost
skipPatternsnumber[]NoSkip patterns for bracketed content

MicrosoftTTS

new MicrosoftTTS<SR extends MicrosoftSampleRate>(options: MicrosoftTTSOptions<SR>)
OptionTypeRequiredDescription
keystringYesAzure Speech API key
regionstringYesAzure region, for example 'eastus'
voiceNamestringYesVoice name, for example 'en-US-JennyNeural'
sampleRate16000 | 24000 | 48000NoAudio sample rate in Hz
speednumberNoSpeaking rate multiplier
volumenumberNoAudio volume
skipPatternsnumber[]NoSkip patterns for bracketed content

OpenAITTS

Fixed at 24,000 Hz — no configurable sample rate.

new OpenAITTS(options: OpenAITTSOptions)
OptionTypeRequiredDescription
apiKeystringUsuallyOpenAI API key
voicestringYesVoice name: 'alloy', 'echo', 'fable', 'onyx', 'nova', or 'shimmer'
modelstringNoModel name. Required (with apiKey and baseUrl) for BYOK
baseUrlstringNoEndpoint URL. Required (with apiKey and model) for BYOK
instructionsstringNoCustom voice instructions
speednumberNoSpeech speed multiplier
skipPatternsnumber[]NoSkip patterns for bracketed content

apiKey is optional only for the reseller-backed tts-1 preset path. If omitted with model: 'tts-1' or no explicit model, AgentKit infers openai_tts_1. If provided, the request stays in BYOK mode.

CartesiaTTS

new CartesiaTTS<SR extends CartesiaSampleRate>(options: CartesiaTTSOptions<SR>)
OptionTypeRequiredDescription
apiKeystringYesCartesia API key
voiceIdstringYesVoice ID (serialized as {"mode": "id", "id": "..."})
modelIdstringYesModel ID
baseUrlstringNoWebSocket URL
languagestringNoTarget language
sampleRate8000 | 16000 | 22050 | 24000 | 44100 | 48000NoAudio sample rate in Hz
skipPatternsnumber[]NoSkip patterns for bracketed content

Other TTS vendors

ClassKey parameters
GoogleTTSkey, voiceName, languageCode?, sampleRate?
AmazonTTSaccessKey, secretKey, region, voiceId, engine
DeepgramTTSapiKey, model, baseUrl?, sampleRate?, additionalParams?
HumeAITTSkey, voiceId, provider, configId?, baseUrl?, speed?, trailingSilence?
RimeTTSkey, speaker, modelId, baseUrl?, credentialMode?. In managed mode (CredentialMode.Managed), only modelId and baseUrl are required
FishAudioTTSkey, referenceId, backend
MiniMaxTTSkey?, groupId?, model, voiceId?, url?
MurfTTSkey, voiceId?, baseUrl?, locale?, rate?, pitch?, model?, sampleRate?
SarvamTTSkey, speaker, targetLanguageCode, pitch?, pace?, loudness?, sampleRate?
GradiumTTSapiKey, url?, modelName?, voiceId?, sampleRate?, additionalParams?
MistralTTSapiKey, model?, voice?, additionalParams?
TypecastTTSapiKey, voiceId, model, additionalParams?
XAiTTSapiKey, language, voiceId?, sampleRate?, additionalParams?
SmallestAITTSapiKey, url?, model?, voiceId?, sampleRate?, speed?, language?, additionalParams?

SmallestAITTS also accepts numberPronunciationLanguage?, mathNotation?, pronunciationDicts?, sessionId?, and requestId?. Voice identifiers are model-specific, so voiceId must belong to the model set in model. For details, see Smallest AI.

For MiniMaxTTS, key is optional only for reseller-backed models: speech-2.6-turbo, speech-2.8-turbo. If key is omitted for one of those models, AgentKit infers the matching session preset. In that preset-backed path, groupId, voiceId, and url are optional overrides rather than required fields. If key is provided, AgentKit uses BYOK, and groupId, voiceId, and url are required.

GenericTTS

Custom OpenAI-compatible HTTP TTS. url is required and must be an absolute HTTP or HTTPS address — the constructor throws if url is missing, badly formatted, or uses a non-HTTP(S) scheme such as ws: or wss:. A valid URL is serialized as tts.vendor = "generic_http".

new GenericTTS(options: GenericTTSOptions)
OptionTypeRequiredDescription
urlstringYesThe HTTP(S) endpoint of your custom TTS service
headersRecord<string, string>NoCustom HTTP headers to forward to the TTS service. Omitted from the request if not set
apiKeystringNoThe API key used to authenticate with the TTS service
modelstringNoThe TTS model name
voicestringNoThe voice name
speednumberNoThe speech rate
sampleRatenumberNoThe sample rate, in Hz, of the output audio. If your TTS service doesn't support multiple sample rates, make sure the returned audio's sample rate matches this value
responseFormat"pcm"NoThe output audio format. Conversational AI Engine currently supports pcm
instructionstringNoInstructions for voice style, emotion, or other playback directives
additionalParamsRecord<string, unknown>NoAdditional parameters passed through to the TTS service. Explicit fields with the same name take precedence
skipPatternsnumber[]NoSkip patterns for bracketed content

STT vendors

Use with withStt().

DeepgramSTT

new DeepgramSTT(options?: DeepgramSTTOptions)

Provide apiKey, or omit it and set model to nova-2 or nova-3 to use Agora-managed credentials. The constructor throws otherwise.

OptionTypeRequiredDescription
apiKeystringNoDeepgram API key. Optional for nova-2 and nova-3 reseller preset usage
modelstringConditionalModel name, for example 'nova-2' or 'enhanced'. Required when apiKey is omitted
languagestringNoLanguage code, for example 'en-US'
keytermstringNoBoost specialized terms and brands
smartFormatbooleanNoEnable smart formatting
punctuationbooleanNoEnable punctuation
additionalParamsRecord<string, unknown>NoAdditional vendor parameters

For nova-2 and nova-3, omit apiKey to use Agora-managed credentials. For all other Deepgram models, apiKey is required.

Other STT vendors

ClassKey parameters
SpeechmaticsSTTkey, language, model?, uri?. apiKey is a deprecated alias of key
MicrosoftSTTkey, region, language
OpenAISTTapiKey, model?, language?, prompt?, inputAudioTranscription?
GoogleSTTprojectId, location, adcCredentialsString, language, model?
AmazonSTTaccessKey, secretKey, region, language
AssemblyAISTTapiKey, language, ws_url?
AresSTTkeywords?, additionalParams?
GeminiSTTapiKey, model?, language?, languageHints?, customVocabulary?, sampleRate? (default 16000), wordTimestamp?, mode? (SMART or VERBATIM), diarization?, additionalParams?
XAiSTTapiKey, language?, baseUrl?, sampleRate?, additionalParams?
SarvamSTTapiKey, language, model?
SmallestAISTTapiKey, language?, url?, sampleRate?, encoding?, keywords?, additionalParams?

SmallestAISTT also accepts the boolean options wordTimestamps?, sentenceTimestamps?, diarize?, vadEvents?, endpointing?, format?, finalizeOnWords?, punctuate?, capitalize?, itnNormalize?, fullTranscript?, redactPii?, and redactPci?, plus eouTimeoutMs? and maxWords?. Boolean options are serialized as the strings "true" or "false" for the Smallest AI wire protocol. For details, see Smallest AI.

For OpenAISTT, the serialized configuration requires a transcription prompt and language — provide them through the prompt and language options or within inputAudioTranscription. model defaults to gpt-4o-mini-transcribe.

MLLM vendors

Use with withMllm() for multimodal end-to-end audio processing without separate STT or TTS steps. Calling withMllm() automatically enables the MLLM module (sets mllm.enable = true); the older advancedFeatures.enable_mllm flag is deprecated.

OpenAIRealtime

new OpenAIRealtime(options: OpenAIRealtimeOptions)
OptionTypeRequiredDescription
apiKeystringYesOpenAI API key
modelstringNoModel name, for example 'gpt-4o-realtime-preview'
voicestringNoVoice identifier
instructionsstringNoSystem instructions
inputAudioTranscriptionRecord<string, unknown>NoAudio transcription settings
urlstringNoWebSocket URL
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage played when the model call fails
inputModalitiesstring[]NoInput modalities, for example ['audio']
outputModalitiesstring[]NoOutput modalities, for example ['text', 'audio']
messagesRecord<string, unknown>[]NoConversation messages for short-term memory
paramsRecord<string, unknown>NoAdditional MLLM parameters
turnDetectionMllmTurnDetectionConfigNoMLLM turn detection configuration; overrides top-level turnDetection

AzureOpenAIRealtime

new AzureOpenAIRealtime(options: AzureOpenAIRealtimeOptions)
OptionTypeRequiredDescription
apiKeystringYesAzure OpenAI API key
urlstringYesAzure OpenAI Realtime WebSocket URL
turnDetectionMllmTurnDetectionConfigYesMLLM turn detection configuration; overrides top-level turnDetection
modelstringNoAzure OpenAI Realtime model or deployment name
voicestringNoVoice identifier
instructionsstringNoSystem instructions
maxHistorynumberNoNumber of conversation history messages to cache
greetingMessagestringNoAgent greeting message
outputModalitiesstring[]NoOutput modalities, for example ['text', 'audio']
messagesRecord<string, unknown>[]NoConversation messages for short-term memory
paramsAzureOpenAIRealtimeParamsNoAdditional Azure OpenAI parameters

GeminiLive

new GeminiLive(options: GeminiLiveOptions)
OptionTypeRequiredDescription
apiKeystringYesGoogle API key
modelstringYesModel name, for example 'models/gemini-3.8-live-extended-thinking'
thinkingLevelGeminiThinkingLevelNoReasoning budget ('low', 'medium', or 'high'), supported only by 'models/gemini-3.8-live-extended-thinking'
urlstringNoWebSocket URL
instructionsstringNoSystem instructions for the model
voicestringNoVoice name, for example 'Aoede' or 'Charon'
affectiveDialogbooleanNoEnable affective dialog
proactiveAudiobooleanNoEnable proactive audio
transcribeAgentbooleanNoTranscribe agent speech
transcribeUserbooleanNoTranscribe user speech
httpOptionsRecord<string, unknown>NoHTTP options
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage played when the model call fails
inputModalitiesstring[]NoInput modalities
outputModalitiesstring[]NoOutput modalities
messagesRecord<string, unknown>[]NoConversation messages for short-term memory
additionalParamsRecord<string, unknown>NoAdditional parameters
turnDetectionMllmTurnDetectionConfigNoMLLM turn detection configuration; overrides top-level turnDetection

VertexAI

new VertexAI(options: VertexAIOptions)
OptionTypeRequiredDescription
modelstringYesModel name, for example 'gemini-live-2.5-flash-preview-native-audio-09-2025'
projectIdstringYesGoogle Cloud project ID
locationstringYesGoogle Cloud location or region
adcCredentialsStringstringYesApplication Default Credentials JSON string
urlstringNoWebSocket URL
instructionsstringNoSystem instructions for the model
voicestringNoVoice name, for example 'Aoede' or 'Charon'
affectiveDialogbooleanNoEnable affective dialog
proactiveAudiobooleanNoEnable proactive audio
transcribeAgentbooleanNoTranscribe agent speech
transcribeUserbooleanNoTranscribe user speech
httpOptionsRecord<string, unknown>NoHTTP options
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage played when the model call fails
inputModalitiesstring[]NoInput modalities
outputModalitiesstring[]NoOutput modalities
messagesRecord<string, unknown>[]NoConversation messages for short-term memory
additionalParamsRecord<string, unknown>NoAdditional parameters
turnDetectionMllmTurnDetectionConfigNoMLLM turn detection configuration; overrides top-level turnDetection

XaiGrok

new XaiGrok(options: XaiGrokOptions)
OptionTypeRequiredDescription
apiKeystringYesxAI API key
urlstringNoWebSocket URL (defaults to xAI Realtime API)
voicestringNoVoice identifier, for example 'eve'
languagestringNoLanguage code
sampleRatenumberNoAudio sample rate in Hz
greetingMessagestringNoAgent greeting message
failureMessagestringNoMessage played when the model call fails
inputModalitiesstring[]NoInput modalities
outputModalitiesstring[]NoOutput modalities
messagesRecord<string, unknown>[]NoConversation messages for short-term memory
paramsRecord<string, unknown>NoAdditional xAI parameters
turnDetectionMllmTurnDetectionConfigNoMLLM turn detection configuration; overrides top-level turnDetection

OpenAIGPTLive

OpenAI GPT-Live MLLM vendor (mllm.vendor: "openai_gpt_live"). The constructor throws if apiKey is empty. withMllm() throws if url isn't a full ws:// or wss:// endpoint, headers isn't a JSON object string, delegation isn't client or responses, or sessionParams overrides a protected field.

Early access

OpenAI GPT-Live is available in early access and isn't intended for production traffic. The SDK automatically routes sessions that use this vendor to the preview endpoint. GPT-Live doesn't support turn detection or input audio transcription.

new OpenAIGPTLive(options: OpenAIGPTLiveOptions)
OptionTypeRequiredDescription
apiKeystringYesOpenAI API key
modelstringNoModel name. Defaults to gpt-live-1
voicestringNoOutput voice. Provider default: marin
promptstringNoSession instructions that define the assistant's behavior
greetingstringNoGreeting the agent speaks when a user joins. Serialized as greeting_message
failureMessagestringNoMessage played when the model call fails
urlstringNoFull ws:// or wss:// endpoint. Defaults to wss://api.openai.com/v1/live/sessions
baseUrlstringNoHost used when url is omitted. Defaults to wss://api.openai.com
pathstringNoWebSocket path used when url is omitted. Defaults to /v1/live/sessions
inputModalitiesstring[]NoInput modalities
outputModalitiesstring[]NoOutput modalities
messagesRecord<string, unknown>[]NoConversation history passed to the model as context
mcpServersMcpServersItem[]NoMCP servers whose tools GPT-Live can call. Requires withTools(true)
toolEnabledbooleanNoAdvertise the agent's tools to GPT-Live. Provider default: false
delegation'client' | 'responses'NoTool delegation mode: client or responses. Provider default: responses. Can't be changed during the session
responsesModelstringNoModel used for delegated tool calls
interruptOnUserTurnbooleanNoInterrupt playback when the user speaks. Provider default: false
outputIdleEndMsnumberNoAgent silence boundary in milliseconds. Provider default: 600. 0 disables inference
inputIdleEndMsnumberNoUser silence boundary in milliseconds. Provider default: 1500
outputSilencePeaknumberNoSpeech amplitude threshold on the 16-bit scale. Provider default: 50
outputSampleRatenumberNoOutput PCM sample rate in Hz. Provider default: 24000
outputBufferMsnumberNoInitial audio cushion in milliseconds. Provider default: 0. A negative value disables pacing
inputBatchMsnumberNoMicrophone audio batching interval in milliseconds
alphaSelectorstringNoOpenAI-Alpha selector for preview contracts. Omitted when not set
headersstringNoExtra provider request headers, as a JSON object string
sessionParamsRecord<string, unknown>NoAdditional session fields. Can't override model, delegation, audio, instructions, or input
paramsRecord<string, unknown>NoAdditional provider parameters. Explicit options take precedence
instructionsstringNoDeprecated. Use prompt instead

inputAudioTranscription and turnDetection are deprecated and accepted only for compatibility. GPT-Live doesn't support them: setting inputAudioTranscription throws, and the SDK ignores turnDetection and logs a warning.

Avatar vendors

Use with withAvatar(). Some avatar vendors require a specific TTS sample rate, enforced at compile time and runtime.

HeyGenAvatar

Deprecated. Use LiveAvatarAvatar for new integrations. HeyGenAvatar still works and serializes vendor: "heygen". Requires TTS at 24,000 Hz.

new HeyGenAvatar(options: HeyGenAvatarOptions)
OptionTypeRequiredDescription
apiKeystringYesHeyGen API key
quality'low' | 'medium' | 'high'YesVideo quality: 360p, 480p, or 720p
agoraUidstringYesRTC UID for the avatar stream
agoraTokenstringNoRTC token for avatar authentication
avatarIdstringNoHeyGen avatar ID
disableIdleTimeoutbooleanNoDisable idle timeout. Default: false
activityIdleTimeoutnumberNoIdle timeout in seconds. Default: 120
enablebooleanNoEnable or disable the avatar. Default: true

AkoolAvatar

Requires TTS at 16,000 Hz.

new AkoolAvatar(options: AkoolAvatarOptions)
OptionTypeRequiredDescription
apiKeystringYesAkool API key
avatarIdstringNoAkool avatar ID
enablebooleanNoEnable or disable the avatar. Default: true

LiveAvatarAvatar

Requires TTS at 24,000 Hz.

new LiveAvatarAvatar(options: LiveAvatarAvatarOptions)
OptionTypeRequiredDescription
apiKeystringYesLiveAvatar API key
quality'low' | 'medium' | 'high'YesVideo quality
agoraUidstringYesRTC UID for the avatar stream
agoraTokenstringNoAvatar token override
avatarIdstringNoAvatar ID
disableIdleTimeoutbooleanNoDisable idle timeout
activityIdleTimeoutnumberNoIdle timeout in seconds
enablebooleanNoEnable or disable the avatar. Default: true

AnamAvatar

new AnamAvatar(options: AnamAvatarOptions)
OptionTypeRequiredDescription
apiKeystringYesAnam API key
avatarIdstringNoAnam avatar ID
enablebooleanNoEnable or disable the avatar. Default: true

GenericAvatar

Generic avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start().

new GenericAvatar(options: GenericAvatarOptions)
OptionTypeRequiredDescription
apiKeystringYesCustom avatar provider API key
apiBaseUrlstringYesAvatar provider API base URL
avatarIdstringYesAvatar ID
agoraUidstringYesRTC UID for the avatar stream
agoraAppIdstringNoAgora App ID override
agoraChannelstringNoAgora channel override
agoraTokenstringNoAvatar token override
enablebooleanNoEnable or disable the avatar. Default: true
additionalParamsRecord<string, unknown>NoAdditional vendor parameters. Explicit options take precedence over matching keys.

Tavus

Tavus avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start(). Serializes vendor: "generic". For the endpoint and a full example, see Tavus.

new Tavus(options: TavusOptions)
OptionTypeRequiredDescription
apiKeystringYesTavus API key
apiBaseUrlstringYesTavus API base URL
avatarIdstringYesTavus avatar ID
agoraUidstringYesRTC UID for the avatar stream
agoraAppIdstringNoAgora App ID override
agoraChannelstringNoAgora channel override
agoraTokenstringNoAvatar token override
enablebooleanNoEnable or disable the avatar. Default: true
additionalParamsRecord<string, unknown>NoAdditional vendor parameters. Explicit options take precedence over matching keys.

Protoface

Protoface avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start(). Serializes vendor: "generic". For the endpoint and a full example, see Protoface.

new Protoface(options: ProtofaceOptions)
OptionTypeRequiredDescription
apiKeystringYesProtoface API key
apiBaseUrlstringYesProtoface API base URL
avatarIdstringYesProtoface avatar ID
agoraUidstringYesRTC UID for the avatar stream
agoraAppIdstringNoAgora App ID override
agoraChannelstringNoAgora channel override
agoraTokenstringNoAvatar token override
enablebooleanNoEnable or disable the avatar. Default: true
additionalParamsRecord<string, unknown>NoAdditional vendor parameters. Explicit options take precedence over matching keys.

LemonSlice

LemonSlice avatars can omit agoraAppId, agoraChannel, and agoraToken. AgentKit fills them from the session at start(). Serializes vendor: "generic". For the endpoint and a full example, see LemonSlice.

new LemonSlice(options: LemonSliceOptions)
OptionTypeRequiredDescription
apiKeystringYesLemonSlice API key
apiBaseUrlstringYesLemonSlice API base URL
avatarIdstringYesAlways lemonslice
agoraUidstringYesRTC UID for the avatar stream
agoraAppIdstringNoAgora App ID override
agoraChannelstringNoAgora channel override
agoraTokenstringNoAvatar token override
enablebooleanNoEnable or disable the avatar. Default: true
additionalParamsRecord<string, unknown>NoAdditional vendor parameters. Explicit options take precedence over matching keys.

Token utilities

Helper functions and classes for generating and managing tokens. Use these when you need control over token lifetime, or when generating tokens outside of a session.

import { generateConvoAIToken, ExpiresIn } from 'agora-agents';

generateConvoAIToken(options)

Generates a Conversational AI token combining RTC and RTM privileges. This is the same token the SDK generates automatically in app-credentials mode. Use this when you need a token outside of a session, or when passing a pre-built token to SessionOptions.token.

generateConvoAIToken(options: GenerateConvoAITokenOptions): string
OptionTypeRequiredDescription
appIdstringYesAgora App ID
appCertificatestringYesAgora App Certificate
channelNamestringYesThe channel the token grants access to
uidnumberYesNumeric RTC UID the token is issued for
tokenExpirenumberNoToken lifetime in seconds. Default: 86400
privilegeExpirenumberNoPrivilege lifetime in seconds. Default: same as tokenExpire

Returns: string — the generated token.

Throws: if appId or appCertificate isn't exactly 32 characters.

const token = generateConvoAIToken({
  appId: 'your-app-id',
  appCertificate: 'your-app-certificate',
  channelName: 'support-room-123',
  uid: 1,
  tokenExpire: ExpiresIn.hours(12),
});

ExpiresIn

Helper for specifying token lifetimes. Use with SessionOptions.expiresIn or generateConvoAIToken. Values are validated and capped at the Agora maximum of 86400 seconds (24 hours).

Value or methodReturnsDescription
ExpiresIn.DAY8640024 hours — the Agora maximum and default
ExpiresIn.MAX86400The Agora maximum
ExpiresIn.HOUR3600One hour
ExpiresIn.seconds(n)numbern seconds, validated and capped at 24 h
ExpiresIn.hours(n)numbern hours in seconds. Throws if n ≤ 0, caps at 24 h with a warning
ExpiresIn.minutes(n)numbern minutes in seconds. Throws if n ≤ 0, caps at 24 h with a warning

Types and enums

Shared types and enums used across AgoraClient, Agent, AgentSession, and vendor classes.

Area

Region used for API routing. Pass to AgoraClient via the area option.

import { Area } from 'agora-agents';
ValueRegion
Area.USUnited States
Area.EUEurope
Area.APAsia-Pacific
Area.CNChina mainland

AgoraAuthMode

The resolved authentication mode on an AgoraClient instance. Read via client.authMode.

type AgoraAuthMode = "app-credentials" | "token" | "basic";
ValueDescription
"app-credentials"App ID and App Certificate provided. SDK auto-generates tokens
"token"Pre-built authToken provided
"basic"customerId and customerSecret provided

AgentSessionEvent

Union type of all valid event names for session.on() and session.off().

type AgentSessionEvent = "started" | "stopped" | "error";

AgentSessionEventHandler

Generic handler type for session event callbacks.

type AgentSessionEventHandler<T = unknown> = (data: T) => void;
EventT
"started"{ agentId: string }
"stopped"{ agentId: string }
"error"unknown (the thrown error)

SpeakPriority

Controls how the agent handles a say() call relative to its current activity.

type SpeakPriority = "INTERRUPT" | "APPEND" | "IGNORE";
ValueDescription
"INTERRUPT"Agent immediately stops current speech and delivers the message
"APPEND"Message is queued and delivered after current speech ends
"IGNORE"Message is discarded if the agent is currently speaking

The SpeakPriorityInterrupt, SpeakPriorityAppend, and SpeakPriorityIgnore constants are also exported.

AgoraError

Thrown when the API returns a 4xx or 5xx response. Catch this to inspect the status code and response body.

import { AgoraError } from 'agora-agents';

try {
  const agentId = await session.start();
} catch (err) {
  if (err instanceof AgoraError) {
    console.error('Status:', err.statusCode);
    console.error('Message:', err.message);
    console.error('Body:', err.body);
  }
}
PropertyTypeDescription
statusCodenumber | undefinedHTTP status code returned by the API
messagestringError message, including the status code and body
bodyunknownRaw response body from the API, if any
rawResponseRawResponse | undefinedThe raw HTTP response