Google Gemini Live
Updated
Integrate Google Gemini Live with the Conversational AI Engine using the Gemini Developer API.
Agora has partnered with Google Gemini Live to bring frontier multimodal large language model capabilities to your app. Gemini Live provides real-time audio processing, enabling natural voice conversations without separate ASR and TTS components. This page explains how to integrate Gemini Live through the Gemini Developer API.
Info
Enabling MLLM automatically disables ASR, LLM, and TTS since the MLLM handles end-to-end voice processing directly.
Sample configuration
The following examples show how to configure Google Gemini Live MLLM (Gemini Developer API) when starting a conversational AI agent.
from agora_agent import Agent
from agora_agent.agentkit.vendors import GeminiLive
# client is your configured Agora client
agent = (
Agent(client)
.with_mllm(GeminiLive(
api_key='your-google-api-key',
model='gemini-3.1-flash-live-preview',
voice='Charon',
instructions='You are a friendly assistant.',
transcribe_agent=True,
transcribe_user=True,
))
)import { Agent, GeminiLive } from 'agora-agents';
// client is your configured Agora client
const agent = new Agent({ client })
.withMllm(new GeminiLive({
apiKey: 'your-google-api-key',
model: 'gemini-3.1-flash-live-preview',
voice: 'Charon',
instructions: 'You are a friendly assistant.',
transcribeAgent: true,
transcribeUser: true,
}));import "github.com/AgoraIO/agora-agents-go/v2/agentkit/vendors"
// client is your configured Agora client
agent := agentkit.NewAgent(client).WithMllm(
vendors.NewGeminiLive(vendors.GeminiLiveOptions{
APIKey: "your-google-api-key",
Model: "gemini-3.1-flash-live-preview",
Voice: "Charon",
Instructions: "You are a friendly assistant.",
TranscribeAgent: agora.Bool(true),
TranscribeUser: agora.Bool(true),
}),
)Use the following mllm configuration in your request:
"mllm": {
"enable": true,
"api_key": "<GOOGLE_GEMINI_API_KEY>",
"messages": [
{
"role": "user",
"content": "<HISTORY_CONTENT>"
}
],
"params": {
"model": "gemini-3.1-flash-live-preview",
"instructions": "You are a friendly assistant.",
"voice": "Charon",
"affective_dialog": false,
"proactive_audio": false,
"transcribe_agent": true,
"transcribe_user": true,
"http_options": {
"api_version": "v1beta"
}
},
"turn_detection": {
// see details below
},
"input_modalities": [
"audio"
],
"output_modalities": [
"audio"
],
"greeting_message": "Hi, how can I assist you today?",
"failure_message": "Sorry, I encountered an issue. Please try again.",
"vendor": "gemini"
}Turn detection
To set up turn detection, add a turn_detection block inside the mllm object when you Start a conversational AI agent.
Info
When mllm.turn_detection is defined, the top-level turn_detection object has no effect.
The following examples show the supported configurations for Google Gemini Live.
-
Server VAD
"turn_detection": { "mode": "server_vad", "server_vad_config": { "prefix_padding_ms": 800, "silence_duration_ms": 640, "start_of_speech_sensitivity": "START_SENSITIVITY_HIGH", "end_of_speech_sensitivity": "END_SENSITIVITY_HIGH" } } -
Agora VAD
"turn_detection": { "mode": "agora_vad", "agora_vad_config": { "interrupt_duration_ms": 160, "prefix_padding_ms": 800, "silence_duration_ms": 640, "threshold": 0.5 } }
Key parameters
enablebooleanEnables the MLLM module. Replaces the deprecated advanced_features.enable_mllm.
api_keystringThe Google Gemini API key used to authenticate requests. You can generate an API key in Google AI Studio.
messagesarray[object]An array of conversation history items passed to the model as context. Each item represents a single message in the conversation history.
rolestringThe role of the message author. For example, system or user.
contentstringThe content of the message.
paramsobjectConfiguration object for the Gemini Live model.
modelstringThe Gemini Live model identifier.
instructionsstringSystem instructions that define the agent's behavior or tone.
voicestringThe voice identifier for audio output. For example, Aoede, Puck, Charon, Kore, Fenrir, Leda, Orus, or Zephyr.
affective_dialogbooleanWhether to enable affective dialog, which allows the model to adapt its tone based on the user's emotional cues.
proactive_audiobooleanWhen enabled, the model may choose not to respond if the user's input does not require a reply, such as background speech or incomplete requests.
transcribe_agentbooleanWhether to transcribe the agent's speech in real time.
transcribe_userbooleanWhether to transcribe the user's speech in real time.
http_optionsobjectHTTP request options for the Gemini Live API.
api_versionstringThe API version to use. For example, v1beta.
turn_detectionobjectTurn detection configuration for the MLLM module. For a full list of turn_detection parameters, see mllm.turn_detection.
modestring- Possible values
agora_vadserver_vad
agora_vad: Agora VAD-based detection.server_vad: Vendor-side VAD-based detection.
agora_vad_configobjectConfiguration for Agora VAD-based turn detection. Applicable when mode is agora_vad.
interrupt_duration_msintegerMinimum duration of speech in milliseconds required to trigger an interruption.
prefix_padding_msintegerDuration of audio in milliseconds to include before the detected speech start.
silence_duration_msintegerDuration of silence in milliseconds required to determine end of speech.
thresholdnumberVAD sensitivity threshold. A higher value reduces false positives.
server_vad_configobjectConfiguration for vendor-side VAD-based turn detection. Applicable when mode is server_vad. Parameters are passed through to the vendor.
prefix_padding_msintegerDuration of audio in milliseconds to include before the detected speech start.
silence_duration_msintegerDuration of silence in milliseconds required to determine end of speech.
start_of_speech_sensitivitystring- Possible values
START_SENSITIVITY_HIGHSTART_SENSITIVITY_LOW
Sensitivity for start of speech detection.
end_of_speech_sensitivitystring- Possible values
END_SENSITIVITY_HIGHEND_SENSITIVITY_LOW
Sensitivity for end of speech detection.
input_modalitiesarray[string]- Default value
- ["audio"]
Input modalities for the MLLM.
["audio"]: Audio-only input["audio", "text"]: Accept both audio and text input
output_modalitiesarray[string]- Default value
- ["audio"]
Output modalities for the MLLM.
["audio"]: Audio-only output["text", "audio"]: Combined text and audio output
greeting_messagestringThe message the agent speaks when a user joins the channel.
failure_messagestringThe message the agent speaks when an error occurs.
vendorstringThe MLLM provider identifier. Set to "gemini" to use Google Gemini Live with the Gemini Developer API.
For comprehensive API reference, real-time capabilities, and detailed parameter descriptions, see the Google Gemini Live API.
