Specifications and compatibility
Updated
Review the performance profile, hardware capabilities, platform compatibility, and audio processing features of Convo AI Device Kit R1.
Performance specifications
The Convo AI Device Kit delivers ultra-low latency performance with robust audio processing capabilities across a global network.
-
Latency
- Conversation latency: As low as 650ms
- Interruption response: As low as 340ms
- Global network end-to-end latency: Median as low as 76ms
-
Audio processing
- Noise suppression: Filters 95% of environmental noise
- Packet loss resistance: Up to 80% packet loss tolerance
-
Coverage
- Network coverage: 200+ countries and regions
- Language support: 35+ languages
Hardware specifications
The R1 kit is based on the RiseLink BK7258 chipset and includes open-source hardware and software resources.
-
Audio capabilities
- Dual-microphone array with local AEC (Acoustic Echo Cancellation) algorithm
- Precise audio capture with echo interference elimination
-
Visual and sensor capabilities
- Integrated camera for visual recognition
- Gyroscope for motion sensing and gesture control
-
Power management
- Battery power support
- Deep sleep mode for mobile scenarios
-
Connectivity
- Bluetooth provisioning
- Wi-Fi 6 support for one-click cloud service connection
-
Display
- Dual-screen collaborative display
-
Interaction modes
- Multi-channel input: voice, touchscreen, and gyroscope (gesture/tilt control)
- Custom wake word support
- Real-time continuous conversation
- LLM real-time visual reasoning
Platform compatibility
The Convo AI Device Kit supports various mainstream communication standards and chipsets.
-
Supported chip manufacturers
- RiseLink
- Espressif
- Unisoc
- Ingenic
- Rockchip (RK)
- Sigmastar
-
Communication standards
- Wi-Fi
- LTE Category 1
-
Additional support
- Image Signal Processor (ISP) chips
For specific supported chip models and compatibility details, contact technical support.
Advanced audio algorithms
The Convo AI Device Kit employs specialized algorithms to ensure accurate voice recognition and natural conversation flow in challenging environments.
AI noise reduction
Filters 95% of environmental noise, enabling accurate recognition even in challenging environments like coffee shops and train stations, preventing interaction errors.
BHVS and voiceprint algorithms
- Background Human Voice Separation (BHVS): Filters background voices in multi-person conversation scenarios
- Voiceprint recognition: Locks onto the primary speaker in multi-person conversations
Graceful interruption algorithm
Enables AI to detect user interruption intent in real-time and precisely determine when to speak and when to stop, restoring natural conversation rhythm.
Weak network resistance
Maintains uninterrupted voice interaction even in challenging network conditions like subways and basements, with 80% packet loss resistance capability to maintain conversation continuity.
