Who is listening?
A person is listening
You need a Human ↔ Human model. Voice quality and naturalness matter.
A machine is listening
You need a Human ↔ Machine model. Relative Word Error Rate (RWERR) reduction and ASR accuracy matter.
Agentic Speech Enhancement Models
Speech Enhancement Models
Accent Translation Models
Language Translation
Hear samples and learn more about each model’s specifications, use cases, and code examples below.
Human ↔ Human
Speech Enhancement · Enhanced (Telco)
SE1.2 — Enhances voice quality for narrowband telephony audio and adds comfort noise. Mainly intended for Telco use cases.Speech Enhancement · Standard
SE2.1 — Restores and enhances voice quality for telephony audio. Low CPU footprint.- Latency: 120ms
- Sample rate: 16kHz → 8kHz
Speech Enhancement · with full-fidelity
SE2.2 — Full-fidelity speech enhancement with bandwidth extension to ultra-fidelity 24kHz.- Latency: 160ms
- Sample rate: 16kHz → 24kHz
Speech Enhancement · Voice Isolation (General)
VI_G_SE — Isolates intended speech by removing background noise and voices. Optimized for human listeners.- Latency: ~40ms
- Sample rate: Up to 24kHz
- Range: Primary speaker within ~1m
Accent Translation
AT5.2 — Accent Translation modifies global accents in real-time, allowing your teams to be instantly understood while preserving what makes every voice unique.- Latency: ~200ms
Language Translation
LT — Real-time language translation that preserves your speakers’ voices, tone, and intent.- Latency: ~3-5s
Human ↔ Machine
Agentic Speech Enhancement · Voice Isolation (General)
AGENTIC_VI_G_SE — Removes background noise and distant voices for complete voice isolation of the primary speaker’s audio stream.- Latency: ~100ms
- Sample rate: 16kHz
- Relative Word Error Rate Reduction (RWERR): 5–30% (average)
Agentic Speech Enhancement · Voice Isolation (Telephony)
AGENTIC_VI_GT_SE— Telephony-optimized variant of Voice Isolation for 8kHz narrowband audio.- Latency: ~100ms
- Sample rate: 8kHz
- Relative Word Error Rate Reduction (RWERR): 5–30% (average)
Agentic Speech Enhancement · Standard
AGENTIC_ST_SE — Removes background noise while preserving all human speech for multi-speaker environments.- Latency: ~100ms
- Sample rate: 16kHz
- Relative Word Error Rate Reduction (RWERR): 5–30% (average)