Speech Enhancement
Restore, reconstruct, and enrich the human voice in real time.Sanas Speech Enhancement (SE) models take degraded voice audio — noisy, reverberant, narrowband, or codec-compressed — and restore it to clear, natural, full-fidelity speech. Unlike noise cancellation, which focuses on removing unwanted sound, Speech Enhancement actively reconstructs and enriches the speech signal itself: cleaning up artifacts, filling in lost detail, and extending bandwidth for a richer, more intelligible result. All Speech Enhancement models are optimized for Human ↔ Human experiences, where a person is on the other end of the line and perceived quality matters most.
Core models
SE Voice Isolation (General)
(SE1.0) VI_G_SE Isolates intended speech by removing background noise and voices. Optimized for human listeners.Speech Enhancement · Enhanced
SE1.2 — Denoises, de-reverberates, and reconstructs degraded speech, with bandwidth extension to 24kHz.Speech Enhancement · Standard
SE2.1 Restores and enhances voice quality for telephony audio. Low CPU footprint.Speech Enhancement · Full-Fidelity
SE2.2 Full-fidelity speech enhancement with bandwidth extension to ultra-fidelity 24kHz.What Speech Enhancement does
- Denoising — removes residual background noise and leaked voices.
- Speech reconstruction — restores degraded or missing detail to improve intelligibility and clarity.
- De-reverberation — reduces room reverb and echo.
- Bandwidth extension — rebuilds high-frequency content to widen narrowband audio.
- Codec restoration — recovers quality lost to low-bitrate compression.
Models
Choose a model
SE Voice Isolation (General)
VI_G_SE (SE1.0) — Isolates intended speech by removing background noise and voices. Optimized for human listeners.- Latency: ~40ms
- Sample rate: Up to 24kHz
- Range: Primary speaker within ~1m
Speech Enhancement · Enhanced
SE1.2 — Denoises, de-reverberates, and reconstructs degraded speech, with bandwidth extension to 24kHz.- Latency: ~40ms
- Sample rate: 16kHz → 24kHz
Speech Enhancement · Standard
SE2.1 — Restores and enhances voice quality for telephony audio. Low CPU footprint.- Latency: 120ms
- Sample rate: 16kHz → 8kHz
Speech Enhancement · Full Fidelity
SE2.2 — Full-fidelity speech enhancement with bandwidth extension to ultra-fidelity 24kHz.- Latency: 160ms
- Sample rate: 16kHz → 24kHz
How to choose
- Pick SE Voice Isolation 1.0 (
VI_G_SE) when you need to isolate the intended speaker — removing background noise and other voices — for a human listener, in contexts like contact centers, conferencing, and gaming. - Pick SE Standard (
SE2.1) when your output path is narrowband telephony (8kHz) and you want the lowest CPU footprint — ideal for high-volume contact center and IVR deployments. - Pick SE Full-Fidelity (
SE2.2) when you need full-fidelity 24kHz output for premium voice experiences such as conferencing and telemedicine. - Pick SE Ehanced (
SE1.2) when input audio is not just narrowband but also noisy, reverberant, or degraded by low-bitrate codecs — Ultra+ restores maximum intelligibility and clarity on top of full 24kHz bandwidth extension.
All Speech Enhancement models accept input up to 16kHz. Input bandwidth above 8kHz is not required, and
SE2.2 and SE1.2 render output at a full 24kHz.