Skip to main content
The Speech Enhancement SE1.2 model delivers full-fidelity speech enhancement for demanding Human ↔ Human audio. It takes input up to 16kHz and outputs ultra-fidelity 24kHz audio, going beyond noise removal to reconstruct and restore degraded speech, remove reverberation, and recover quality lost to low-bitrate codecs — for a richer, clearer, more natural listening experience. Speaker identity and accent are preserved, without the perceptible shift seen in some other Speech Enhancement models.

Key Features

Bandwidth Extension

Extends audio to ultra-fidelity 24kHz for a richer, more natural voice experience beyond standard telephony quality.

Denoising

Removes background noise and non-primary voices, including residual noise and background speech left by earlier models.

Speech Reconstruction

Reconstructs and restores degraded speech, improving intelligibility and signal clarity — even without upsampling.

De-reverberation

Removes room reverberation from the signal for a cleaner, more present voice.

Codec Restoration

Recovers quality lost to low-bitrate codec compression.

Cleaner High Frequencies

Improved reproduction of sibilants and fricatives (reduced lisping) and reduced scratchiness in very low signal-to-noise conditions.

Specifications

Model ID: SE1.2
Category: Speech Enhancement
Type: Human ↔ Human

40ms–90ms

Streaming latency

16kHz → 24kHz

Input / Output sample rate
Input constraint: input bandwidth above 8kHz is not supported, so the maximum input sampling rate is 16kHz. The model operates and outputs at 24kHz.Latency: 40ms when running noise cancellation only; up to 90ms when using the full enhancement pipeline (bandwidth extension, de-reverberation, speech reconstruction, and codec restoration). CPU footprint is comparable to SE 1.0, with only a modest increase.

Use Cases

Contact Centers

Premium voice quality for agent-customer calls where clarity and presence matter most.

Conferencing

Ultra-fidelity audio for video conferences and virtual meetings.

Telemedicine

Crystal-clear audio for doctor-patient consultations where every word counts.

Degraded & Codec-Compressed Audio

Restore intelligibility on noisy, reverberant, or low-bitrate codec-degraded audio.
Known Limitations
  • Some residual lisping can remain on speech with bandwidth below ~3kHz.
  • Some scratchiness and artifacts can remain on audio compressed with very low-bitrate codecs.

Code Example

Create an audio processor with the SE 1.2 model:
For full setup and initialization, see the Quick Start →

Next Steps

Quick Start

Get up and running with Sanas SDK in under 5 minutes.

API Reference

Full SDK documentation for classes, enums, and callbacks.

Processing Multiple Streams

Handle multiple concurrent audio streams.