Speech and Audio Analytics

Feature Grid Preview

Speech-to-Text (ASR)

Optimised for regional ASEAN languages and accents, supporting post-processing and near-real-time transcription. Supports challenging environments such as RF recorded speech.

Code-Switching Speech-to-Text

Designed for Singapore's multilingual context, supporting natural switching between English, Mandarin, Malay, and Tamil. Available for post-processing and near-real-time use.

Speaker Verification & Identification

Verify or identify speakers from voice recordings. Available for post-processing.

Speaker Diarization

Identify and separate individual speakers within multi-speaker audio. Available for post-processing.

Speech Translation

Translate spoken content across supported languages, with post-processing and near-real-time capabilities.

Sound Event Detection

Detects and classifies a wide range of non-speech sounds. Also identifies ambient context such as indoor, outdoor, and environmental noise.

Flexible Deployment of Speech Translation

Android-ready speech translation capability for integration into mobile and IoT applications.

Localisation of Speech Models

KLASS develops localised TTS models that accurately transcribe Singapore street and place names, such as Tanjong Pagar and Joo Chiat, addressing limitations of generic pretrained models through locally adapted speech data.