Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models
Developing Automatic Speech Recognition (ASR) for morphologically rich, low resource languages such as Assamese is challenging due to insufficient annotated speech data. The pretrained Whisper model performs poorly on Assamese speech recognition tasks. This paper presents a controlled, fine tuned Whisper based Assamese ASR system trained on the Mozilla Common Voice 24.0 Assamese corpus. A hardware aware optimized ...