NVIDIA's AI Models Master Najdi and Hijazi Dialects: Reducing Speech Recognition Errors by More Than Half

2026-10-04T06:09:01.668Z

NVIDIA advances AI models to understand Saudi dialects, reducing speech recognition errors significantly, paving the way for more accurate voice applications.

NVIDIA has made significant progress in enhancing its AI models' ability to understand Saudi dialects, following the training of its speech recognition model Nemotron 3.5 ASR on audio data in the Najdi and Hijazi dialects.

The company explained that the customization process relied on approximately 133.7 hours of audio recordings from the SADA dataset, aiming to improve the model's ability to handle local linguistic and dialectical variations. (NVIDIA Developer)

According to the results published by NVIDIA, the word recognition error rate for the Najdi and Hijazi dialects decreased from 55.05% before training to 29.96% after training, while the character-level error rate dropped from 31.63% to 12.18%.

The training process took approximately 4.5 hours using two graphics processing units, involving more than 12,000 training steps, with updates to the model layers to enhance its understanding of phonetic and dialectical features.

This development comes at a time when the applications of voice AI are increasing in customer service, digital assistants, and converting interviews and meetings into text, alongside its growing use in media, broadcasting, and content creation.

Improving the understanding of local dialects paves the way for more user-friendly voice applications for Saudi users, particularly in automated transcription services, voice assistants, and search within audio and visual content.