🔍 Read the full analysis: Global South Language Makes Its Debut On The Open ASR Leaderboard on ThorstenMeyerAI.com
TL;DR
The Open ASR Leaderboard has added Hindi and Indian English, making them the first Indic and Global South languages on the platform. This expands benchmarking to over half a billion speakers and introduces new metadata for bias analysis.
The Open ASR Leaderboard has officially added two new evaluation sets, Monsoon en-IN for Indian English and Monsoon hi-IN for Hindi, making them the first Indic and Global South languages to appear on the platform. This development broadens the scope of benchmarking in automatic speech recognition (ASR) from predominantly European languages to include languages spoken by over half a billion people, highlighting a significant step toward more inclusive AI evaluation.
The new datasets are part of a collaboration between Voice Arena and Hugging Face, designed to assess ASR models across diverse populations. The sets consist of recordings from a total of 4,888 speakers, with each clip capturing spontaneous, unscripted conversations in various acoustic environments. The Indian English set includes approximately 11 hours of audio split between public and private segments, sourced from over 1,400 speakers across India’s districts. The Hindi dataset comprises roughly 6.8 hours of speech from 468 speakers, with recordings collected from multiple regions, ages, genders, and device types. Each clip records 12 speaker attributes, including age, gender, occupation, income, and location, enabling detailed bias and fairness analyses.
These datasets are designed with diversity in mind, capturing speech from various devices, environments, and speech types—ranging from opinion and narration to disagreement and recall. Unlike previous datasets, which often relied on scripted or studio-recorded speech, the Monsoon sets are based on real-world, spontaneous conversations, making them more representative of natural speech patterns in the Global South. The Hindi data introduces a novel approach by providing a lattice of multiple valid spellings for each transcript span, addressing the challenge of spelling variation in Hindi and other Indic languages.
Impact of Including Hindi and Indian English on ASR Benchmarking
The inclusion of Hindi and Indian English on the Open ASR Leaderboard marks a pivotal shift toward more inclusive and representative benchmarking in speech recognition AI. Historically, the platform has focused on European languages, which limits the assessment of model performance across diverse accents, dialects, and socio-economic backgrounds. By integrating datasets that reflect the linguistic and acoustic diversity of the Global South, this development encourages the creation of more equitable ASR systems.
Furthermore, the detailed speaker metadata allows researchers and developers to analyze biases and disparities in model performance across different demographic groups, such as age, gender, region, and income. This could lead to targeted improvements, reducing disparities in speech recognition accuracy and fostering more trustworthy AI applications in multilingual, multicultural contexts.
In addition, the datasets’ design—featuring spontaneous speech and multiple spellings—pushes the development of models capable of handling real-world variability, which is crucial for deploying ASR in practical settings across the Global South. Overall, this expansion broadens the scope of benchmarking, potentially influencing the direction of future research and commercial deployment of multilingual ASR systems.
automatic speech recognition device
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of ASR Benchmarking and Language Diversity
Until now, the Open ASR Leaderboard has primarily included European languages such as English, French, and German, with datasets focused on scripted or controlled speech. This focus has limited the ability to evaluate model performance across diverse accents, dialects, and socio-economic backgrounds, which are critical for deploying ASR globally.
The platform’s recent efforts to incorporate metadata and diverse datasets aim to address these limitations by revealing biases and disparities that are often hidden in aggregate error rates. Prior research, such as the ‘Racial Disparities in Automated Speech Recognition’ and ‘Quantifying Bias in Automatic Speech Recognition’ studies, demonstrated that commercial ASR systems perform significantly worse for certain demographic groups, especially Black speakers and those with regional accents.
The addition of Hindi and Indian English datasets signifies a step toward rectifying these gaps, as these languages are spoken by over half a billion people. The datasets are designed to reflect real-world, spontaneous speech collected from various regions, ages, and devices, thus providing a more representative benchmark for evaluating models in diverse contexts. The datasets also introduce novel normalization techniques, especially for Hindi, where multiple spellings are accepted for the same transcript segment, addressing language-specific challenges.
“Our goal is to push for more equitable evaluation metrics that reveal biases and disparities in speech recognition systems across different populations.”
— Hugging Face spokesperson
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About Dataset Impact and Model Performance
Several key questions remain open. The Hindi dataset, comprising only 1.33 hours of speech, is relatively small, raising concerns about the stability and reliability of model rankings on this set. It is unclear how existing top-performing models will perform on these new datasets or whether current benchmarks are sufficiently challenging for Hindi and Indian English.
Additionally, the effectiveness of the lattice normalization approach for Hindi, compared to traditional normalization methods, has not yet been validated through published performance comparisons. Whether models trained on European languages can generalize effectively to these new datasets without significant adaptation remains uncertain.
Finally, it is not yet known whether leaderboard participants will analyze and publish disaggregated results based on the detailed speaker attributes, which is essential for understanding and addressing biases explicitly.
As an affiliate, we earn on qualifying purchases.
Future Directions for Multilingual ASR Benchmarking
The datasets are now available for public self-scoring, with private splits scheduled for release later, allowing researchers to evaluate their models’ performance on Indian English and Hindi. The next steps include baseline evaluations by existing models, which will establish reference points for future improvements.
More comprehensive analyses, including bias assessments across demographic attributes, are expected as researchers begin to publish disaggregated results. Additionally, further data collection efforts are likely to expand the datasets, increasing hours of speech and speaker diversity, especially for Hindi, to improve stability and robustness.
The inclusion of these languages may also inspire the development of new normalization and modeling techniques tailored to Indic languages, further advancing multilingual ASR capabilities. Overall, this marks a significant milestone in making speech recognition technology more globally representative and equitable.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why are Hindi and Indian English important additions to the leaderboard?
They are spoken by over half a billion people, and their inclusion helps evaluate and improve ASR systems for diverse, underrepresented populations, promoting fairness and inclusivity in AI.
What challenges are associated with the Hindi dataset?
The Hindi dataset is relatively small, with only 1.33 hours of speech, and features multiple valid spellings, which complicates model evaluation and normalization approaches.
How does the new metadata help address bias?
By recording attributes like age, gender, region, and income, researchers can analyze how models perform across different demographic groups, identifying and reducing disparities.
Will existing models perform well on these new datasets?
This remains unknown until baseline evaluations are conducted, but current models trained on European languages may require adaptation to perform effectively on Hindi and Indian English.
What is the significance of the lattice approach for Hindi transcripts?
It allows multiple correct spellings for each transcript segment, addressing language-specific spelling variations and improving recognition accuracy.
Primary source: Hugging Face · via ThorstenMeyerAI.com