Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages

TL;DR

This paper identifies four structural barriers—web presence gap, token scarcity, tokenization penalty, connectivity exclusion—that hinder AI support for underrepresented languages like Bengali.

cs.CL 🔴 Advanced 2026-08-13 1 citations 93 views
Avijit Roy Proma Roy
low-resource languages NLP AI infrastructure digital divide linguistic equity

Key Findings

Methodology

The study employs multi-source data analysis and case studies, integrating large-scale corpora such as Sangraha to quantify indicators like web content share, token counts, tokenization efficiency, and internet penetration. Comparative analysis between Bengali and English reveals stark disparities: Bengali accounts for less than 0.5% of web content versus 49.5% for English; token ratios are approximately 67:1; Bengali's alphasyllabary script causes high token fertility; rural internet access is only 36.5%. These metrics are correlated with model performance gaps evaluated through benchmarks like BenLLM-Eval. The research systematically identifies how these structural factors interact, creating a compounded disadvantage for Bengali NLP.

Key Results

  • Bengali's web presence is less than 0.5%, despite representing nearly 4% of the global population, leading to severe data scarcity for training large language models.
  • Major multilingual corpora allocate approximately 30 billion tokens to Bengali, compared to about 2 trillion tokens for English, resulting in a 67:1 ratio that hampers model performance.
  • Bengali's alphasyllabary script results in higher token fertility, requiring more subword units to represent the same content, which impacts training efficiency and model accuracy.
  • Internet penetration in rural Bangladesh is only 36.5%, limiting access to cloud-based AI tools, especially in low-connectivity environments, exacerbating digital inequality.

Significance

This research underscores the deep-rooted structural biases embedded in AI infrastructure, highlighting how historical resource allocation, design defaults, and socio-economic factors systematically exclude low-resource languages like Bengali. By framing data scarcity as a structural barrier, it advocates for offline-first deployment strategies and interdisciplinary approaches to promote linguistic equity. The findings challenge the prevailing narrative that performance gaps are solely technical, urging policymakers, researchers, and developers to rethink foundational assumptions and prioritize inclusive infrastructure development, ultimately contributing to global digital inclusion and cultural preservation.

Technical Contribution

The paper offers a comprehensive analysis of four interconnected structural failures—web content disparity, token scarcity, tokenization inefficiency, and connectivity barriers—using quantitative metrics and linguistic insights. It introduces a multi-dimensional framework for diagnosing language inequity in AI, emphasizing resource distribution and design biases. The proposal of offline-first architectures, leveraging local inference models, represents a significant shift from cloud-dependent paradigms, enabling scalable, sustainable, and equitable AI deployment for low-resource languages. The integration of linguistic typology and tokenization analysis advances understanding of script-specific challenges, paving the way for tailored NLP solutions.

Novelty

This work is pioneering in systematically connecting macro-level infrastructural disparities with micro-level technical challenges in low-resource language AI. Unlike prior studies focusing solely on data augmentation or model optimization, it emphasizes the structural roots—web content distribution, tokenization inefficiencies, and connectivity constraints—that perpetuate inequity. The concept of structural silence, as a systemic exclusion mechanism, offers a novel lens for analyzing language bias. The advocacy for offline-first design as an equity-oriented strategy marks a paradigm shift in AI deployment for marginalized languages, setting a new direction for inclusive NLP research.

Limitations

  • The analysis centers on Bengali as a case study; other low-resource languages may face different or additional barriers, requiring further contextual research.
  • Quantitative performance assessments rely on existing benchmarks; real-world educational and social impacts need longitudinal studies and user feedback.
  • Implementing offline models involves logistical challenges, such as hardware costs, updates, and maintenance, which are not fully addressed in this study.

Future Work

Future research should expand to diverse low-resource languages, developing culturally and linguistically tailored datasets and benchmarks. Emphasis on local inference models, energy efficiency, and adaptive architectures will enhance deployment in remote areas. Interdisciplinary collaborations between linguists, AI engineers, and educators are essential to refine tokenization schemes, corpus construction, and pedagogical integration. Policy initiatives should prioritize resource reallocation, infrastructure development, and community involvement to foster sustainable digital ecosystems for underrepresented languages. Long-term, these efforts aim to democratize AI access, preserve linguistic diversity, and empower marginalized communities worldwide.

AI Executive Summary

The rapid advancement of artificial intelligence has transformed many aspects of society, especially in education and language support. AI-powered tools promise scalable solutions to bridge access gaps in under-resourced communities, offering personalized tutoring, translation, and content generation. However, beneath this optimistic outlook lies a critical challenge: the infrastructure supporting these tools is inherently biased towards high-resource languages like English, inadvertently marginalizing speakers of underrepresented languages such as Bengali.

This disparity is rooted in historical resource allocation, technological defaults, and socio-economic factors. Bengali, despite being one of the most spoken languages globally with over 285 million speakers, remains vastly underrepresented in digital content and AI datasets. The web presence of Bengali content accounts for less than 0.5%, a stark contrast to its population share of nearly 4%. This web content gap translates into a severe shortage of training data—only about 30 billion tokens are allocated to Bengali in major multilingual corpora, compared to 2 trillion tokens for English. Such data scarcity hampers the ability of language models to learn and generalize effectively.

Moreover, the unique script of Bengali, an alphasyllabary, introduces a tokenization penalty. Standard tokenizers like Byte Pair Encoding (BPE) and WordPiece, optimized for Latin scripts, fragment Bengali words into a higher number of subword units, increasing token fertility and computational overhead. This structural inefficiency further exacerbates the data deficit, making model training less effective.

Adding to these technical barriers is the digital divide in rural Bangladesh, where internet penetration is only 36.5%. Cloud-dependent AI tools become inaccessible to many learners, deepening educational inequities. These interconnected barriers form what the authors term 'structural silence'—a systemic exclusion of low-resource languages from mainstream AI development, not through explicit policies, but through default design choices that favor dominant languages.

Addressing these issues requires a paradigm shift. The authors advocate for offline-first AI architectures, emphasizing local inference models that operate independently of internet connectivity. Such models are not only more accessible but also more sustainable, reducing energy consumption and infrastructure costs. They argue that designing AI systems with equity in mind involves rethinking resource distribution, linguistic analysis, and deployment strategies.

This research underscores the importance of interdisciplinary collaboration—combining insights from linguistics, computer science, and education—to create inclusive AI ecosystems. It calls for a reevaluation of publication norms to recognize foundational work in corpus construction and system design for low-resource languages. Ultimately, the goal is to democratize AI access, preserve linguistic diversity, and empower marginalized communities, ensuring that technological progress benefits everyone, not just the privileged few.

Deep Dive

Abstract

Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-resourced communities. Yet the infrastructure underlying these tools, including training corpora, tokenization schemes, evaluation benchmarks, and deployment architectures, can systematically disadvantage speakers of underrepresented languages before a model is trained. This paper examines these structural barriers through Bengali, one of the world's most widely spoken languages, focusing on AI-assisted education in low-connectivity environments. We identify four interlocking failures: a severe web presence gap, with Bengali accounting for less than 0.5% of global web content despite representing nearly 4% of the global population; a 67:1 training-token deficit between English and Bengali in major multilingual corpora; a tokenization penalty associated with Bengali's alphasyllabary script that compounds the data deficit through higher token fertility; and connectivity exclusion, with individual internet penetration at 36.5% in rural areas compared with 71.4% in urban areas. These failures reflect longstanding resource-allocation decisions, institutional priorities, and design defaults that did not center underrepresented languages in mainstream AI development. We argue that dataset scarcity should be understood as a structural barrier rather than an isolated technical limitation, and that offline-first design should be treated as an equity-oriented infrastructure strategy. We conclude with directions for linguistics and AI research aimed at reducing these structural inequalities.

cs.CL cs.AI cs.CY

References (19)

Overview of BLP-2025 Task 2: Code Generation in Bangla

Nishat Raihan, M. Jawad, Md Mezbaur Rahman et al.

2025 14 citations ⭐ Influential

Ethnologue

R. Collin

2010 152 citations ⭐ Influential

XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages

Tahmid Hasan, Abhik Bhattacharjee, Md Saiful Islam et al.

2021 541 citations View Analysis →

The United Nations

L. Fagerlund

1993 12551 citations

Cognitive Load Theory

M. Abkemeier

2020 2978 citations

Reliably exploring the presence of languages on the Internet

Daniel Pimienta

2024 2 citations

Learning subject content through a foreign language should not ignore human cognitive architecture: A cognitive load theory approach

Stéphanie Roussel, Danielle Joulia, A. Tricot et al.

2017 122 citations

Does Native Language Play a Role in Learning a Programming Language?

Adalbert Gerald Soosai Raj, Kasama Ketsuriyonk, J. Patel et al.

2018 42 citations

The State and Fate of Linguistic Diversity and Inclusion in the NLP World

Pratik M. Joshi, Sebastin Santy, A. Budhiraja et al.

2020 1435 citations View Analysis →

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla

Abhik Bhattacharjee, Tahmid Hasan, Kazi Samin Mubasshir et al.

2021 329 citations View Analysis →

Evaluating LLMs' Multilingual Capabilities for Bengali: Benchmark Creation and Performance Analysis

Shimanto Bhowmik, Tawsif Tashwar Dipto, Md Sazzad Islam et al.

2025 10 citations View Analysis →

QLoRA: Efficient Finetuning of Quantized LLMs

Tim Dettmers, Artidoro Pagnoni, Ari Holtzman et al.

2023 5335 citations View Analysis →

BenLLM-Eval: A Comprehensive Evaluation into the Potentials and Pitfalls of Large Language Models on Bengali NLP

M. Kabir, Mohammed Saidul Islam, Md Tahmid Rahman Laskar et al.

2023 37 citations View Analysis →

IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

Mohammed Safi Ur Rahman Khan, Priyam Mehta, A. Sankar et al.

2024 71 citations View Analysis →

Improving Bengali and Hindi Large Language Models

A. Shahriar, Denilson Barbosa

2024 4 citations

Goldfish: Monolingual Language Models for 350 Languages

Tyler A. Chang, Catherine Arnett, Zhuowen Tu et al.

2024 28 citations View Analysis →

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

Pierre-Carl Langlais, C. Hinostroza, Mattia Nee et al.

2025 18 citations View Analysis →

The Teacher's Dilemma: Balancing Trade-Offs in Programming Education for Emergent Bilingual Students

Emma R. Dodoo, Tamara Nelson-Fromm, M. Guzdial

2025 2 citations View Analysis →

Bridging the Last Mile: Unpacking the Rural Digital Divide in Bangladesh

Rayhan Rashed, Muhammad Ali, Sadia Sharmin et al.

2025 4 citations

Cited By (1)

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities