Who Uses Open-Weight Models? China and the Shifting Geography of AI in Science
Analyzing 21 million papers with NLP, found Chinese researchers favor open-source models like Qwen, reaching 44% usage in 2026, indicating regional shifts in AI ecosystems.
Key Findings
Methodology
This study employs a mixed NLP pipeline, including dictionary extraction, SciBERT classifiers, and statistical models (logistic regression, multinomial) to identify and verify model mentions and usage in full-text papers from S2ORC. The process involves training classifiers on silver-standard annotations, distinguishing between mere mentions and actual use, and linking model mentions to specific families. Author affiliation and name data from OpenAlex enable regional analysis. The models quantify the heterogeneity of open-weight adoption, especially emphasizing Chinese-developed models like Qwen and DeepSeek.
Key Results
- By 2026, 44.0% of single-model papers used open-source models, with Chinese models accounting for 60.5% of this share. Chinese institutional affiliation increased the odds of using open models by 2.23 times, contributing 44% of the growth since 2023. Qwen alone appeared in 22% of single-model papers and 62.1% of multi-model papers, dominating open-weight choices in China.
- In multi-model papers, open-source models appeared in 87.2%, with Chinese models like Qwen and DeepSeek prevalent. The adjusted usage rate of Chinese open-weight models among Chinese-affiliated papers was 37.1%, compared to 9.2% in non-Chinese papers, highlighting regional ecosystem shifts.
- Statistical models confirm that Chinese researchers are disproportionately inclined toward Chinese open-weight models, with logistic regression showing OR=2.23 for Chinese institutional affiliation, and multinomial analysis indicating strong concentration of Chinese models in the ecosystem. These findings reflect a geopolitical and cultural influence on AI model adoption.
Significance
This research uncovers the geopolitical and cultural factors shaping AI model ecosystems, challenging the notion that open-source equates to open science. The regional preference for Chinese open-weight models signifies a broader realignment driven by policy, market, and sociocultural influences. It highlights the importance of understanding local contexts in global AI deployment, informing policymakers and researchers about the evolving landscape of AI-enabled science and innovation. The findings emphasize that model ecosystem shifts are not merely technical but deeply embedded in geopolitical strategies, affecting scientific reproducibility and access.
Technical Contribution
The study introduces a comprehensive NLP framework integrating dictionary matching, SciBERT classifiers, and statistical modeling to accurately identify and analyze model usage in scientific literature. It innovates by quantifying regional biases and ecosystem shifts, providing a scalable approach to study AI adoption patterns. The combination of large-scale text mining with advanced statistical analysis offers a novel methodology for understanding the socio-technical dynamics of AI in science, with potential applications in policy and ecosystem design.
Novelty
This is the first large-scale, data-driven investigation into the geographic and institutional biases in AI model adoption within scientific literature. It uniquely combines NLP techniques with statistical modeling to reveal how geopolitical factors influence model ecosystems, especially highlighting the dominance of Chinese-developed open-source models. This work advances understanding beyond traditional case studies, providing a quantitative, global perspective on AI’s integration into science.
Limitations
- The reliance on text matching and classification may lead to misclassification, especially in ambiguous contexts or background mentions. The regional attribution based on author affiliation and names may be affected by institutional changes or name ambiguities. Data is current only up to mid-2026, and rapid ecosystem changes may not be captured in future analyses.
Future Work
Future research will incorporate model performance metrics and real-world applications to understand decision factors behind model choice. Expanding regional analysis to include other countries, and integrating policy and market data, will deepen insights into ecosystem dynamics. Longitudinal studies tracking ecosystem evolution and impact on scientific reproducibility are also planned.
AI Executive Summary
The rapid proliferation of large language models (LLMs) has transformed scientific research, prompting a need to understand how researchers choose among competing models. This study leverages a sophisticated NLP pipeline to analyze over 21 million full-text scientific articles from the Semantic Scholar corpus, identifying mentions and actual usage of various LLMs. The findings reveal a significant shift over time: while GPT-family models dominated early research, their share declined as a broader array of models, especially open-source Chinese models like Qwen and DeepSeek, gained prominence. By 2026, 44% of single-model papers employed open-source models, with Chinese models accounting for over 60% of this share. Statistical analysis shows that Chinese researchers are 2.23 times more likely to adopt Chinese open-weight models, contributing substantially to the ecosystem's reconfiguration. This regional preference underscores the influence of geopolitical and cultural factors on AI deployment in science, challenging simplistic notions equating openness with transparency. The research methodology combines dictionary extraction, SciBERT classifiers, and advanced statistical models to quantify these dynamics, providing a comprehensive picture of the evolving AI landscape. The implications extend beyond technical innovation, highlighting how market, policy, and sociocultural contexts shape scientific instrument choices. Future work aims to incorporate performance metrics and broader regional analyses, fostering a nuanced understanding of AI’s role in global scientific progress. Overall, this work emphasizes that AI ecosystem shifts are deeply intertwined with geopolitical strategies, affecting scientific reproducibility, access, and innovation trajectories.
Deep Dive
Abstract
As LLMs have become a flashpoint for scientific research, computer scientists and STS scholars have advocated the use of open-weight models. Since LLM research has matured and more high-quality model families are available, have researchers adopted open-weight models? We present the first systematic study of model selection in scientific research, analyzing 21 million full-text articles through June 2026 from the Semantic Scholar Open Research Corpus (S2ORC). We employ a mixed NLP pipeline to extract model occurrences in article full text and determine whether they are used or merely mentioned by researchers. We divide our corpus into single- and multi-model family studies, which we take as a proxy for applied and foundational AI research. We find GPT-family models dominate both single- and multi-family research, but that both areas are becoming more diverse over time. In single-family papers, open-weight model use rises steadily, reaching 44.0% in 2026. However, we find that recent growth is driven by the availability of high-quality open-weight Chinese models. Further, a logistic regression model finds that open-weight adoption is heterogeneously distributed, estimating that researchers at Chinese institutions have 2.23 times the odds of using an open-weight model, accounting for 44.0% of the increase in open-weight adoption since 2023. A complementary multinomial model shows this association is concentrated in Chinese open-weight models: in 2026, their adjusted use is 37.1% among papers with Chinese affiliations, a 27.9 percentage-point over papers with no observed China link. These findings suggest that open-weight adoption in science is not a general turn toward open science, but part of a broader realignment of model ecosystems in which platforms and markets, and the sociocultural and geopolitical contexts which shape them, determine which AI systems become scientific instruments.