Large language models create an uneven informational layer over cities
Comparative analysis of three major LLMs across 304 neighborhoods in five US cities reveals systematic fabrication and omission biases, influenced by socioeconomic factors, affecting urban information distribution.
Key Findings
Methodology
Using synthetic user profiles (income, age, sex, residence) across five US cities, the study evaluates GPT-4, Llama-3.3-70b, and Gemini-2.0 in recommending restaurants. Combining SafeGraph, Yelp, Cuebiq data, it measures hallucination (nonexistent venues) and coverage gaps (missing real venues). Regression models analyze spatial and socioeconomic correlates, revealing shared biases across models and the impact of verified data on reducing hallucinations but not omissions.
Key Results
- Average hallucination rate is 36.8%, higher in neighborhoods with weaker digital footprints; verified data reduces hallucinations to near zero, but 47.5% of real venues are never recommended, with 31.9% shared across models, indicating systemic biases rooted in shared training corpora.
- Spatial analysis shows hallucination and omission correlate with neighborhood socioeconomic features: affluent, low-density areas exhibit higher hallucination; high omission in fragmented online footprints. Models favor popular, online-discussed venues, reinforcing existing hierarchies.
- User-level analysis reveals higher-income and tourist profiles are directed toward higher-priced, less frequented, more socially diverse venues, reflecting internalized social stratification. These patterns mirror real urban behaviors, suggesting models encode latent social structures.
- Consumer demand simulations indicate a shift from fast food and chain restaurants toward independent, full-service venues, with economic redistribution favoring higher-end dining, potentially reshaping urban food economies.
Significance
This research uncovers the structural biases embedded in LLMs as urban information layers, highlighting risks of reinforcing spatial and social inequalities. Understanding these biases is crucial for designing fairer AI systems and guiding urban policy, ensuring that digital urban infrastructure promotes equitable access and representation, rather than amplifying existing disparities.
Technical Contribution
The study introduces a framework combining verified data calibration and socioeconomic regression to distinguish epistemic (knowledge-based) from allocative (attention-based) biases. It demonstrates shared biases across models, emphasizing systemic issues rooted in training corpora. The approach offers a pathway for developing fairer recommendation algorithms and understanding AI-driven urban information dynamics.
Novelty
This is the first comprehensive comparison of multiple LLMs at the neighborhood scale, distinguishing hallucination from omission biases, and linking these biases to socioeconomic and spatial factors. It uniquely combines multi-source data and demographic analysis to reveal systemic, shared biases, advancing understanding of AI's role in urban information landscapes.
Limitations
- The focus on restaurant data limits generalization; other urban domains like transportation or housing require further study.
- Reliance on platform-specific data (Yelp, SafeGraph) may introduce platform bias, affecting broader applicability.
- Calibration with verified data addresses epistemic errors but does not fully uncover the internal model bias mechanisms, which need further exploration.
Future Work
Future research should explore causal mechanisms behind biases, incorporate multi-modal data, and develop interventions to mitigate unfair recommendations. Extending analysis to other urban sectors and integrating policy frameworks will be vital for ensuring equitable urban digital infrastructure.
AI Executive Summary
Cities are complex socio-technical systems where information plays a pivotal role in shaping urban experiences. The advent of large language models (LLMs) has introduced a new layer of urban information, influencing how people discover and consider places. This study systematically compares three prominent LLMs—GPT-4, Llama-3.3-70b, and Gemini-2.0—in recommending restaurants across 304 neighborhoods in five US cities. Using synthetic user profiles and multi-source data, the research uncovers significant biases: models frequently fabricate venues, especially in neighborhoods with weaker digital footprints, and systematically overlook real venues, with 47.5% of establishments never recommended. These biases are spatially patterned, correlating with socioeconomic factors such as income, density, and online activity, revealing that the models encode and reinforce existing urban inequalities.
Further analysis shows that even when models are constrained to verified data, hallucinations vanish but omissions persist, indicating that the biases are rooted in the models’ attention mechanisms rather than knowledge gaps. The recommendations also vary systematically across user demographics: higher-income and tourist profiles are directed toward more expensive, less frequented, and socially diverse venues, reflecting latent social stratification embedded in the models. These recommendation biases lead to a redistribution of consumer spending, favoring independent and full-service restaurants over fast-food chains, potentially reshaping urban economic patterns.
Overall, this work highlights the role of LLMs as a new form of urban infrastructure that unevenly distributes visibility, with profound implications for urban inequality and economic resilience. It calls for targeted interventions to address systemic biases, ensuring that AI-driven urban information promotes fairness, inclusivity, and sustainable city development.
Deep Analysis
Background
城市信息系统如GPS、点评平台已深刻影响城市空间认知。近年来,LLMs作为新型信息基础设施崛起,能在无监督条件下生成丰富的城市地点信息。前期研究揭示了偏见与不平等的数字表现,但缺乏邻里尺度的系统分析。本研究填补此空白,结合多源数据,分析模型偏差的空间与社会结构根源,为城市信息公平提供新视角。
Core Problem
现有推荐系统多依赖用户行为和显式评价,存在偏见积累。LLMs虽能生成丰富信息,但其偏差机制尚不明晰。虚构与遗漏的偏差可能导致信息失真,影响城市空间公平。特别是在邻里层面,偏差可能加剧社会经济差异,影响城市治理与经济分配,亟需系统性分析。
Innovation
本研究创新点在于:1)系统性比较三大模型在邻里尺度的推荐偏差;2)区分虚构(虚假推荐)与遗漏(未推荐真实地点);3)结合空间、社会经济与用户属性分析偏差根源;4)验证验证数据校准策略的效果。通过多源数据融合,揭示偏差的空间分布和共享性,为未来模型公平性设计提供理论基础。
Methodology
- �� 采集五城304区的真实餐厅数据(SafeGraph);• 设计合成用户画像(收入、年龄、性别、居住状态);• 使用三大LLMs生成推荐,结合验证数据筛查虚构;• 计算虚构率(虚假推荐占比)与遗漏率(未推荐真实地点比例);• 采用回归分析空间与社会经济特征对偏差的影响;• 比较不同模型偏差共享程度,分析偏差根源。
Experiments
采用合成画像在五城不同社区测试三大模型,评估虚构率与遗漏率。虚构率平均36.8%,遗漏率47.5%。通过空间回归分析,发现偏差与社区的线上线下信息碎片化密切相关。模型校准实验验证验证数据能消除虚构,但遗漏偏差仍存。模型偏差的空间与人口属性关系被系统揭示。
Results
模型虚构率平均36.8%,在信息碎片化区更高,验证数据校准后虚构几乎消除,但遗漏偏差持续存在,47.5%的真实地点未被推荐。偏差空间分布与社区的社会经济特征显著相关:高收入、低密度区偏向虚构,遗漏集中在信息缺失区域。用户偏好分析显示,模型偏向高收入、旅游用户推荐高价、少人流、社会多样性强的地点。这些偏差导致消费重分布,从快餐向高端餐厅转变,影响城市餐饮格局。
Applications
此研究为城市数字治理提供警示,推动模型偏差的识别与修正。可应用于优化城市信息平台,减少偏见,促进公平分配。未来,结合偏差分析与政策制定,可实现城市空间的公平与可持续发展。
Limitations & Outlook
仅分析餐厅类别,未来应扩展到交通、住宿等场景。数据来源有限,可能存在平台偏差。模型校准未深入模型训练机制,偏差根源仍待揭示。
Plain Language Accessible to non-experts
想象你在一个巨大的城市里,有很多餐厅,但你不知道所有的地方。现在,有个智能机器人(就像一个超级聪明的朋友),可以告诉你哪些餐厅值得去。可是,这个机器人有时候会告诉你一些不存在的餐厅(虚假信息),或者遗漏掉一些真正的好地方(遗漏)。它偏向推荐那些在网上讨论多、很有名的场所,而忽略了那些隐藏在角落的小店。不同的机器人(模型)虽然都这样,但它们的偏差有些相似。更有趣的是,机器人会根据你的背景(比如你是学生、富人、游客)推荐不同的餐厅,比如富人喜欢远一点的高档餐厅,学生喜欢便宜的快餐。这样一来,城市中的消费也会被引导到不同的地方,可能让一些小店变得更火,而一些快餐店变得冷清。这个研究帮我们理解这些智能“朋友”是怎么“看”城市的,也提醒我们要注意它们的偏见,避免让城市变得不公平。
ELI14 Explained like you're 14
想象你在一个超级大的城市里,想找个好吃的地方,但不知道所有餐厅在哪。于是你问一个特别聪明的机器人(就像一个超级厉害的助手),它会告诉你一些餐厅推荐。但是,这个机器人有时候会编造一些不存在的餐厅(虚假推荐),或者忘记告诉你一些真实存在的好地方(遗漏)。更神奇的是,它偏爱推荐那些在网上很火、很多人讨论的餐厅,而忽略了那些冷门的小店。不同的机器人虽然都这样,但它们的偏差很像,说明这是它们“学”到的共同偏见。更有趣的是,机器人会根据你的背景,比如你是学生、富人、还是游客,推荐不同的餐厅:富人喜欢远一点的高档餐厅,学生喜欢便宜的快餐,而游客会被带到更热闹、更有特色的地方。这就像城市的“看法”被机器人带偏了,可能让一些小店变得更火,而快餐店变得冷清。这个研究帮我们理解这些“机器人”是怎么“看”城市的,也提醒我们要注意它们的偏见,避免让城市变得不公平。
Abstract
Large language models (LLMs) are emerging as a new informational layer over cities, shaping which places people discover, consider, and ultimately visit. Yet little is known about which places they surface, which they ignore, and whether these patterns vary across communities and users and translate into real-world economic consequences. Here, we audit restaurant recommendations from three major LLMs across 304 neighborhoods in five U.S. cities using 320 synthetic user profiles spanning income, age, sex, and residential status. We find that LLMs both fabricate venues and systematically overlook real ones. Fabrication is concentrated in neighborhoods with weaker digital and physical footprints and disappears when models are provided with verified venue lists. In contrast, invisibility persists: even when choosing from a fixed set of real venues, 47.5% of establishments are never recommended, and 31.9% of these blind spots are shared across all three model families, indicating that uneven visibility reflects not only missing knowledge but also stable patterns of selective attention rooted in shared patterns of visibility rather than model-specific errors. The same selectivity extends to users. Within identical venue pools, higher-income users receive more expensive and less popular venues, while tourists are directed toward costlier but more socially diverse establishments than local residents. Simulating the resulting shifts in consumer demand suggests that widespread reliance on LLM recommendations would redirect visits and revenue away from chain and quick-service restaurants toward independent and full-service dining. Together, our findings show that LLMs act as a selective layer of urban information that unevenly distributes visibility across places and people, with potential consequences for local economies and urban inequality.