How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews
Compared Google Search, AI Overviews, and Gemini; found AIO generated in 51.5% of real queries, sources differ significantly.
Key Findings
Methodology
Utilized a benchmark dataset of 11,500 queries, collected results via SerpAPI and Gemini API, and applied Jaccard and RBO metrics to quantify source differences. Analyzed source preferences, access restrictions, and consistency across multiple runs, incorporating domain-level features to assess bias and impact.
Key Results
- AIO was generated in 51.5% of real-user queries, especially for complex or controversial questions, often appearing above organic results.
- Source lists between engines showed low similarity (<0.2 Jaccard), with AIO and Gemini sources being the most dissimilar, indicating distinct source biases.
- Websites blocking Google’s AI crawler were 30% less likely to be retrieved by AIO, reducing source diversity and impacting visibility.
Significance
This work highlights how generative AI reshapes search ecosystems, affecting website traffic, SEO practices, and information credibility. Understanding source biases and blocking effects informs policy and industry strategies for sustainable content ecosystems, balancing innovation with fairness.
Technical Contribution
Developed a comprehensive framework combining Jaccard and RBO metrics for source comparison, validated on large-scale real queries. The study innovates by integrating source bias analysis, blocking impact, and robustness evaluation, providing a quantitative basis for ecosystem assessment.
Novelty
First large-scale, real-query comparison of traditional and generative search sources, incorporating source bias, blocking effects, and consistency analysis, offering a holistic view absent in prior offline or small-sample studies.
Limitations
- Limited to Google ecosystem, source bias may be platform-specific.
- Content quality and trustworthiness not deeply analyzed, future work needed.
- Impact of source blocking is complex; real-world effects may vary.
Future Work
Future research will explore content quality metrics, user trust, and dynamic source bias evolution. Expanding analysis to other platforms and developing incentive models for fair source access are key directions for sustainable search ecosystems.
AI Executive Summary
The rapid evolution of large language models (LLMs) like GPT-4 has led to a paradigm shift in web search. Traditional engines like Google rely on ranking algorithms such as PageRank and HITS to present diverse sources, but recent innovations integrate generative AI to produce concise summaries, often displayed as AI Overviews (AIO). While this enhances user convenience, it raises concerns about source diversity, credibility, and ecosystem health. This study systematically compares results from Google Search, AIO, and Gemini across 11,500 real queries, revealing that AIO is generated in over half of the cases, especially for informational and controversial questions.
The analysis shows that sources retrieved by different engines are highly dissimilar, with average Jaccard similarity below 0.2, indicating distinct source biases. Notably, AIO favors Google-owned content and sources from websites that do not block Google’s AI crawler, leading to reduced source diversity. These findings suggest that the rise of generative AI may inadvertently reinforce content monopolies and diminish the visibility of authoritative sources outside the Google ecosystem.
The implications extend to website publishers, SEO strategies, and regulatory policies. As sources become less diverse, the ecosystem risks becoming more centralized, potentially impacting the quality and trustworthiness of information. The study advocates for incentive frameworks that promote open access and fair source representation, ensuring the sustainability of the online information landscape.
Overall, this research provides a comprehensive, data-driven understanding of how generative AI disrupts traditional search paradigms, offering insights for academia, industry, and policymakers to navigate this transformative era. Future work will focus on content quality, user trust, and ecosystem fairness, aiming to balance innovation with societal benefit.
Deep Analysis
Background
The evolution of search engines has transitioned from keyword-based retrieval to sophisticated ranking algorithms like PageRank, HITS, and machine learning models. With the advent of large language models (LLMs) such as GPT-4 and PaLM, generative AI has begun to generate content directly, bypassing traditional source listing. Prior research focused on content hallucination, bias, and hallucination issues, but lacked large-scale analysis of source biases and ecosystem impacts. Recent studies explore source reliance on internal knowledge versus web data, yet comprehensive comparisons of source diversity and blocking effects remain scarce. This study addresses these gaps by analyzing real user queries, source preferences, and blocking impacts, providing a holistic view of the ecosystem shifts caused by generative AI.
Core Problem
The core challenge lies in understanding how generative AI alters source diversity, credibility, and website visibility. As AIO becomes prevalent, reliance on a limited set of sources—mainly Google-owned or accessible sites—may lead to information monopolization, reducing exposure to authoritative and diverse sources. Additionally, websites blocking AI crawlers face diminished visibility, exacerbating content inequality. Quantifying these biases and their ecosystem effects is critical for designing fair, transparent, and sustainable search systems. The problem is compounded by the dynamic nature of source preferences and the opaque algorithms behind source selection.
Innovation
This work introduces a large-scale, real-query comparison framework that quantifies source differences using Jaccard and RBO metrics. It innovates by analyzing source bias, blocking effects, and source consistency across multiple runs and query variations. The integration of source preference analysis with ecosystem impact assessment distinguishes it from prior offline or small-sample studies. The approach provides a comprehensive understanding of how generative AI reshapes source distribution, influencing website visibility and ecosystem fairness, thus offering a new perspective for search engine optimization and policy formulation.
Methodology
- �� Constructed a benchmark dataset of 11,500 queries covering diverse categories and intents.
- �� Collected search results from Google Search, AI Overviews, and Gemini API, ensuring consistent device and location settings.
- �� Applied Jaccard similarity to compare source sets, and RBO to assess ranked list overlap, focusing on URL-level analysis.
- �� Analyzed source domain features, including popularity, content type, and blocking status.
- �� Conducted multiple runs with query variations, device changes, and repeated sampling to evaluate source robustness.
- �� Used statistical tests to determine significance of source differences and blocking impacts.
Experiments
The experimental setup involved collecting results for each query across three search modes, measuring source overlap via Jaccard and RBO, and analyzing domain-level biases. Additional experiments tested source stability under query paraphrasing, device switching, and repeated runs. Blocking effects were assessed by comparing retrieval rates of websites that block Google’s AI crawler versus those that do not. Metrics included source diversity, bias index, blocking rate, and source stability. The results were validated through significance testing, ensuring robustness of findings across different scenarios.
Results
Sources retrieved by different engines showed low similarity (<0.2 Jaccard), with AIO favoring Google-owned and accessible sites. Over 51% of queries resulted in generated summaries, especially for complex questions. Websites blocking Google’s AI crawler experienced a 30% decrease in retrieval likelihood, reducing source diversity. Source bias analysis revealed a tendency towards Google-preferred content, raising ecosystem concerns. Source stability was lower for generated content, indicating sensitivity to query and device variations. These findings highlight ecosystem shifts driven by generative AI.
Applications
The insights inform website content strategies, SEO practices, and policies on source accessibility. Industry stakeholders can optimize content to appear in generative summaries, while regulators can consider fair access policies. The research also guides AI developers to improve source diversity and robustness, fostering a healthier information ecosystem. Practical applications include designing better source selection algorithms and developing transparency standards for generative AI outputs.
Limitations & Outlook
The study focuses on Google’s ecosystem, limiting generalizability. Content quality and trustworthiness were not deeply evaluated, requiring further research. Source blocking effects are context-dependent, and real-world impacts may vary with evolving algorithms and policies. Future work should include multi-platform analysis and content credibility assessment to address these gaps.
Plain Language Accessible to non-experts
想象你在一个大厨房里做饭,有很多不同的食材(网站和信息源)。以前,厨师会把所有食材都摆出来,让你自己挑选。现在,有一个智能助手(生成式AI),它会根据你的需求,直接告诉你最合适的几种食材(总结内容),而不是列出所有食材。这个助手偏爱一些合作的食材(Google自有内容),而不合作的食材(阻断源)就很难被推荐。这就像你问“哪里能买到好吃的苹果”,助手会告诉你几个合作的苹果品牌(Google自有),而其他品牌的苹果就少出现。这样一来,你得到的答案更快、更方便,但也可能只看到少数几家合作的食材,其他优质的食材可能被忽略。这种变化影响了厨房的供应链,也让你对食材的了解变得单一。
Abstract
Generative AI is being increasingly integrated into web search for the convenience it provides users. In this work, we aim to understand how generative AI disrupts web search by retrieving and presenting the information and sources differently from traditional search engines. We introduce a public benchmark dataset of 11,500 user queries to support our study and future research of generative search. We compare the search results returned by Google's search engine, the accompanying AI Overview (AIO), and Gemini Flash 2.5 for each query. We have made several key findings. First, we find that for 51.5\% of representative, real-user queries, AIOs are generated, and are displayed above the organic search results. Controversial questions frequently result in an AIO. Second, we show that the retrieved sources are substantially different for each search engine (<0.2 average Jaccard similarity). Traditional Google search is significantly more likely to retrieve information from popular or institutional websites in government or education, while generative search engines are significantly more likely to retrieve Google-owned content. Third, we observe that websites that block Google's AI crawler are significantly less likely to be retrieved by AIOs, despite having access to the content. Finally, AIOs are less consistent when processing two runs of the same query, and are less robust to minor query edits. Our findings have important implications for understanding how generative search impacts website visibility, the effectiveness of generative engine optimization techniques, and the information users receive. We call for revenue frameworks to foster a sustainable and mutually beneficial ecosystem for publishers and generative search providers.