Evaluating Verifiability in Generative Search Engines
Proposes citation recall and precision metrics to evaluate verifiability; finds only 51.5% sentences fully supported in top search engines.
Key Findings
Methodology
This study conducts human evaluations on Bing Chat, NeevaAI, perplexity.ai, and YouChat, analyzing responses to diverse queries from sources like Google logs and Reddit. It introduces citation recall and citation precision metrics, measuring the extent and correctness of source support. Responses are segmented into sentences, and each sentence's support is judged via the AIS framework, enabling objective quantification of verifiability. Human annotators assess fluency, utility, and support levels, providing a comprehensive evaluation of each system's performance.
Key Results
- The average sentence support rate is only 51.5%, indicating many responses lack full backing. Citation support accuracy stands at 74.5%, revealing prevalent support errors. High fluency and perceived usefulness do not correlate with verifiability, with a strong negative correlation (r=-0.96). Copying or paraphrasing web content boosts citation precision but diminishes content utility, highlighting a trade-off. These findings underscore the need for better source validation in deployed systems.
- Different systems show significant variation; Bing Chat and YouChat perform slightly better but overall support remains low. The low support rate increases the risk of misinformation, especially when users cannot verify claims. The citation F1 metric effectively combines recall and precision, offering a balanced evaluation tool.
- The study emphasizes that current commercial systems often prioritize fluency and utility over source support, risking user deception. The metrics and evaluation framework proposed serve as essential tools for advancing trustworthy AI, guiding future improvements in verifiability and transparency.
Significance
This research exposes critical gaps in the trustworthiness of popular generative search engines, highlighting that high fluency and perceived helpfulness do not guarantee factual accuracy or source support. The metrics introduced provide industry and academia with a standardized way to quantify and improve content verifiability. Addressing these issues is vital for deploying AI systems in sensitive domains like healthcare, finance, and legal advice, where misinformation can have serious consequences. The work pushes the field toward more transparent, source-supported generation, fostering user trust and reducing misinformation spread.
Technical Contribution
The paper pioneers the formalization of citation recall and precision as core metrics for evaluating verifiability in generative models. It combines human annotation with AIS-based binary judgments, establishing a robust framework for support assessment. The introduction of citation F1 as a harmonic mean metric balances the trade-off between coverage and correctness. The analysis of system behaviors reveals how copying and paraphrasing strategies influence support metrics, providing insights for future model training and evaluation. This work sets a new standard for transparency in generative AI evaluation.
Novelty
This is the first comprehensive attempt to quantify the verifiability of commercial generative search engines using explicit citation-based metrics. Unlike prior work focusing solely on fluency or relevance, this study emphasizes source support, addressing a critical gap in trustworthy AI. The combination of human annotation, AIS framework, and new metrics offers a novel, practical approach to benchmarking and improving system reliability, marking a significant step forward in responsible AI development.
Limitations
- The reliance on human annotation introduces subjective bias despite multiple annotators. Future automation could improve scalability. The sentence-level support assessment may overlook finer-grained support nuances. The sample size, though diverse, remains limited; broader validation is needed. Additionally, copying content from web pages, while improving citation support, may mask underlying factual inaccuracies, requiring further investigation.
Future Work
Future research should focus on automating support detection to scale evaluations. Integrating multi-modal sources like images and videos could enhance verification robustness. Developing standardized benchmarks and datasets for verifiability will facilitate industry-wide adoption. Exploring methods to incentivize truthful source citation during model training and deployment is also crucial. Ultimately, establishing transparent, source-supported generation will be key to trustworthy AI systems.
AI Executive Summary
The rapid adoption of generative search engines such as Bing Chat and YouChat promises a transformative shift in how users access information online. These systems generate fluent, seemingly informative responses with embedded citations, aiming to enhance user trust and utility. However, underlying issues of verifiability—whether statements are fully supported by sources and whether citations are accurate—remain largely unaddressed. This gap poses significant risks, especially as users often rely on these responses for critical decisions.
To systematically evaluate these concerns, this study introduces two metrics: citation recall, measuring the proportion of statements fully supported by citations, and citation precision, assessing the correctness of citations supporting each statement. Through human annotation of responses from four leading systems across diverse query types, the research reveals a troubling reality: only about half of the generated sentences are fully supported, and just three-quarters of citations are accurate. Interestingly, responses that appear more helpful tend to have lower support levels, indicating a trade-off between perceived utility and verifiability.
The findings highlight that current commercial systems often prioritize fluency and user satisfaction over source support, inadvertently increasing misinformation risks. Copying or paraphrasing web content can boost citation precision but at the expense of content utility, underscoring the complexity of balancing these aspects. The proposed metrics and evaluation framework provide a foundation for future development, emphasizing the importance of source validation in trustworthy AI.
This work calls for industry-wide standards and further research into automated support detection, multi-modal verification, and incentivizing source accuracy. By addressing these challenges, the AI community can develop more transparent, reliable systems that truly serve as trustworthy tools for information seeking, ultimately fostering greater user confidence and reducing the spread of false information.
Deep Dive
Plain Language Accessible to non-experts
想象你在厨房里做饭,食谱上写着每道菜的步骤和用料。现在,假设有一个超级厨师,他可以根据你的问题,快速做出一道菜,还会告诉你用的材料来自哪里。可是,有时候这个厨师会随意拼凑材料,或者引用的资料根本不靠谱。这就像生成式搜索引擎,虽然看起来回答得很流畅、很有用,但其实可能没有真正的依据。为了确保答案是真的,研究人员设计了一套“支持度”检查机制,就像你检查厨师用的材料是不是新鲜、是不是对的。结果显示,大部分回答都没有充分的出处支持,容易让人误信。未来,希望这个“支持度”机制能让厨师(系统)做得更靠谱,让我们吃得更放心。
ELI14 Explained like you're 14
想象你问一个超级聪明的朋友关于宇宙的秘密,他用很酷的方式回答你,还会告诉你资料来自哪里。但是,有时候这个朋友会胡乱说话,或者引用的资料根本不靠谱。虽然他的回答听起来很酷、很有用,但你不知道是不是可信。科学家们也遇到这个问题,他们希望这些“聪明的朋友”说的话都能有确凿的出处。于是,他们设计了一套“支持度”系统,就像你检查朋友说的话是不是有证据一样。研究发现,大部分回答都没有充分的证据支持,容易误导我们。未来,他们希望这个系统能帮这些“聪明的朋友”变得更可靠,让我们更放心相信他们说的话。
Abstract
Generative search engines directly generate responses to user queries, along with in-line citations. A prerequisite trait of a trustworthy generative search engine is verifiability, i.e., systems should cite comprehensively (high citation recall; all statements are fully supported by citations) and accurately (high citation precision; every cite supports its associated statement). We conduct human evaluation to audit four popular generative search engines -- Bing Chat, NeevaAI, perplexity.ai, and YouChat -- across a diverse set of queries from a variety of sources (e.g., historical Google user queries, dynamically-collected open-ended questions on Reddit, etc.). We find that responses from existing generative search engines are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations: on average, a mere 51.5% of generated sentences are fully supported by citations and only 74.5% of citations support their associated sentence. We believe that these results are concerningly low for systems that may serve as a primary tool for information-seeking users, especially given their facade of trustworthiness. We hope that our results further motivate the development of trustworthy generative search engines and help researchers and users better understand the shortcomings of existing commercial systems.