Sycamore: Characterizing Synthetic Personas for Evaluating Genomics Visualization Retrieval
Sycamore framework compares ungrounded, grounded synthetic personas and real experts in genomics visualization evaluation, revealing biases and limitations.
Key Findings
Methodology
Using a three-condition probe design, Sycamore evaluates Geranium, a multimodal genomics search engine, with ungrounded (LLM-only), grounded (based on van den Brandt et al.'s interview-coded evidence), and expert feedback. Grounded personas incorporate evidence retrieval via PersonaCite, aligning responses with documented user concerns. Ungrounded personas rely solely on generic LLM priors. The study analyzes differences in language, focus, and modality preferences, revealing how grounding influences feedback realism and bias. The framework enables systematic comparison of feedback styles and biases across conditions.
Key Results
- Grounded synthetic personas produce feedback more aligned with documented user language and concerns, while ungrounded personas tend to focus on operational details, with both missing the image modality preference observed in experts. All synthetic conditions converge on a 'find-and-adapt' approach but fail to reflect the expert preference for image modality. The abstention rate in grounded personas reaches about 50%, indicating reliance on retrieved evidence, whereas ungrounded personas always respond, showing lack of evidence checking.
- In modality ranking, ungrounded personas uniformly prefer Specification > Image > Text, whereas grounded personas show variability. Both synthetic conditions diverge from expert preferences, highlighting the importance of grounding for realistic feedback. The analysis underscores the potential and limitations of LLM-based synthetic personas in domain-specific evaluation.
- The experiments confirm that grounding improves the alignment of synthetic feedback with real user concerns but also exposes biases and gaps, especially in multimodal preference modeling. These insights inform future development of more accurate and diverse synthetic user models for scientific visualization assessment.
Significance
This work advances the understanding of how LLM-generated synthetic personas can supplement or partially replace real user feedback in domain-specific visualization evaluation. By systematically comparing grounded and ungrounded models against expert feedback, it highlights the importance of evidence grounding for realism and bias mitigation. The framework offers scalable, cost-effective alternatives to traditional user studies, especially in resource-scarce fields like genomics. It paves the way for more nuanced, multi-dimensional evaluation methods that can adapt to complex scientific data environments, ultimately accelerating the development and validation of visualization tools in biomedical research and beyond.
Technical Contribution
The paper introduces a novel three-condition evaluation framework combining large language models, evidence retrieval, and expert feedback comparison. It innovates by integrating PersonaCite-based evidence retrieval with a modular RAG architecture, enabling the generation of more realistic, evidence-grounded synthetic user feedback. The framework systematically analyzes feedback differences, revealing biases and focus shifts induced by grounding. It also provides a scalable methodology for domain-specific evaluation, bridging the gap between static personas and real user diversity. The approach demonstrates how to leverage retrieval-augmented generation to improve the fidelity of synthetic user models in complex, multimodal scientific contexts.
Novelty
This is the first systematic comparison of ungrounded versus grounded synthetic personas in a high-stakes genomics visualization context. The integration of evidence retrieval with persona modeling represents a significant innovation, addressing limitations of prior static or overly generic user simulations. The study uniquely highlights how grounding shifts feedback towards documented user concerns, providing empirical validation for evidence-based persona generation. It advances the field by demonstrating that combining domain knowledge with LLMs can produce more realistic and useful synthetic feedback, setting a new standard for automated evaluation in scientific visualization.
Limitations
- The reliance on evidence retrieval may limit the diversity of responses, especially when evidence is sparse or biased, potentially constraining the range of simulated user behaviors.
- The current framework is validated only on Geranium, and its generalizability to other complex visualization systems remains to be tested.
- Despite improvements, synthetic feedback still diverges from expert preferences in multimodal biases, indicating room for further refinement in modeling user diversity and preferences.
Future Work
Future research will focus on integrating multi-source real user data to enhance diversity and realism of synthetic personas. Developing adaptive learning mechanisms, such as reinforcement learning, could enable models to better capture evolving user preferences. Extending the framework to other scientific domains and more interactive scenarios will test its scalability and robustness. Additionally, incorporating multi-turn dialogues and dynamic context understanding could further improve the fidelity of synthetic feedback, making it a more reliable tool for automated evaluation and iterative design of complex visualization systems.
AI Executive Summary
In the rapidly evolving field of genomics data visualization, traditional user studies face significant challenges due to the scarcity of domain experts and the diversity of user needs. This bottleneck hampers comprehensive evaluation and iterative improvement of visualization tools. To address this, the present study introduces Sycamore, a novel framework that systematically compares ungrounded, grounded synthetic personas, and real expert feedback within a unified evaluation protocol.
Sycamore leverages Geranium, a multimodal genomics visualization retrieval system, as its evaluation platform. It employs a three-condition probe design: ungrounded personas generated solely from large language models (LLMs), grounded personas based on evidence retrieved from van den Brandt et al.'s interviews via PersonaCite, and actual expert feedback from prior studies. This setup enables a detailed analysis of how grounding influences feedback quality, language, and focus. The grounded personas incorporate evidence retrieval, which aligns their responses with documented user concerns, while ungrounded personas rely on generic priors, often drifting toward operational details.
Results reveal that grounding shifts synthetic feedback toward realistic user concerns, improving the fidelity of simulated responses. However, both synthetic conditions tend to miss the expert preference for image modality, indicating areas for further refinement. The grounded personas exhibit higher abstention rates (~50%), reflecting their reliance on retrieved evidence, whereas ungrounded personas always respond, revealing their lack of evidence validation.
This work provides valuable insights into the strengths and limitations of synthetic personas in domain-specific visualization evaluation. It demonstrates that evidence-grounded models can better emulate real user concerns, offering a scalable alternative to costly expert studies. The framework's modular design allows adaptation to other scientific fields and complex interactive systems, promising broad applicability. Future efforts will focus on enhancing diversity, integrating multi-source data, and extending to more dynamic, multi-turn interactions.
Overall, Sycamore advances the methodology of automated, evidence-based user simulation, contributing to more reliable, scalable evaluation processes in scientific visualization. It highlights the importance of grounding in domain knowledge, paving the way for smarter, more realistic synthetic user models that can accelerate tool development and validation in biomedical research and beyond.
Deep Dive
Abstract
Evaluating visualization systems in niche domains such as genomics is challenging due to scarcity of domain experts and difficulty recruiting a representative user base. While LLM-based synthetic personas are increasingly used to ease evaluation bottlenecks, they face well-founded skepticism. Rather than weighing synthetic personas as substitutes for real users, we ask a fundamental open question: when synthetic personas evaluate a real visualization system, what do they actually produce, and how does that output change when grounded in documented human contexts? We present Sycamore, an exploratory three-condition probe design using Geranium, a search engine for multimodal genomics visualization, as a case study. Sycamore evaluates Geranium using: (1) ungrounded synthetic personas from generic LLM priors; (2) grounded synthetic personas constrained by voice-of-customer artifacts from a prior interview study; and (3) a published baseline study of real domain experts. We observe that grounding shifts synthetic feedback toward the language and concerns of documented users, while ungrounded evaluators drift toward operational specifics that real participants did not raise; both synthetic conditions, however, converge on a find-and-adapt frame and miss the image-modality preference observed in the expert study. We discuss what these observations imply for where synthetic personas might fit alongside expert studies in domain-specific visualization evaluation. All supplemental materials are available at https://osf.io/kdfr3/.