Silicon Sampling via Cross-Survey Transfer

TL;DR

Proposed cross-survey transfer framework; zero-shot LLMs achieve 52% accuracy on unseen items, nearing supervised models.

cs.AI 🔴 Advanced 2026-07-03 28 views
Chan-Tung Ku Chan Hsu Pei-Cing Huang Frank Cheng-shan Liu I-Ling Cheng Yihuang Kang
silicon sampling large language models survey simulation cross-survey transfer political attitudes

Key Findings

Methodology

This study proposes a cross-survey transfer framework, testing the individual-level predictive ability of large language models (LLMs) by partitioning survey items into context (Set A) and prediction (Set B) sets. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, the performance of three open-weight LLMs (27B-120B parameters) and supervised machine learning baselines was evaluated.

Key Results

  • Zero-shot LLMs achieved 52% accuracy on genuinely unseen items, only 6 percentage points lower than a supervised random forest, demonstrating strong predictive capabilities without training data.
  • A construct predictability hierarchy emerged, with 67% accuracy for partisan attitudes and only 23% for sovereignty, indicating varying prediction difficulty across constructs.
  • Variance collapse and safety alignment effects were observed in both LLMs and supervised models, with significant differences across model families.

Significance

This study reveals the potential and limitations of silicon sampling in individual-level prediction through the cross-survey transfer framework. It holds significant academic value and provides policymakers with a new tool for rapidly assessing public opinion, especially when traditional methods are costly and slow.

Technical Contribution

Contributions include a novel individual-level evaluation framework, validation of LLM performance in a non-WEIRD, Mandarin political context, and analysis of variance collapse and alignment effects across different models.

Novelty

This research is the first to validate the cross-cultural transferability of silicon sampling in a non-WEIRD context and provides a more rigorous test of LLMs' individual-level predictive capabilities through the cross-survey transfer framework.

Limitations

  • LLMs perform poorly on sovereignty issues, possibly due to safety alignment influences.
  • The study only uses open-weight models, not evaluating closed-source models.
  • 0-10 scale items are disadvantaged in exact-match evaluation.

Future Work

Future work could explore fine-tuning on target population data, richer prompting strategies, and cross-cultural replication in other non-WEIRD contexts.

AI Executive Summary

Silicon sampling, using large language models (LLMs) to simulate human survey respondents, offers a novel approach to traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level predictions. This paper proposes a cross-survey transfer framework, where an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, the study evaluates three open-weight LLMs and supervised machine learning baselines. Results show that zero-shot LLMs achieve 52% accuracy on unseen items, nearing the performance of supervised random forest models. The study also reveals significant differences in predictability across constructs, such as partisan attitudes and sovereignty issues, and variance collapse and safety alignment effects across different models.

These findings clarify both the promise and boundaries of silicon sampling, providing policymakers with a new tool for rapidly assessing public opinion, especially when traditional methods are costly and slow. However, the study also highlights LLMs' poor performance on sovereignty issues, possibly influenced by safety alignment. Future work could explore fine-tuning on target population data, richer prompting strategies, and cross-cultural replication in other non-WEIRD contexts.

In conclusion, the cross-survey transfer framework offers a more rigorous evaluation method for silicon sampling, revealing both the potential and limitations of LLMs in individual-level prediction and providing new insights and directions for future survey research.

Deep Analysis

Background

Silicon sampling uses large language models to simulate human survey respondents, offering a novel approach to traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level predictions. Existing studies focus mainly on English-language surveys from WEIRD populations, lacking validation in other cultural and linguistic contexts.

Core Problem

Traditional survey research faces declining response rates and rising costs, while policymakers and researchers need rapid, fine-grained public opinion measurements. Silicon sampling offers a potential solution, but its individual-level predictive ability remains underexplored.

Innovation

This paper proposes a cross-survey transfer framework, testing LLMs' individual-level predictive ability by partitioning survey items into context and prediction sets. This method not only validates silicon sampling's cross-cultural transferability in a non-WEIRD context but also provides a more rigorous test of LLMs' individual-level predictive capabilities.

Methodology

  • �� Use Taiwan Election and Democratization Study (TEDS) 2024 data
  • �� Partition survey items into context (Set A) and prediction (Set B) sets
  • �� Evaluate three open-weight LLMs and supervised machine learning baselines
  • �� Analyze variance collapse and safety alignment effects

Experiments

The experiment uses TEDS 2024 data, including 5 demographic variables, 28 context items, and 14 prediction items. Three open-weight LLMs and two supervised machine learning baselines (logistic regression and random forest) were evaluated, with alignment removal experiments conducted.

Results

Zero-shot LLMs achieved 52% accuracy on unseen items, nearing the performance of supervised random forest models. The study also reveals significant differences in predictability across constructs, such as partisan attitudes and sovereignty issues, and variance collapse and safety alignment effects across different models.

Applications

Silicon sampling can be used for rapid public opinion assessment, especially when traditional methods are costly and slow. Future exploration could involve combining LLM-generated priors with smaller human samples.

Limitations & Outlook

LLMs perform poorly on sovereignty issues, possibly due to safety alignment influences. The study only uses open-weight models, not evaluating closed-source models. 0-10 scale items are disadvantaged in exact-match evaluation.

Plain Language Accessible to non-experts

Imagine you have a smart assistant that can predict your opinions on new questions based on your previous answers. It's like telling the assistant your favorite foods and movies, then it guesses the books you might like. This assistant uses a technology called large language models (LLMs), which learn from vast amounts of data to mimic human thinking. In this study, researchers tested how well this assistant performs in different cultural contexts, finding it predicts well on some issues like political attitudes but poorly on others like sovereignty. It's like the assistant guessing your book preferences but sometimes getting it wrong due to lack of information. The study also found that this assistant can be overly cautious, leading to less diverse predictions.

ELI14 Explained like you're 14

Imagine you have a super smart robot friend that can guess your opinions on new questions based on your previous answers. Like, you tell it your favorite foods and movies, and it tries to guess the books you might like. That's what large language models (LLMs) do. Researchers wanted to see how this robot performs in different cultures. Turns out, it's great at some things, like political attitudes, but not so good at others, like sovereignty. It's like your robot friend guessing your book preferences but messing up sometimes because it doesn't have enough info. The study also found that this robot can be too cautious, making its guesses less varied.

Glossary

Silicon Sampling

Using large language models to simulate human survey respondents.

Used for rapid public opinion assessment.

Large Language Model (LLM)

A model that learns from vast data to mimic human language behavior.

Used to simulate survey respondents.

Cross-Survey Transfer

Testing model prediction ability by partitioning survey items into context and prediction sets.

Used to evaluate LLMs' individual-level prediction ability.

Variance Collapse

Model predictions are narrower in distribution than real data.

Observed in both LLMs and supervised models.

Safety Alignment

Mechanism to prevent models from generating inappropriate or biased content.

Affects LLM prediction performance.

Open Questions Unanswered questions from this research

  • 1 How to enhance LLM distribution diversity without affecting prediction accuracy?
  • 2 How to optimize LLM prediction ability in non-WEIRD contexts?

Applications

Immediate Applications

Rapid Public Opinion Assessment

Policymakers can use LLMs to quickly assess public opinion, especially when traditional methods are costly.

Long-term Vision

Cross-Cultural Survey Research

Improving LLM cross-cultural transferability for broader global survey research.

Abstract

Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.

cs.AI cs.CL cs.CY cs.MA stat.ME