Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

TL;DR

FlexPension-LLM predicts pension enrollment among China's flexible workers, achieving 0.9316 F1.

cs.CL 🔴 Advanced 2026-09-04 101 views
Yumiao Li Peixin Liu Donglin Di Chen Li Runhuan Feng
large language models social policy pension flexible employment China

Key Findings

Methodology

The paper introduces FlexPension-LLM, leveraging the DKI-RDistill framework to predict pension enrollment among flexible workers in China. It integrates policy-grounded cues and error-filtered supervision, validated on the CHFS 2019 dataset, and fine-tuned using LoRA/SFT to generate structured decision traces.

Key Results

  • FlexPension-LLM achieved a Composite F1 of 0.9316 on the CHFS 2019 blind test, outperforming 15 baselines and statistically indistinguishable from Claude Opus 4.6.
  • It averaged a Composite F1 of 0.7549 across four external surveys, demonstrating stable performance.
  • Policy cue injection and error-filtered supervision were key to performance gains.

Significance

This study demonstrates the potential of large language models in social policy assessment, particularly in predicting pension enrollment among flexible workers. It offers a new tool for policy simulation, reducing the high costs of pilot programs and uncertainties of econometric methods.

Technical Contribution

FlexPension-LLM combines policy cue injection and error-filtered supervision, significantly improving prediction accuracy. Compared to existing models, it provides more interpretable decision traces, better simulating individual policy responses.

Novelty

This is the first application of large language models to predict pension enrollment among China's flexible workers, innovatively combining policy cue injection and error-filtered supervision.

Limitations

  • The model may underperform in specific policy contexts, especially when policy rules are complex and variable.
  • Dependence on the dataset may limit applicability to other countries or regions.

Future Work

Future research could extend to other social policy domains, further optimize the model's policy cue injection mechanism, and explore cross-national applicability.

AI Executive Summary

Assessing the impacts of social policy changes is a challenge for policymakers. Traditional econometric methods are unreliable for hypothetical scenarios, while field pilot programs are costly. This paper proposes using large language models (LLMs) as policy assessment tools, specifically for predicting pension enrollment among China's flexible workers. FlexPension-LLM is the first domain-specialized LLM for this task, leveraging the DKI-RDistill framework to enhance prediction accuracy through policy cue injection and error-filtered supervision.

In the CHFS 2019 dataset's blind test, FlexPension-LLM achieved a Composite F1 of 0.9316, outperforming 15 baselines and statistically indistinguishable from Claude Opus 4.6. The model also excelled across four external surveys, averaging a Composite F1 of 0.7549, demonstrating strong cross-survey generalization. Policy cue injection and error-filtered supervision were key to these performance gains.

This study not only showcases the potential of LLMs in social policy assessment but also provides a new tool for policy simulation. Future research could extend to other social policy domains, further optimize the model's policy cue injection mechanism, and explore cross-national applicability.

Deep Analysis

Background

Social policy assessment is a complex field where traditional econometric methods often fall short in predicting the impacts of policy changes. With the advent of big data and machine learning, there is growing interest in leveraging these technologies to improve the accuracy and efficiency of policy assessments.

Core Problem

Predicting pension enrollment among flexible workers is challenging due to the decision-making process being influenced by multiple factors, including economic status, household context, historical participation, and hukou policy.

Innovation

FlexPension-LLM innovatively combines the DKI-RDistill framework, applying policy cue injection and error-filtered supervision to pension enrollment prediction. This approach not only improves prediction accuracy but also provides more interpretable decision traces.

Methodology

  • �� Trained and validated using the CHFS 2019 dataset.
  • �� Injected policy cues through the DKI-RDistill framework.
  • �� Fine-tuned the model using LoRA/SFT.
  • �� Generated structured decision traces.

Experiments

Experiments were conducted using the CHFS 2019 dataset for blind testing and validated across four external surveys. The model achieved a Composite F1 of 0.9316 in blind testing and demonstrated stable performance across external surveys.

Results

FlexPension-LLM outperformed 15 baselines in blind testing and was statistically indistinguishable from Claude Opus 4.6. It averaged a Composite F1 of 0.7549 across external surveys, showing strong cross-survey generalization.

Applications

The model can be used to predict pension enrollment among flexible workers, helping policymakers better understand the potential impacts of policy changes and conduct more effective policy simulations.

Limitations & Outlook

The model may underperform in specific policy contexts, particularly when policy rules are complex and variable. Additionally, its dependence on the dataset may limit applicability to other countries or regions.

Plain Language Accessible to non-experts

Imagine a complex jigsaw puzzle where each piece represents a policy factor like income, household context, and hukou policy. FlexPension-LLM acts like a smart assistant that quickly identifies where each piece fits and predicts the final picture. It's like being in a large supermarket where FlexPension-LLM can predict what items shoppers will buy based on their shopping list, budget, and preferences.

ELI14 Explained like you're 14

Imagine you're playing a strategy game where you need to decide whether to join an alliance. FlexPension-LLM is like a super helper in the game, using your resources, friends' advice, and game rules to help you make the best decision. It's like deciding which club to join at school, with FlexPension-LLM giving you the best advice based on your interests, schedule, and school rules.

Glossary

Large Language Model (LLM)

An AI model capable of understanding and generating natural language, often used for complex language tasks.

Used in this paper to predict social policy impacts.

DKI-RDistill

A framework combining policy cue injection and error-filtered supervision to enhance model prediction accuracy.

Used for training the FlexPension-LLM model.

Composite F1

A metric for evaluating model prediction performance, combining precision and recall.

Used to assess FlexPension-LLM's prediction performance.

LoRA/SFT

A technique for model fine-tuning by keeping base weights frozen and training only low-rank adapter parameters.

Used for fine-tuning FlexPension-LLM.

Policy Cue Injection

Injecting policy-related information as input cues to improve prediction accuracy.

Used in the DKI-RDistill framework.

Open Questions Unanswered questions from this research

  • 1 How can similar models be applied in other countries or regions? What specific policy and social factors need consideration?
  • 2 How does the model perform in dynamically changing policy rules? Is a new adaptation mechanism needed?

Applications

Immediate Applications

Policy Simulation

Policymakers can use the model to simulate the impact of different policy changes on pension enrollment among flexible workers, optimizing policy design.

Social Research

Researchers can utilize the model's data analysis capabilities to study the impact of social policies on different populations.

Long-term Vision

Global Applicability

The model could be expanded to other countries, aiding global policymakers in more effective policy assessments.

Abstract

Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometric methods can be unreliable when extrapolating to hypothetical scenarios, while field pilot programs are highly costly. In this paper, we propose using large language models (LLMs) as policy-assessment tools adapted from general-purpose models. We present FlexPension-LLM, the first domain-specialized large language model for a hierarchical pension-enrollment prediction task among flexible workers in China, and introduce DKI-RDistill, which injects policy-grounded cues into the prompt, including Probit-derived marginal effects and hukou-province pension rules. The method then uses LoRA/SFT to distill rationale-augmented supervision into an open-weight MoE student, with teacher errors corrected by regenerating those cases under ground-truth labels. On a CHFS 2019 blind split, FlexPension-LLM achieves 0.9316 Composite F1, surpassing its Claude Sonnet 4.5 teacher and 15 of 17 baselines, and is statistically indistinguishable from Claude Opus 4.6. Across four external surveys, it averages 0.7549 Composite F1 and shows the narrowest performance range among the strongest systems. Component analysis shows that gains come mainly from policy-grounded cue injection and error-filtered supervision, while rationales provide decision traces that can be checked against policy rules.

cs.CL