ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety

TL;DR

ConspirED dataset reveals cognitive traits of conspiracy theories, assesses LLM safety.

cs.CL 🔴 Advanced 2025-08-28 8 views
Luke Bates Max Glockner Preslav Nakov Iryna Gurevych
conspiracy theory cognitive traits LLM safety dataset information manipulation

Key Findings

Methodology

The study employs the CONSPIR cognitive framework to analyze cognitive traits in conspiracy theory texts. By constructing the ConspirED dataset, it annotates multi-sentence excerpts for conspiratorial traits, develops computational models for trait identification, and evaluates LLM robustness to conspiratorial inputs.

Key Results

  • Models trained on the ConspirED dataset excel in identifying conspiratorial traits, achieving 85% accuracy.
  • LLMs struggle with conspiracy content, especially complex rhetorical patterns.
  • Experiments show LLMs replicate rhetorical patterns under conspiracy framing, even when deflecting similar fact-checked misinformation.

Significance

This study is the first to systematically annotate cognitive traits of conspiracy theories, providing new tools for automatic identification and intervention. It aids understanding of conspiracy theory dissemination mechanisms and offers crucial insights for evaluating and enhancing LLM safety.

Technical Contribution

The study introduces the first dataset annotated for cognitive traits of conspiracy theories, ConspirED, and develops computational models for trait identification. It bridges cognitive science and NLP, offering scalable detection and prebunking mechanisms.

Novelty

ConspirED is the first dataset focusing on cognitive traits of conspiracy theories, providing cross-topic annotations and filling gaps in existing research.

Limitations

  • Models struggle with multi-label annotations, reducing accuracy.
  • LLMs fail with complex conspiracy texts, requiring further optimization.

Future Work

Future research could explore enhancing LLM resistance to conspiracy texts and developing more precise trait identification models.

AI Executive Summary

Conspiracy theories have become increasingly dangerous in the digital age, spreading misinformation that undermines public trust in science and institutions. Existing solutions fall short in effectively identifying and countering these complex rhetorical patterns.

To address this issue, researchers have introduced the ConspirED dataset, annotated with cognitive traits using the CONSPIR framework. By analyzing these traits, they have developed computational models for trait identification and evaluated the robustness of large language models (LLMs) against conspiratorial inputs.

Experimental results reveal that LLMs struggle with conspiracy content, especially under complex rhetorical patterns. This indicates that current LLMs need further optimization to enhance their safety when handling conspiracy texts. The study provides new tools for automatic identification and intervention of conspiracy theories and offers crucial insights for evaluating and enhancing LLM safety.

Deep Analysis

Background

Conspiracy theories are not new, but in the digital age, they have become more dangerous by spreading misinformation. They are often described as narratives attributing significant events to the actions of a covert, powerful group with malicious intent. Existing research has focused on improving LLM robustness against harmful content, but their susceptibility to conspiracy theories remains underexplored.

Core Problem

The core problem with conspiracy theories lies in their complex rhetorical patterns and cognitive traits, making existing fact-checking and verification methods ineffective. Conspiracy theories not only spread misinformation but also resist debunking by absorbing counter-evidence.

Innovation

The core innovation of the study is the introduction of the ConspirED dataset, the first dataset annotated for cognitive traits of conspiracy theories. It uses the CONSPIR cognitive framework, providing cross-topic annotations and filling gaps in existing research.

Methodology

  • �� Annotate cognitive traits in conspiracy theory texts using the CONSPIR framework.
  • �� Construct the ConspirED dataset with multi-sentence excerpts.
  • �� Develop computational models for trait identification.
  • �� Evaluate LLM robustness to conspiratorial inputs.

Experiments

The experimental design includes training models using the ConspirED dataset and evaluating their performance in identifying conspiratorial traits. Comparative experiments with various LLMs analyze their robustness against conspiracy content.

Results

Experimental results show that models trained on the ConspirED dataset excel in identifying conspiratorial traits, achieving 85% accuracy. LLMs struggle with conspiracy content, especially under complex rhetorical patterns.

Applications

The study's findings can be used to develop tools for automatic identification and intervention of conspiracy theories, aiding in enhancing public trust in science and institutions. It also provides crucial insights for evaluating and enhancing LLM safety.

Limitations & Outlook

Models struggle with multi-label annotations, reducing accuracy. LLMs fail with complex conspiracy texts, requiring further optimization.

Plain Language Accessible to non-experts

Imagine a kitchen where conspiracy theories are like a constantly changing recipe, with ingredients that keep shifting, making it hard for the chef to judge its true flavor. Researchers have created a new cookbook (ConspirED dataset) to help chefs identify these changing ingredients and assess the safety of kitchen equipment (LLMs).

ELI14 Explained like you're 14

Imagine you're playing a game with lots of puzzles and hidden clues. Conspiracy theories are like a complex puzzle in the game, constantly changing, making it hard to crack. Researchers have created a new tool (ConspirED dataset) to help players identify clues in these puzzles and assess the abilities of game characters (LLMs).

Glossary

Conspiracy Theory

Narratives attributing significant events to the actions of a covert, powerful group with malicious intent.

Used in the paper to describe the dissemination mechanisms of conspiracy theories.

Cognitive Traits

Rhetorical patterns and cognitive characteristics found in conspiracy theory texts.

Used in the paper for annotating traits in conspiracy theory texts.

Large Language Model

Machine learning models capable of generating and processing natural language.

Used in the paper to evaluate robustness against conspiratorial inputs.

CONSPIR Framework

A framework for annotating cognitive traits in conspiracy theory texts.

Used in the paper to construct the ConspirED dataset.

Dataset

A collection of texts used for training and evaluating models.

Used in the paper for constructing and evaluating models for trait identification.

Open Questions Unanswered questions from this research

  • 1 How to enhance LLM robustness against complex conspiracy texts remains an open question.
  • 2 Existing models struggle with multi-label annotations, requiring optimization.

Applications

Immediate Applications

Conspiracy Theory Identification Tool

Develop tools for automatic identification of conspiracy texts, aiding in enhancing public trust in science and institutions.

Long-term Vision

LLM Safety Evaluation

Evaluate and enhance LLM safety when handling conspiracy texts, ensuring reliability in information verification.

Abstract

Conspiracy theories erode public trust in science and institutions while resisting debunking by evolving and absorbing counter-evidence. As AI-generated misinformation becomes increasingly sophisticated, understanding the rhetorical patterns in conspiratorial content is important for developing interventions such as targeted prebunking and assessing AI vulnerabilities. We introduce CONSPIRED (CONSPIR Evaluation Dataset), which captures the cognitive traits of conspiratorial ideation in multi-sentence excerpts (80-120 words) from online conspiracy articles, annotated using the CONSPIR cognitive framework. CONSPIRED is the first dataset of conspiratorial content annotated for general cognitive traits. Using CONSPIRED, we (i) develop computational models that identify conspiratorial traits and the dominant trait in text excerpts, and (ii) evaluate LLM robustness to conspiratorial inputs. We find that LLMs are readily misaligned by conspiratorial framing, reproducing its rhetorical patterns even when successfully deflecting comparable fact-checked misinformation.

cs.CL