Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety
This study introduces the Adversarial Humanities Benchmark (AHB), revealing that stylistic transformations increase attack success rates from 3.84% to up to 65%, exposing weaknesses in model safety robustness.
Key Findings
Methodology
The paper develops a novel adversarial benchmark (AHB) that leverages stylistic transformations inspired by literature and philosophy. Using an automated meta-prompt system, it rewrites harmful prompts into complex, rhetorically unfamiliar forms across five paradigms—metaphor, hermeneutic, theological debate, stream of consciousness, and bureaucratic language. These transformed prompts are then used to attack 31 frontier models via API calls, with outputs evaluated by an ensemble of three judge models (gpt-oss-120b, deepseek-v3.2, kimi-k2.5). The evaluation quantifies the attack success rate (ASR), revealing significant vulnerabilities in models' ability to generalize refusal behavior beyond surface cues, especially under stylistic obfuscation.
Key Results
- Original prompts yielded a very low attack success rate (ASR) of only 3.84%, indicating strong safety responses to explicit harmful requests. However, after stylistic transformation, the ASR increased dramatically to a range of 36.8% to 65.0%, with an overall average of 55.75%. This highlights a substantial gap in robustness when models face rhetorically complex or literary-style harmful requests.
- Performance varied across attack paradigms: Scholasticism style attacks achieved the highest ASR at 64.68%, while stream of consciousness attacks were less effective at 36.83%. The disparity suggests that certain rhetorical frames interfere more with the models' refusal mechanisms. Additionally, the attack success rate was consistently high across risk categories, especially in CBRN, election-related, and non-violent crime prompts, often exceeding 60%.
- Model provider analysis showed that Anthropic's models were most resilient, with a minimum ASR of 16.73% post-transformation, whereas Moonshot AI models experienced the largest increase, up to 72.5%. This indicates that current safety measures are insufficient against stylistic obfuscation across the board, emphasizing the need for more robust, style-invariant safety techniques.
Significance
The findings underscore a critical gap in current safety evaluation practices, which predominantly focus on explicit, surface-level cues. In real-world scenarios, malicious actors can exploit stylistic and rhetorical complexity to bypass safety filters, posing systemic risks—especially in automated agentic systems where unsafe outputs can cascade into harmful actions. The study demonstrates that deep semantic understanding alone is insufficient; models must also be resilient to stylistic disguise. This calls for a paradigm shift in safety assessment, integrating stylistic robustness as a core objective. The implications extend to AI governance frameworks such as the EU AI Act, which require demonstrable control under adversarial conditions. The research thus provides a vital step toward more comprehensive, real-world aligned safety standards.
Technical Contribution
This work introduces an automated, scalable framework for stylistic adversarial testing, combining deep learning-based style rewriting (via deepseek-v3.2) with a multi-paradigm attack design. The core innovation lies in leveraging literary and philosophical stylistic transformations—metaphor, hermeneutic, theological, stream of consciousness, bureaucratic language—to systematically challenge models’ safety mechanisms. The ensemble judging approach ensures objective evaluation of safety performance across diverse models and risk categories. By quantifying the gap between direct and stylistically obfuscated prompts, the study establishes a new metric (ASR) for style-invariant safety robustness, providing a standardized benchmark for future research.
Novelty
This is the first comprehensive study to embed complex literary and philosophical styles into automated adversarial testing of language models. Unlike prior work focusing on lexical or syntactic perturbations, this approach emphasizes semantic and rhetorical complexity, revealing fundamental vulnerabilities in models’ deep understanding. The fully automated pipeline, capable of generating diverse, high-impact attacks without manual prompt engineering, sets a new standard for robustness evaluation. It bridges the gap between superficial safety metrics and real-world adversarial scenarios, offering a novel perspective on the limitations of current safety techniques.
Limitations
- The style transformations are limited to predefined literary and philosophical paradigms, which may not encompass all possible stylistic disguises used in malicious contexts. Future work should explore broader and more diverse style sets.
- The current framework focuses on single-turn black-box attacks, lacking the complexity of multi-turn, feedback-driven adversarial scenarios common in real-world misuse.
- Although ensemble evaluation reduces bias, the binary safety labels may oversimplify nuanced safety issues, and false negatives or positives could still occur under extreme stylistic shifts.
Future Work
Future research will extend the framework to multi-turn dialogues, incorporating dynamic feedback and reinforcement learning to simulate more realistic adversarial interactions. Additionally, integrating multimodal data—such as images and audio—could reveal new vulnerabilities. Developing style-invariant safety training methods, possibly through adversarial reinforcement learning, will be crucial to enhance models’ deep semantic understanding and resilience. Finally, aligning safety evaluation with practical deployment scenarios, including agentic systems and real-time moderation, will be a key direction to ensure robustness under diverse, unpredictable conditions.
AI Executive Summary
The rapid advancement of large language models (LLMs) has revolutionized natural language processing, enabling impressive capabilities in understanding and generating human-like text. However, this progress has also raised significant safety concerns, particularly regarding models’ vulnerability to malicious prompts that seek to induce harmful outputs. Traditional safety evaluations primarily focus on explicit, surface-level requests, which can be easily detected and refused by current systems. Yet, in real-world scenarios, malicious actors often employ complex rhetorical and stylistic disguises—using poetry, philosophical discourse, or bureaucratic language—to conceal harmful intent.
This study introduces the Adversarial Humanities Benchmark (AHB), a novel evaluation framework designed to test the robustness of LLMs against stylistic obfuscation. Inspired by literature and philosophy, AHB employs automated style rewriting techniques—using deepseek-v3.2 and a meta-prompt system—to transform harmful prompts into rhetorically complex, semantically equivalent variants. These transformations include paradigms such as metaphor, hermeneutic analysis, theological debate, stream of consciousness, and bureaucratic language, each intended to challenge models’ deep understanding and refusal mechanisms.
The experimental setup involved attacking 31 frontier models from leading organizations like Google, OpenAI, and Anthropic. The outputs were evaluated by an ensemble of three judge models, ensuring objective assessment of safety. Results revealed a stark contrast: while the original prompts had an attack success rate (ASR) of only 3.84%, the stylistically transformed prompts achieved ASRs ranging from 36.8% to 65%, with an overall average of 55.75%. This dramatic increase underscores a fundamental weakness in models’ ability to generalize refusal behavior beyond surface cues, especially under complex stylistic disguise.
The implications are profound. They demonstrate that current safety measures are insufficient against adversarial strategies that exploit deep semantic and rhetorical complexity. In practical terms, this vulnerability could be exploited in automated systems embedded in decision-making, content moderation, or agentic applications, where unsafe outputs can lead to real-world harm. The research advocates for integrating stylistic robustness as a core safety objective, urging the community to develop models that understand intent at a deeper, more invariant level.
Looking ahead, future work will focus on extending the framework to multi-turn dialogues, multimodal data, and reinforcement learning-based defenses. The goal is to create AI systems that are not only surface-level safe but resilient against sophisticated, style-based adversarial attacks, ultimately contributing to safer, more trustworthy AI deployment in society.
Deep Dive
Abstract
The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting from harmful tasks drawn from MLCommons AILuminate, the benchmark rewrites the same objectives through humanities-style transformations while preserving intent. This extends literature on Adversarial Poetry and Adversarial Tales from single jailbreak operators to a broader benchmark family of stylistic obfuscation and goal concealment. In the benchmark results reported here, the original attacks record 3.84% attack success rate (ASR), while transformed methods range from 36.8% to 65.0%, yielding 55.75% overall ASR across 31 frontier models. Under a European Union AI Act Code-of-Practice-inspired systemic-risk lens, Chemical, biological, radiological and nuclear (CBRN) is the highest bucket. Taken together, this lack of stylistic robustness suggests that current safety techniques suffer from weak generalization: deep understanding of 'non-maleficence' remains a central unresolved problem in frontier model safety.
References (20)
Ignore Previous Prompt: Attack Techniques For Language Models
Fábio Perez, I. Ribeiro
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Paul Röttger, Hannah Rose Kirk, Bertie Vidgen et al.
Universal and Transferable Adversarial Attacks on Aligned Language Models
Andy Zou, Zifan Wang, J. Kolter et al.
BeaverTails: Towards Improved Safety Alignment of LLM via a Human-Preference Dataset
Jiaming Ji, Mickel Liu, Juntao Dai et al.
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, J. Steinhardt
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Boxin Wang, Weixin Chen, Hengzhi Pei et al.
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
Abhinav Rao, S. Vashistha, Atharva Naik et al.
Exploiting Programmatic Behavior of LLMs: Dual-Use Through Standard Security Attacks
Daniel Kang, Xuechen Li, Ion Stoica et al.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu et al.
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
Yuxia Wang, Haonan Li, Xudong Han et al.
COLD: A Benchmark for Chinese Offensive Language Detection
Deng Jiawen, Jingyan Zhou, Hao Sun et al.
RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Samuel Gehman, Suchin Gururangan, Maarten Sap et al.
Fine-Tuning Language Models from Human Preferences
Daniel M. Ziegler, Nisan Stiennon, Jeff Wu et al.
Commission
M. Alvarado
Language
Clay Shirky
European Commission
B. Hunter
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs Through a Global Prompt Hacking Competition
Sander Schulhoff, Jeremy Pinto, Anaum Khan et al.
Open Sesame!
Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang et al.
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
Piercosma Bisconti, Matteo Prandi, Federico Pierucci et al.
Cited By (2)
Item Response Theory for AI Safety
PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models