A Psycholinguistic Evaluation of Language Models' Sensitivity to Argument Roles
Systematic psycholinguistic evaluation of language models reveals partial sensitivity to argument roles, with models distinguishing plausible from implausible verbs but lacking human-like selective patterns.
Key Findings
Methodology
Using psycholinguistic paradigms, minimal sentence pairs were constructed to control for semantic and syntactic factors. Models (GPT-2, BERT, RoBERTa) were evaluated via surprisal, layer probing, and attention analysis across conditions involving argument swapping, verb changes, and argument replacement. Neural and behavioral data were compared to assess structural sensitivity. This multi-faceted approach allowed detailed analysis of how models encode and utilize argument roles during sentence processing.
Key Results
- Models showed some sensitivity to argument roles, especially in the change-verb condition with surprisal differences exceeding 1.2 bits and accuracy over 85%. In swap-arguments, sensitivity dropped significantly; GPT-2 small was nearly insensitive. Layer probing revealed that middle layers encode argument information more strongly than final layers. Attention analysis indicated correct focus on subjects but limited structural utilization. Overall, models partially capture argument structures but differ from human processing, especially in dynamic integration.
- RoBERTa outperformed GPT-2, demonstrating better structural sensitivity, likely due to bidirectionality. Surprisal and probing results suggest models rely heavily on lexical cues in some conditions, while structural cues are weakly encoded or lost at later layers. These findings highlight the gap between current models and human-like understanding of argument relations, emphasizing the need for explicit structural representations.
- The experiments collectively indicate that although models can distinguish some plausibility differences, their internal mechanisms differ from humans. They tend to depend on surface features rather than deep structural understanding, limiting their ability to generalize across varied sentence structures. This underscores the importance of integrating syntactic modules and multi-modal data to improve structural comprehension in future models.
Significance
This research introduces a novel, systematic framework combining psycholinguistic methods with deep model analysis, providing insights into the structural understanding capabilities of state-of-the-art language models. Demonstrating their partial sensitivity to argument roles, the study highlights both progress and limitations in modeling human sentence processing. The findings inform future directions for improving model architecture, emphasizing the integration of explicit syntactic information and dynamic reasoning. This work bridges cognitive science and AI, offering a pathway toward more human-like natural language understanding systems, with implications for NLP applications such as dialogue, translation, and information extraction. It also advances theoretical understanding of how neural networks encode complex syntactic-semantic relations, fostering interdisciplinary collaboration.
Technical Contribution
The study develops a comprehensive evaluation framework combining surprisal measures, layer-wise probing classifiers, and attention analysis, tailored to assess structural sensitivity in language models. It introduces a novel experimental paradigm inspired by psycholinguistics, enabling fine-grained analysis of argument role encoding. The approach reveals that middle layers encode argument plausibility more robustly, while final layers tend to lose this information, highlighting the importance of intermediate representations. The integration of neural and behavioral metrics offers a new lens for understanding model cognition, setting a foundation for future structural enhancements. This methodology can be extended to evaluate other syntactic phenomena, fostering more cognitively plausible NLP models.
Novelty
This is the first comprehensive application of psycholinguistic experimental paradigms—surprisal, layer probing, and attention analysis—to evaluate the structural sensitivity of multiple large language models. Unlike prior works focusing on lexical or sentence-level metrics, this study isolates argument role effects through minimal pairs, controlling for confounding factors like animacy. It systematically compares unidirectional and bidirectional models, revealing distinct encoding patterns across layers. The multi-method approach offers a nuanced understanding of how models process syntactic relations, providing a new benchmark for future research and a deeper insight into the cognitive plausibility of neural language models.
Limitations
- The experiments are limited to English sentences, and cross-linguistic generalization remains untested. Different languages with varied syntactic structures may pose additional challenges.
- Models show weak sensitivity in swap-arguments conditions, indicating that they do not fully grasp argument structure, especially in complex or nested sentences.
- The reliance on large-scale training data and computational resources limits applicability in resource-constrained settings. Moreover, current models depend heavily on surface cues, lacking robust dynamic structural reasoning.
Future Work
Future research should explore integrating explicit syntactic modules, such as dependency parsers or graph-based representations, into transformer architectures. Multi-modal training incorporating visual or contextual cues could enhance structural understanding. Extending experiments to multiple languages and more complex syntactic phenomena will test model robustness. Combining neural and neurophysiological data may also refine models to better mimic human sentence processing, ultimately leading to more cognitively plausible NLP systems.
AI Executive Summary
This study systematically evaluates large language models’ sensitivity to argument roles using psycholinguistic paradigms. By constructing minimal sentence pairs with controlled manipulations—such as argument swapping, verb changes, and argument replacements—the research probes models’ ability to distinguish plausible from implausible verbs based on argument structure. Surprisal analysis reveals that models like RoBERTa outperform GPT-2 in structural sensitivity, with surprisal differences exceeding 1.2 bits in some conditions. Layer probing shows that middle layers encode argument information more effectively than final layers, which tend to lose this information. Attention analysis confirms that models can correctly identify subjects but do not utilize structural cues dynamically, unlike humans. Overall, models partially capture argument relations but rely heavily on lexical cues, indicating a gap in structural understanding. The findings emphasize the importance of incorporating explicit syntactic representations and multi-modal data to enhance model cognition. Limitations include language scope and complexity of syntactic phenomena, guiding future research toward more cognitively aligned architectures. This work bridges cognitive science and NLP, advancing toward models that process language more like humans do, with broader implications for AI applications requiring deep structural comprehension.
Deep Analysis
Background
Recent advances in deep learning, especially Transformer-based models like BERT, GPT, and RoBERTa, have significantly improved NLP tasks. Psycholinguistic studies have shown that humans process argument roles with a delay, integrating structural information dynamically. Prior research has explored models’ ability to handle syntax, but often focusing on lexical cues or sentence-level metrics. Recent efforts have used probing and surprisal measures to evaluate structural sensitivity, yet a comprehensive understanding remains elusive. This study builds on these foundations, aiming to systematically compare models’ internal representations with human sentence processing, particularly regarding argument roles, a core aspect of syntactic comprehension.
Core Problem
Despite high performance on many NLP benchmarks, current models struggle with understanding and generalizing argument structures, especially in manipulated sentences where roles are reversed or verbs change. This gap affects their interpretability and robustness, limiting real-world applications like translation, question answering, and dialogue systems. The core challenge is whether models encode structural relations explicitly or rely on surface cues, and how this impacts their ability to mimic human-like processing. Addressing this requires detailed, multi-faceted evaluation methods that can dissect internal representations and behavioral responses, revealing the mechanisms underlying structural understanding.
Innovation
The main innovations include: 1) adopting psycholinguistic experimental paradigms—minimal pairs, neural and behavioral metrics—to evaluate structural sensitivity systematically; 2) integrating surprisal, layer probing, and attention analysis to dissect how models encode argument roles across layers; 3) designing controlled manipulations (argument swap, verb change, argument replacement) to isolate structural effects from lexical cues; 4) comparing unidirectional and bidirectional models, revealing differences in internal encoding and processing. This comprehensive framework advances the understanding of how neural models process syntactic relations, bridging cognitive science and AI.
Methodology
- �� Construct minimal sentence pairs with controlled argument manipulations, including swap-arguments, change-verb, and replace-argument conditions. • Compute surprisal values for target verbs, analyzing differences between plausible and implausible contexts. • Use layer-wise probing classifiers trained on verb representations to assess the encoding of argument roles at each layer. • Perform attention analysis to identify whether models correctly focus on subjects and objects, indicating structural recognition. • Compare neural responses with human electrophysiological data (N400 amplitudes) to validate cognitive plausibility. • Evaluate multiple models (GPT-2, BERT, RoBERTa) across different sizes, analyzing how architecture influences structural sensitivity.
Experiments
The experiments involve three key analyses: 1) surprisal measurement to quantify prediction difficulty across conditions, revealing models’ sensitivity to argument plausibility; 2) layer probing to analyze where in the network argument information is encoded, with high accuracy indicating strong internal representation; 3) attention analysis to verify whether models correctly allocate focus to syntactic roles, especially subjects. The datasets are based on psycholinguistic stimuli from Chow et al. (2016) and Kim et al. (2005), with balanced conditions to isolate structural effects. Results are compared across models and layers, providing a comprehensive picture of structural encoding and its limitations.
Results
Models show partial sensitivity: surprisal differences in change-verb conditions exceed 1.2 bits with >85% accuracy, but sensitivity drops sharply in swap-arguments, with GPT-2 small near chance. Layer probing indicates middle layers encode argument roles more robustly than final layers, especially in RoBERTa. Attention analysis confirms correct focus on subjects but limited structural integration. Overall, models rely on lexical cues and surface features, with weaker dynamic structural understanding compared to humans. These findings highlight the need for architecture improvements to better capture syntactic relations.
Applications
The evaluation framework can be used to benchmark and improve NLP systems in tasks requiring deep syntactic understanding, such as machine translation, question answering, and dialogue systems. Incorporating explicit structural modules or multi-modal data could enhance robustness and interpretability. The insights also inform cognitive modeling, aiding the development of AI that mimics human sentence processing more closely. Practical deployment in real-world applications demands models with stronger generalization to complex syntactic phenomena, which this research aims to facilitate.
Limitations & Outlook
The current study focuses on English, limiting cross-linguistic generalization. Models show weak sensitivity in complex or nested structures, and reliance on surface cues restricts their interpretability. Computational costs of probing and attention analysis are high, hindering scalability. Future work should explore integrating explicit syntactic representations, multi-modal training, and cross-lingual evaluations to address these limitations, moving toward more cognitively plausible models.
Plain Language Accessible to non-experts
想象你在厨房做饭,食材代表句子中的词。你需要根据食材的搭配判断菜是否好吃。比如,把番茄和鸡肉放在一起,味道不错,但如果交换位置,可能就不合适。厨师(模型)也是这样,它试图理解句子中词的关系,判断动词是否符合前面提到的角色。虽然它能识别一些合理的搭配,但在复杂的句子中还会迷糊,就像厨师不知道哪个食材应该放在哪里一样。这项研究就像让厨师试试不同的食材组合,看看它是否能正确判断菜的味道,从而了解它对句子结构的理解程度。
ELI14 Explained like you're 14
想象你在玩拼图游戏,每块拼图代表句子中的词。你要把拼图拼在一起,组成一幅完整的画。有时候,你会把一只猫和一只狗的位置互换,画面就变得怪怪的。人们在拼图时,能很快发现哪些交换会让画变奇怪,哪些不会。而这个研究就像是在测试电脑(模型)是否也能像我们一样,判断拼图的合理性。科学家用特别设计的句子,让模型判断动词是否符合前面的角色,就像检测拼图是否拼对了。结果显示,模型在某些情况下能识别出不合理的拼图,但在复杂的情况下还不够聪明,不能像人类一样灵活应对。未来,模型需要学会更多关于句子结构的秘密,才能变得像你一样聪明!
Abstract
We present a systematic evaluation of large language models' sensitivity to argument roles, i.e., who did what to whom, by replicating psycholinguistic studies on human argument role processing. In three experiments, we find that language models are able to distinguish verbs that appear in plausible and implausible contexts, where plausibility is determined through the relation between the verb and its preceding arguments. However, none of the models capture the same selective patterns that human comprehenders exhibit during real-time verb prediction. This indicates that language models' capacity to detect verb plausibility does not arise from the same mechanism that underlies human real-time sentence processing.