MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

TL;DR

Introduces MultivationBench, a benchmark based on Maslow and Reiss models, to evaluate multimodal sequential motivation reasoning; models perform poorly.

cs.AI 🔴 Advanced 2026-07-29 44 views
Kawai Chung Chunkit Chan Yauwai Yim Yuxuan Liu Haochen Shi Weiqi Wang Qing Zong Tianshi Zheng Yixuan Fu Kai Chung Wong Hao Liang Yifan Gao Xi Yang Janet Hui-wen Hsiao Yangqiu Song
multimodal large models motivation reasoning psychological frameworks benchmarking sequential consistency

Key Findings

Methodology

This study develops MultivationBench, grounded in Maslow’s hierarchy and Reiss’s 16 desires, combining visual and textual data. It involves behavior extraction, multi-stage candidate generation, and human verification to ensure psychological validity. The dataset includes 1000 stories, 4023 behaviors, and 16092 questions. Models like Gemini-3-Flash and Grok-4.1-Fast are evaluated in zero-shot settings using EM and F1 metrics, focusing on dynamic motivation inference across story sequences.

Key Results

  • All models show limited performance, with average EM below 20% and F1 around 50%. Performance drops significantly on longer stories, with EM as low as 6%. Multimodal inputs outperform unimodal, especially visual cues for motivation correction. Fine-grained Reiss-based reasoning is more challenging than broad Maslow categories, indicating the difficulty in modeling complex psychological states.
  • Short stories yield better results; performance declines with story length, highlighting the challenge of continuous motivation tracking. Different models excel in definition tasks but struggle with practical, context-specific motivation inference. The gap between static recognition and dynamic reasoning remains large.
  • Humans outperform models significantly, with EM over 60% and F1 near 73%. Models need improvements in multimodal fusion, long-sequence reasoning, and fine-grained psychological state modeling. The findings reveal substantial room for advancement in enabling models to perform human-like social reasoning over extended narratives.

Significance

This work pioneers a systematic evaluation of multimodal models’ capacity for continuous psychological motivation inference, addressing a core challenge in social AI. By integrating psychological theories, it enhances understanding of human behavior behind actions, facilitating more natural human-computer interactions and social robotics. The benchmark exposes current limitations, guiding future research toward models capable of sustained, nuanced social reasoning, crucial for deploying AI in real-world social environments.

Technical Contribution

The paper introduces a novel benchmark combining Maslow’s and Reiss’s models for multimodal, long-sequence motivation reasoning. It employs a multi-stage candidate generation process with human validation, along with a new long-sequence consistency metric. The experimental framework rigorously tests models’ ability to track and revise motivations dynamically, emphasizing the importance of visual cues in refining psychological inference, thus pushing the frontier of social AI evaluation.

Novelty

This is the first work to systematically combine psychological theories with multimodal, long-sequence reasoning benchmarks. Unlike prior static or short-term assessments, it emphasizes continuous, dynamic inference of mental states in realistic narratives, representing a significant step forward in modeling human-like social cognition in AI systems.

Limitations

  • Models still struggle with maintaining motivation consistency over long narratives, indicating limited memory and reasoning capabilities. The dataset’s cultural bias may limit generalization, and current models lack robust multi-step reasoning mechanisms. Further, the zero-shot evaluation setting does not explore potential gains from fine-tuning or reinforcement learning.
  • The complexity of psychological states and subtlety of visual cues pose challenges for current architectures. The benchmark’s scope, while comprehensive, cannot fully capture real-world social variability. Computational costs remain high, and model interpretability in reasoning processes needs enhancement.
  • Future work should focus on integrating memory-augmented architectures, expanding dataset diversity, and developing explainability tools to better understand model reasoning processes.

Future Work

Future directions include integrating reinforcement learning and multi-task training to improve long-term motivation tracking, expanding datasets with diverse cultural contexts, and developing models with better memory and reasoning capabilities. Additionally, exploring explainability techniques will help interpret model decisions, fostering trust and transparency. Ultimately, the goal is to develop AI systems capable of human-like, continuous social understanding, applicable in complex real-world scenarios such as social robotics, mental health support, and personalized virtual agents.

AI Executive Summary

The rapid development of multimodal large language models (MLLMs) has opened new horizons in AI’s social cognition capabilities. However, existing benchmarks primarily evaluate static understanding or short-term reasoning, leaving a significant gap in assessing models’ ability to perform continuous, dynamic motivation inference in complex narratives. Recognizing this challenge, Kawai et al. introduce MultivationBench, a comprehensive benchmark designed to evaluate how well models can infer evolving psychological states over long, multimodal story sequences.

Grounded in well-established psychological frameworks—Maslow’s hierarchy of needs and Reiss’s 16 desires—this benchmark incorporates behaviors, motivation labels, and story context to simulate real-world social scenarios. The dataset comprises 1000 stories, with over 4000 behaviors and 16,000+ questions, enabling detailed analysis of models’ reasoning over extended narratives. The process involves automatic candidate generation, human validation, and multi-stage reasoning tasks, ensuring both scalability and psychological validity.

Experimental results reveal that even state-of-the-art models like Gemini-3-Flash and Grok-4.1-Fast perform poorly in long-sequence motivation reasoning, with average EM scores below 20% and significant performance drops as story length increases. Multimodal inputs improve performance compared to text-only or image-only settings, emphasizing the importance of visual cues in refining motivation inference. Notably, models excel in broad categories like Maslow’s hierarchy but struggle with finer-grained Reiss desires, highlighting the complexity of modeling nuanced psychological states.

This work underscores the gap between static recognition and dynamic reasoning in current AI systems, providing a vital benchmark for future research. Its implications extend to social robotics, virtual assistants, and mental health applications, where understanding human motivation over time is crucial. The authors advocate for integrating advanced memory, multi-task learning, and explainability techniques to bridge these gaps, aiming toward AI that can truly understand and engage in human-like social reasoning over extended interactions.

Deep Analysis

Background

多模态大模型在视觉和语言理解方面取得巨大突破,但在社会认知和心理状态推理方面仍有限。早期研究如SocialIQA、MotiveBench关注静态文本,缺乏动态、多模态的连续性评估。近年来,视频理解和多模态推理逐步发展,但多聚焦短时、事件导向,难以模拟人类在长时间、多场景中的心理状态变化。心理学理论如Maslow层级和Reiss欲望模型,为理解复杂行为提供理论基础,但在多模态场景中的应用尚未系统化。现有基准多偏重静态或短序列,难以评估模型在真实社会场景中的连续推理能力。

Core Problem

当前多模态大模型在连续心理动机推理方面表现不足,主要原因在于模型难以整合多模态信息、追踪长序列中的心理状态变化。静态评估无法反映模型在动态场景中的推理能力,长序列中信息的不断更新和修正对模型提出更高要求。缺乏系统性评估工具限制了对模型能力的全面理解,也阻碍了模型在社会智能中的实际应用。解决这一问题需要设计更贴近人类认知的评估框架,结合心理学理论,强化模型的连续推理能力。

Innovation

本研究创新性地结合Maslow层级和Reiss欲望模型,构建了多模态连续动机推理基准——MultivationBench。区别于以往静态或短序列评估,强调模型在长时间、多模态情境中的心理状态追踪。引入多阶段候选方案生成与人类校验机制,确保标签的心理学合理性。设计了长序列一致性指标,系统性测试模型在复杂场景中的推理能力。实验验证视觉信息在修正动机中的关键作用,推动多模态社会认知研究向更真实、更复杂的场景发展。

Methodology

  • �� 数据采集:结合SSID、StoryReasoning、MovieBench,覆盖多样故事和情境。
  • �� 行为识别:利用多模态模型自动提取角色行为链,标注行为对应的潜在动机。
  • �� 动机方案生成:基于Maslow和Reiss模型,自动生成候选动机,结合人类校验确保合理性。
  • �� 任务设计:包括定义任务(识别行为动机类别)和实际推理任务(在具体情境中选择合适动机),多标签预测。
  • �� 长序列评估:引入连续性指标,衡量模型在整个故事中的动机追踪能力。
  • �� 实验评估:采用零样本评估,比较不同模型(如Gemini-3-Flash、Grok-4.1-Fast)在不同故事长度和模态下的表现。

Experiments

采用1000个故事、4023行为实例,分短中长故事段落。模型包括闭源(如Gemini-3-Flash)和开源(如Llama-4系列),指标为EM和F1。通过不同模态输入(多模态、文本、图像)评估性能。还设计长序列连续性和理论类别识别的专项测试。实验验证多模态信息对动机推理的提升作用,分析模型在细粒度和长序列中的表现差异。

Results

所有模型在连续动机推理中表现有限,EM平均不足20%,F1在50%左右。长故事中性能显著下降,模型难以持续追踪角色心理状态。多模态输入优于单一模态,视觉信息对修正动机尤为关键。模型在细粒度Reiss模型上的表现远低于Maslow层级,反映出复杂心理状态推理的难度。人类标注准确率明显优越,模型仍需在多模态融合和长序列推理方面突破。

Applications

该基准可用于评估和提升多模态模型在社会认知、情感理解、虚拟助手等场景中的能力。未来,结合此评估框架,开发更具连续性和细粒度的模型,有望实现更自然的人机交互、社会机器人等应用,推动智能系统更贴近人类认知。

Limitations & Outlook

模型在长序列中的连续性和心理状态追踪仍不足,未充分解决多模态信息整合和记忆问题。数据集偏向特定文化背景,泛化能力有限。实验主要基于零样本评估,未来需结合微调和强化学习提升性能。模型在细粒度心理状态识别和复杂场景理解方面仍有较大差距。

Abstract

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.

cs.AI