MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

TL;DR

MedRealMM evaluates multimodal medical consultations using MCCP framework, featuring 5,620 real cases.

cs.AI 🔴 Advanced 2026-07-10 3 views
Runhan Shi Quan Zhou Yuqian Xu Shuai Yang Xin Wu Zitong Zhou Hui Liu Bin Zha Zheming Wang Liya Li Wei Wei Jinru Ding Wenrao Pang Mouxiao Bian Haoyuan Hu Jun Xu Jie Xu
multimodal medical consultation LLMs dataset safety

Key Findings

Methodology

MedRealMM employs a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically critical moments from real patient-doctor interactions collected from JD Health and convert them into standardized next-response generation tasks. Each instance is paired with a case-specific rubric refined by physicians, rewarding clinically desirable behaviors and penalizing unsafe, unsupported, or contradictory responses.

Key Results

  • Image information is crucial for reliable clinical performance, and current frontier models still face significant bottlenecks in safety-sensitive error avoidance.
  • Some frontier models meet or exceed physicians on positive clinical criteria but trigger more negative criteria.
  • MedRealMM provides a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultations.

Significance

MedRealMM bridges the gap between existing benchmarks and real clinical practice, offering a more realistic evaluation framework. It not only advances research in multimodal medical reasoning but also provides a critical reference standard for future AI medical applications.

Technical Contribution

This study achieves single-turn evaluation instance extraction from real multi-turn multimodal consultations via the MCCP framework and introduces a physician-in-the-loop iterative rubric refinement protocol, combining LLM-initialized criteria with physician revisions for case-specific evaluation grounded in real clinical judgment.

Novelty

MedRealMM is the first benchmark to extract multimodal challenge points from real online medical consultations, distinguishing it from previous studies relying on synthetic dialogues or patient simulators.

Limitations

  • Current models perform poorly in safety-sensitive error avoidance, potentially leading to unsafe medical advice.
  • The dataset's diversity may not cover all possible clinical scenarios.
  • The rubric construction relies on physician feedback, which may introduce subjectivity.

Future Work

Future work may include expanding dataset diversity, improving model safety and robustness, and developing more automated rubric generation methods.

AI Executive Summary

MedRealMM is a real-world multimodal benchmark for Chinese online medical consultation, addressing the misalignment between existing benchmarks and actual clinical practice. Built from de-identified patient-doctor interactions collected from a nationwide internet hospital, MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments and convert them into standardized next-response generation tasks. Each instance is paired with a case-specific rubric refined by physicians, rewarding clinically desirable behaviors and penalizing unsafe, unsupported, or contradictory responses.

The current release contains 5,620 real-world multimodal cases spanning 64 clinical departments. The study evaluates 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Results show that image information is critical for reliable clinical performance, and current frontier models remain below the online physician response.

Although some frontier models satisfy as many or more positive clinical criteria than physicians, they trigger more negative criteria, indicating that safety-sensitive error avoidance remains a central bottleneck. MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation. The dataset will be publicly available on Hugging Face.

Deep Analysis

Background

Recent advances in large language models (LLMs) have made them promising assistants for online medical consultation. However, existing benchmarks often rely on synthetic dialogues or patient simulators, lacking consideration of patient-uploaded medical images, or use multiple-choice or lexical-overlap metrics to evaluate open-ended clinical responses, which poorly reflect clinical quality.

Core Problem

Existing benchmarks poorly align with real clinical practice, often relying on synthetic dialogues or patient simulators, omitting patient-uploaded medical images, or using multiple-choice or lexical-overlap metrics to evaluate open-ended clinical responses, which poorly reflect clinical quality.

Innovation

MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments from real patient-doctor interactions and convert them into standardized next-response generation tasks. Each instance is paired with a case-specific rubric refined by physicians, rewarding clinically desirable behaviors and penalizing unsafe, unsupported, or contradictory responses.

Methodology

  • �� Use MCCP framework to identify clinically demanding moments
  • �� Convert each moment into a standardized next-response generation task
  • �� Pair each instance with a case-specific rubric refined by physicians
  • �� Reward clinically desirable behaviors, penalize unsafe, unsupported, or contradictory responses

Experiments

The study evaluates 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Results show that image information is critical for reliable clinical performance, and current frontier models remain below the online physician response.

Results

Image information is crucial for reliable clinical performance, and current frontier models still face significant bottlenecks in safety-sensitive error avoidance. Some frontier models meet or exceed physicians on positive clinical criteria but trigger more negative criteria.

Applications

MedRealMM provides a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultations. The dataset will be publicly available on Hugging Face.

Limitations & Outlook

Current models perform poorly in safety-sensitive error avoidance, potentially leading to unsafe medical advice. The dataset's diversity may not cover all possible clinical scenarios. The rubric construction relies on physician feedback, which may introduce subjectivity.

Plain Language Accessible to non-experts

Imagine visiting a doctor who not only listens to your symptoms but also examines your reports and photos. MedRealMM acts like a smart assistant, helping doctors better understand this information. It learns from real doctor-patient conversations, identifying critical moments to assist doctors in making better decisions. It's like an experienced tour guide who knows when to alert tourists about important sights and when to explain the background stories.

ELI14 Explained like you're 14

Hey, imagine you're playing a super complex game where doctors are the players, and they need to make decisions based on patient descriptions and photos. MedRealMM is like a game assistant, helping doctors make the right choices at critical moments. It learns from real doctor-patient conversations, just like learning from pro gamers' recordings, helping doctors complete tasks faster and better.

Glossary

Multimodal

Involves multiple forms of information, such as text and images.

Used to describe patient-uploaded text and image information in the paper.

Large Language Model

An AI model capable of understanding and generating natural language.

Used to evaluate online medical consultations.

Clinical Challenge Point

Moments in consultations requiring significant doctor decision-making.

Used to identify moments for focused evaluation.

Rubric

Criteria used to evaluate model performance.

Refined by physicians to evaluate model-generated responses.

De-identification

The process of removing personal identifiable information.

Used to protect patient privacy.

Open Questions Unanswered questions from this research

  • 1 How to improve model performance in safety-sensitive error avoidance?
  • 2 How to expand the dataset to cover more clinical scenarios?

Applications

Immediate Applications

Online Medical Consultation

Helps doctors better understand patient information in online consultations, improving diagnostic accuracy.

Long-term Vision

Intelligent Medical Assistant

Develop into a comprehensive intelligent medical assistant, providing broader medical support.

Abstract

Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clinical practice. Many rely on synthetic conversations or patient simulators, omit patient-uploaded medical images, or evaluate open-ended clinical responses using multiple-choice or lexical-overlap metrics that poorly reflect clinical quality. We introduce \textbf{MedRealMM}, a large-scale benchmark for multimodal online medical consultation built from de-identified patient-doctor interactions collected from a nationwide Chinese internet hospital. MedRealMM uses a Multimodal Clinical Challenge Point (MCCP) extraction framework to identify clinically demanding moments in authentic consultation trajectories and converts each into a standardized next-response generation task while preserving the preceding text-image context. Each instance is paired with a case-specific rubric refined by physicians that rewards clinically desirable behaviors and penalizes unsafe, unsupported, or contradictory responses. The current release contains 5,620 real-world multimodal cases spanning 64 clinical departments. We evaluate 19 general-purpose and medical-specialized LLMs, including text-only and multimodal systems. Our results show that image information is critical for reliable clinical performance and that current frontier models remain below the online physician response. Although some frontier models satisfy as many or more positive clinical criteria than physicians, they trigger more negative criteria, indicating that safety-sensitive error avoidance remains a central bottleneck. MedRealMM offers a realistic and reproducible benchmark for evaluating multimodal medical reasoning in real-world online consultation. The dataset will be publicly available on Hugging Face at https://huggingface.co/datasets/jdh-algo/MedRealMM.

cs.AI cs.CL cs.CV