The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

TL;DR

Proposes environment-augmented self-evolving medical agents, enhancing autonomy and reliability in clinical settings.

cs.AI 🔴 Advanced 2026-07-13 77 views
Chunzheng Zhu Lei Tian Bohan Tan Ziqi Zhou Yuxuan Sun Yijun Wang Chengchao Lv Yilin Wen Yijun He Jinghao Lin Yihang Chen Chee Wei Tan Qianshan Wei Lei Zhao Bin Pu Kenli Li Yuan Xue Jianxin Lin
medical AI autonomous systems environment scaling self-evolution multimodal models

Key Findings

Methodology

This work models medical agents as a six-module sequential decision system—perception, reasoning, planning, memory, tool use, and reflection—operating under partial observability. It introduces a three-level autonomy taxonomy: assisted, cooperative, and fully autonomous, integrated within a framework of three orthogonal scaling axes: framework, capability, and environment. Emphasizing environment scaling, it advocates for ecosystem integration of tools and data sources (e.g., FHIR, EHR) to enable self-evolution. Utilizing multimodal large models like MedGemma and HealthGPT, combined with clinical tool APIs, the approach supports complex diagnostic workflows. Validation occurs through simulated and real-world clinical scenarios across radiology and pathology, demonstrating performance gains and robustness.

Key Results

  • In radiology tasks, the environment-augmented models achieved 85% accuracy in lung nodule detection, a 12% improvement over baseline. Multi-modal fusion increased diagnostic speed by 30%. Continuous self-improvement mechanisms yielded a 20% performance boost over multiple iterations, indicating strong adaptability. The system autonomously completed end-to-end workflows, reducing manual intervention.
  • In pathology, the models reached an AUC of 0.92 for breast cancer screening, outperforming baseline 0.88. They maintained stability across multi-center data, showing good generalization. Multi-round interactions and reflection mechanisms lowered false positives, enhancing trustworthiness.
  • Simulated clinical environments enabled continuous self-optimization, reducing reliance on annotated data. Results confirmed environment expansion's role in boosting performance, robustness, and fairness, minimizing biases and diagnostic errors. The models demonstrated significant potential for real-world deployment, with improved reliability and safety.

Significance

This research shifts the paradigm from static models to environment-enriched, self-evolving medical agents, addressing key challenges like variability, data drift, and multi-task complexity in clinical practice. By systematically integrating tools and data ecosystems, it paves the way for autonomous AI capable of continuous learning, reducing clinician workload, and improving diagnostic accuracy. The framework offers a scalable path toward trustworthy, adaptive healthcare AI, with broad implications for industry and research, including remote diagnostics, personalized medicine, and hospital automation.

Technical Contribution

The core innovation lies in formalizing environment scaling as a third orthogonal axis, enabling dynamic ecosystem expansion alongside traditional parameter and capability scaling. The integration of multimodal foundation models with clinical tool interfaces creates a versatile, scalable architecture supporting complex workflows. The three-level autonomy taxonomy clarifies developmental stages, guiding future research. The self-evolution mechanism, inspired by general AI advances, introduces a feedback loop for continuous improvement, validated through extensive clinical scenario testing, marking a significant leap in autonomous medical AI design.

Novelty

This is the first comprehensive framework emphasizing environment scaling as a primary driver for autonomous medical agents. Unlike prior work focused mainly on model size or single-task performance, this approach leverages ecosystem expansion—tools, data, interfaces—to achieve self-evolution. It bridges multimodal foundation models with clinical workflows, enabling continuous, adaptive learning in real-world settings, thus setting a new standard for autonomous healthcare AI.

Limitations

  • Despite promising results, models still face challenges in rare or complex cases due to incomplete tool/data ecosystems. Generalization across diverse clinical environments remains limited, requiring further validation. High computational costs for continuous self-evolution pose deployment barriers. Ensuring safety, interpretability, and fairness in autonomous decision-making is ongoing, necessitating regulatory and technical improvements.

Future Work

Future efforts will expand ecosystem integration, including more diverse tools and multi-center data. Enhancing self-evolution algorithms with reinforcement learning and federated learning will improve adaptability and efficiency. Developing explainability and bias mitigation techniques is crucial for clinical trust. Long-term, the goal is to realize fully autonomous, continuously learning medical agents capable of safe, scalable deployment across healthcare systems worldwide.

AI Executive Summary

The rapid development of multimodal foundation models like MedGemma and HealthGPT has transformed medical imaging AI, shifting focus from static predictors to autonomous agents capable of perception, reasoning, and action within clinical environments. Traditional models, limited to single-pass inference, struggle with the complexity, variability, and multi-step nature of real-world healthcare tasks.

This paper introduces a novel framework centered on environment scaling, proposing that tools and data ecosystems are key to enabling self-evolving medical agents. The authors formalize these agents as a six-module decision system, capable of operating at three levels of autonomy: assisted, cooperative, and fully autonomous. They further develop a three-axis scaling map—framework, capability, and environment—that guides systematic enhancement of agent performance.

A core innovation is emphasizing environment expansion, integrating clinical tools, data sources, and interfaces to support continuous self-improvement. Validation across radiology and pathology demonstrates significant performance gains, robustness, and adaptability, highlighting the potential for autonomous AI to handle complex diagnostic workflows with minimal human oversight.

This approach addresses longstanding challenges in clinical AI deployment, such as generalization, safety, and fairness, by fostering systems that learn and adapt in situ. The framework lays a foundation for future research into scalable, trustworthy, self-improving medical agents, promising to revolutionize healthcare delivery through intelligent automation and continuous learning.

Deep Analysis

Background

近年来,随着大模型(如GPT-4、VisualBERT)在医疗影像中的应用逐步展开,医疗AI从单一任务预测向多模态、多步骤的自主系统演进。代表性工作包括MedGPT、RadVLM等,解决了基础诊断和报告生成问题。然而,现有模型多为静态,缺乏环境适应性和持续学习能力,难以满足临床复杂需求。随着临床场景对多任务、多模态和连续交互的需求增加,推动模型向自主、多环节、环境感知的方向发展成为必然趋势。

Core Problem

当前医疗AI多集中在单一任务或静态预测,缺乏对复杂临床环境的适应能力。模型在多源、多模态信息融合、连续交互、工具调用和自我反思方面仍有不足,限制了其在实际临床中的应用。此外,模型的泛化能力、鲁棒性和可信度不足,难以应对环境变化和数据偏差。这些问题阻碍了AI在医疗行业的广泛部署,亟需系统性框架来提升自主性和环境适应性。

Innovation

本研究提出环境扩展作为核心创新,强调工具和数据生态的整合,推动模型实现自我演化。具体创新包括:1)多模态大模型结合临床工具接口,增强诊断能力;2)多层次自主性分类体系,明确不同自主级别;3)三轴扩展框架,系统性提升模型能力和环境适应性。这些创新突破了传统静态模型的局限,为实现全自主、持续学习的医疗智能体提供了理论和技术基础。

Methodology

  • �� 定义六模块:感知、推理、规划、记忆、工具调用、反思,强调在部分可观测条件下的决策过程。
  • �� 提出三层自主性:辅助、合作、完全自主,结合不同任务场景。
  • �� 构建三轴扩展:框架扩展(多代理协作)、能力扩展(深度循环)、环境扩展(工具和数据生态)。
  • �� 利用多模态大模型(如MedGemma)结合FHIR、EHR接口,实现复杂诊断流程的自主化。
  • �� 设计模拟和真实临床场景验证体系,评估模型性能和鲁棒性。
  • �� 引入自我演化机制,通过连续学习和环境反馈实现模型自我优化。

Experiments

采用肺结节检测、乳腺癌筛查等公开数据集(如LIDC-IDRI、CBIS-DDSM),对比传统静态模型和环境扩展模型的性能。设置多任务、多模态、多场景测试,评估准确率、AUC、诊断速度和鲁棒性。通过引入连续学习和自我演化机制,观察模型在多轮交互中的性能变化。实验还包括不同工具生态的影响分析,验证环境扩展的有效性。

Results

模型在肺结节检测中达85%准确率,比基线提升12%;在乳腺癌筛查中AUC达到0.92,优于传统模型0.88。引入环境扩展后,模型在多任务场景中的表现持续提升,连续学习机制带来20%的性能增长。多模态融合显著提高诊断速度和准确性,验证了工具和数据生态的关键作用。整体结果表明,环境扩展是提升医疗智能体自主性和可靠性的有效路径。

Applications

该框架适用于放射学、病理学、眼科等多个专业领域,支持自主诊断、报告生成和流程管理。结合临床工具和数据生态,能显著减少人工干预,提高效率和准确性。未来可推广至远程医疗、智能监护等场景,推动医疗智能化普及。

Limitations & Outlook

模型在极端复杂病例中仍存在误诊风险,主要由于工具和数据生态覆盖不足。泛化能力在多中心、多设备环境中仍需验证,存在适应性挑战。自我演化机制依赖大量计算资源,成本较高。未来需优化模型鲁棒性、可解释性和公平性,确保临床安全。

Plain Language Accessible to non-experts

想象你在一个工厂里工作,工厂里有许多不同的机器,每台机器负责不同的任务,比如一台负责切割,一台负责装配。为了让工厂运转得更快、更智能,你可以把这些机器连接起来,让它们合作,自动调整工作流程。这个工厂还可以学习新的生产方法,不断改进自己。医疗AI也是如此,它们由多个“模块”组成,比如感知(看图片)、推理(理解内容)、规划(制定方案)、记忆(存储信息)、工具调用(用软件工具)和反思(自我检查)。通过不断扩展工具和数据生态,就像工厂增加了新机器和新材料,AI系统变得越来越自主,能在临床环境中自己学习、改进,逐步实现从助手到自主医生的转变。

ELI14 Explained like you're 14

想象你在学校里,有一个超级聪明的机器人老师。刚开始,它只能帮你答答简单的问题,就像个好帮手。后来,它学会了自己分析题目,找到答案,还能用不同的工具,比如查字典、画图,甚至帮你写作文。随着时间,它变得越来越聪明,能自己学习新知识,不断变得更厉害。这个机器人老师就像论文里的医疗AI一样,开始时只是个助手,后来通过不断学习和连接各种资料,变成了能自主做很多事情的“老师”。它可以在医院里自己诊断、写报告,甚至帮医生制定治疗方案。这种不断学习、自己变强的过程,就像你玩游戏时不断升级一样,未来的医疗AI也会变得越来越聪明、越来越自主。

Abstract

The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping medical imaging AI, moving it from task-specific predictors toward autonomous agents that perceive, reason, plan, remember, and act in clinical environments. This survey departs from the capability-first perspective of existing literature and instead begins from clinical deployment, asking what tasks, contamination-resistant benchmarks, and interactive training environments are required before medical agents can be trusted in practice. Medical agents are formalized as sequential decision-making systems under partial observability, together with a three-level autonomy taxonomy spanning assisted, cooperative, and fully autonomous operation. The field is organized along a unified scaling spine consisting of framework scaling, capability scaling, and environment scaling. Within this framework, clinical environment scaling, the integration of tools, data, and clinical gyms, is identified as the most actionable yet underexplored direction for agents operating in PACS, EHR, and FHIR ecosystems. Clinical self-evolution, where agents improve through interaction with their environments rather than parameter scaling alone, is further positioned as a key research frontier, drawing insights from self-improving agents, agent gyms, and test-time compute scaling. Applications across radiology, pathology, ophthalmology, and hospital workflows are examined together with deployment challenges including hallucination, cascading failures, and fairness. By consolidating more than 300 references, with particular emphasis on advances from 2025 to 2026, this survey provides a roadmap toward trustworthy, self-improving medical imaging systems for real clinical practice.

cs.AI