Foundation-Model-Based Agents in Industrial Automation: Purposes, Capabilities, and Open Challenges
Foundation-model-based industrial agents leverage large language models for autonomous decision-making, enhancing human interaction (+37%) and uncertainty handling (+35%).
Key Findings
Methodology
Following PRISMA 2020, a systematic review screened 2,341 publications, selecting 88 relevant studies. These were analyzed through structured coding to assess technological maturity, application domains, capabilities, and limitations. A working definition of FM-based industrial agents was established, integrating agent theory and automation standards. Quantitative metrics evaluated improvements in human interaction, uncertainty management, and negotiation. Data from diverse industrial contexts validated the capability profiles and identified key bottlenecks. The analysis included comparison with baseline traditional agents, highlighting capability shifts and persistent challenges.
Key Results
- Most systems (75%) are at TRL 4-6, with only 9.1% deployed in real-world environments. Operational goals focus on user assistance, monitoring, and process optimization, with less emphasis on planning and scheduling. Capabilities show a 37% increase in human interaction and 35% in uncertainty handling, but negotiation capabilities declined by 39%. Major limitations include poor generalization, hallucination, unstable outputs, data scarcity, and inference latency.
- Compared to traditional agents, FM-based systems excel in multimodal perception and adaptive reasoning, especially in complex environments. However, high inference latency and data dependency hinder real-time applications. The ability to handle dynamic disturbances and integrate external knowledge sources remains a challenge, requiring further research into model robustness and efficiency.
- Future work should focus on improving model generalization, reducing inference latency, and enhancing safety and explainability. Cross-domain adaptation, lightweight architectures, and standardized evaluation metrics are key directions to accelerate industrial adoption.
Significance
This review consolidates the current state of FM-based industrial agents, emphasizing their potential to revolutionize manufacturing, energy, and logistics. By quantifying capability improvements and limitations, it provides a roadmap for industry practitioners and researchers. The integration of large language models into industrial workflows addresses long-standing challenges in human-machine collaboration, decision support, and process automation. Establishing a clear framework and evaluation standards will facilitate broader adoption, fostering smarter, more autonomous industrial systems. This work bridges academic innovation and practical deployment, contributing to the evolution of Industry 4.0.
Technical Contribution
The paper introduces a formal definition of FM-based industrial agents, combining agent theory with automation standards. It develops a capability assessment framework that quantifies improvements in interaction and uncertainty handling, grounded in specific metrics. The integration of retrieval-augmented generation and multimodal perception represents a significant advancement over prior rule-based or purely symbolic approaches. The comparative analysis with traditional agents highlights the transformative potential of foundation models, providing a systematic basis for future research and standardization efforts.
Novelty
This is the first comprehensive systematic review explicitly focusing on FM-based agents in industrial contexts. It uniquely combines capability profiling, maturity assessment, and limitations analysis within a unified framework. The work bridges the conceptual gap between classical agent theories and modern foundation models, offering a standardized definition and evaluation approach. Its novelty lies in quantifying capability shifts across domains and identifying persistent bottlenecks, thus setting a foundation for future innovations and industry standards.
Limitations
- Despite significant improvements, models still struggle with generalization in highly variable environments, leading to inconsistent outputs and decision reliability.
- Inference latency remains high, limiting real-time control applications, especially in high-frequency scenarios.
- Dependence on large datasets and high computational costs pose barriers to widespread deployment and continuous learning.
- Safety, robustness, and explainability are underdeveloped, raising concerns about deployment in safety-critical systems.
Future Work
Future research should prioritize enhancing model robustness and generalization, optimizing inference speed, and developing lightweight architectures. Establishing industry-wide evaluation benchmarks and safety standards is crucial. Cross-domain adaptation and multimodal integration will further expand applicability. Advances in explainability and safety mechanisms are essential for deployment in critical sectors. Collaboration between academia and industry will accelerate these developments, shaping the next generation of autonomous industrial systems.
AI Executive Summary
The landscape of industrial automation is undergoing a transformative shift driven by foundation models, especially large language models (LLMs). These models, exemplified by GPT-4 and LLaMA, are increasingly embedded into intelligent agents that support decision-making, process monitoring, and engineering automation. Unlike traditional rule-based systems, FM-based agents interpret unstructured data, engage in natural language conversations, and orchestrate heterogeneous tools through flexible reasoning mechanisms.
This review systematically analyzed 88 recent studies, revealing that most systems are still at early to mid-stage technological maturity (TRL 4-6), with only a small fraction reaching deployment. Despite limited deployment, these agents demonstrate substantial gains in human interaction (+37%) and uncertainty management (+35%), addressing key industrial challenges. However, limitations such as hallucination, unstable outputs, and high inference latency persist, constraining real-time applications.
The significance of this work lies in its comprehensive capability profiling and the establishment of a formal definition bridging classical agent concepts with modern foundation models. It highlights the transformative potential of these systems in manufacturing, energy, and logistics, paving the way for smarter, more autonomous industrial ecosystems. Future research directions include improving model generalization, reducing latency, and enhancing safety and explainability, aiming for broader industrial adoption.
Overall, this work provides a crucial roadmap for academia and industry to harness foundation models' capabilities, fostering innovation in Industry 4.0 and beyond. It underscores that while challenges remain, the trajectory points toward increasingly intelligent, flexible, and autonomous industrial systems that will reshape manufacturing and related sectors in the coming decade.
Deep Dive
Applications
What is the real-world impact?
Limitations & Outlook
What gaps remain?
Abstract
Foundation models, particularly large language models, are increasingly integrated into agent architectures for industrial tasks such as decision support, process monitoring, and engineering automation. Yet evidence on their purposes, capabilities, and limitations remains fragmented across domains. This work examines how mature foundation-model-based agent systems are in industrial contexts, how their functional profile differs from conventional agent systems, and which limitations persist. A systematic literature survey following the PRISMA 2020 guideline is presented, screening 2,341 publications and synthesising a corpus of 88 publications through a structured coding scheme. The results show that reported systems are predominantly at prototype and early validation stages (75.0% at TRL 4-6), with deployment-oriented evidence remaining rare (9.1%). Operational goals are most frequently positioned in user assistance, monitoring, and process optimisation, while conventional production-control purposes such as planning and scheduling are less prominent. Compared with an established baseline for industrial agent systems, the capability profile reveals substantial gains in human interaction (+37%) and dealing with uncertainty (+35%), but a pronounced deficit in negotiation (-39%). The most widely reported limitations concern lack of generalization, hallucination and output instability, data scarcity, and inference latency. A working definition of foundation-model-based industrial agents is also proposed, bridging conventional agent theory, automation-engineering standards, and the foundation-model paradigm.