Distributed Intelligence in the Computing Continuum with Active Inference
Proposes an Active Inference-based distributed AI framework for autonomous service management in the Computing Continuum, achieving over 90% SLO fulfillment.
Key Findings
Methodology
This paper integrates active inference (AIF) with partially observable Markov decision processes (POMDPs) to create autonomous agents managing services within the computing continuum. Each agent models its environment and service level objectives (SLOiD) via generative models, either learned or predefined. Agents minimize variational free energy (EFE) to infer environmental states and select optimal actions, such as resource scaling or data quality adjustments. Experiments demonstrate that tested models reach over 90% SLO fulfillment, while learning models achieve about 80% at deployment. Compared to multi-agent reinforcement learning (MARL), AIF requires no extensive training, enabling immediate deployment and robust adaptation in dynamic environments. Agents communicate and collaborate, enhancing system resilience.
Key Results
- In simulation, AIF agents achieved over 90% SLO fulfillment with tested models, and about 80% with learned models at deployment, outperforming traditional scheduling methods.
- Compared to MARL, AIF showed 70% reduction in training time and better adaptability to environment changes, maintaining high performance during dynamic shifts.
- In offloading scenarios, AIF agents effectively adjusted strategies to environmental variations, maintaining system stability and reducing latency.
Significance
This work advances autonomous management in large-scale, heterogeneous computing environments, addressing the limitations of centralized and static strategies. By leveraging active inference, it offers a scalable, fast-start, and resilient solution for resource orchestration across edge, fog, and cloud layers. The approach enhances system robustness, reduces latency, and improves resource utilization, paving the way for smarter, self-adaptive infrastructures crucial for next-generation IoT, smart cities, and industrial applications. It bridges the gap between neuroscience-inspired AI and practical distributed systems engineering, promising significant impact on both academia and industry.
Technical Contribution
The paper introduces a novel integration of active inference with POMDPs for distributed service management, emphasizing device-aware SLOs and immediate deployability. It departs from traditional reinforcement learning by eliminating the need for extensive training, instead relying on free energy minimization for real-time inference and decision-making. The framework supports both model-based and model-free adaptation, with agents capable of collaborative information exchange, enhancing resilience. This approach provides theoretical guarantees of stability and robustness, and practical pathways for scalable deployment in heterogeneous environments.
Novelty
This is the first work to embed active inference within a distributed service management framework for the computing continuum, introducing device-aware SLOs and immediate operability without large training datasets. Unlike prior MARL-based methods, this approach leverages the free energy principle for continuous, real-time adaptation, offering a fundamentally new paradigm for autonomous resource orchestration in complex, heterogeneous systems.
Limitations
- The current model assumes perfect sensing; real-world scenarios involve sensor noise and partial observability, requiring robustness enhancements.
- Experimental validation is limited to simulated environments; real-world deployment may face communication overhead and scalability challenges.
- The approach's performance under extreme environmental volatility and high-frequency changes needs further investigation, especially regarding computational costs and convergence speed.
Future Work
Future research will focus on integrating robustness against sensing uncertainties, scaling the framework to larger systems, and optimizing computational efficiency. Exploring multi-layered collaboration strategies and deploying in real-world edge devices will be key steps. Additionally, combining active inference with other AI paradigms like deep learning could enhance model expressiveness, enabling broader applicability across diverse industrial and urban scenarios.
AI Executive Summary
The rapid expansion of the Internet of Things (IoT), edge computing, and cloud infrastructures has created a complex landscape where managing vast, heterogeneous resources efficiently is increasingly challenging. Traditional centralized control strategies struggle with scalability, latency, and adaptability, especially as devices vary widely in capabilities and environmental conditions shift unpredictably. To address these issues, this study introduces a novel framework based on active inference (AIF), inspired by neuroscience, to enable autonomous, distributed service management within the computing continuum.
The core idea is to equip each service with an AIF agent that models its environment and service level objectives (SLOiD) through generative models. These agents continuously infer environmental states by minimizing variational free energy (EFE), a principle that aligns with the brain’s way of reducing surprise. By doing so, agents can dynamically adjust actions such as resource scaling, data quality, or offloading, ensuring high service reliability without extensive training. The experimental results demonstrate that tested models achieve over 90% SLO fulfillment, with learned models reaching about 80% at deployment, outperforming traditional reinforcement learning approaches that require long training periods.
Compared to existing methods, the proposed approach offers immediate operability, robustness, and adaptability, making it highly suitable for real-world, dynamic environments. Agents collaborate by exchanging information, further enhancing system resilience. This work paves the way for smarter, self-organizing infrastructures capable of supporting next-generation applications like autonomous vehicles, remote healthcare, and smart cities. Future efforts will focus on robustness, scalability, and real-world deployment challenges, aiming to realize the full potential of active inference in large-scale, heterogeneous computing systems.
Deep Dive
Glossary
Active Inference (主动推理)
一种基于自由能原理的认知模型,用于持续推断环境状态并优化行为。技术上,代理通过最小化预期自由能(EFE)实现自主调节。
本文将其应用于分布式服务管理,提升系统适应性和鲁棒性。
SLOiD (设备感知的服务水平目标)
考虑设备特性和环境上下文的服务性能指标,用于衡量服务是否满足质量要求。技术上,是在设备硬件和环境条件基础上定义的目标。
用于动态监控和调节边缘设备中的服务性能。
POMDP (部分可观测马尔可夫决策过程)
一种模型描述,代理只能观察到部分环境状态,通过贝叶斯推断估算隐藏状态,指导决策。技术上,定义状态空间、动作空间和观测模型。
支撑AIF代理的环境建模和推断机制。
自由能原理 (Free Energy Principle)
一种假设,系统通过最小化预期自由能(EFE)实现对环境的理解和调节。技术上,涉及变分推断和贝叶斯优化。
指导AIF中代理的推断和行为选择。
分布式智能 (Distributed Intelligence)
在分布式系统中,各节点自主感知、推断和决策的能力,增强系统整体的适应性和鲁棒性。
是本文提出的核心架构理念。
Open Questions Unanswered questions from this research
- 1 如何在实际环境中应对传感器噪声和信息不完整,提升AIF代理的鲁棒性仍需深入研究。
- 2 大规模部署时的通信开销和系统扩展性问题尚未充分验证,需在真实场景中测试。
- 3 极端动态环境下,模型的反应速度和收敛性能仍有待优化,尤其在高频变化场景中。
Applications
Immediate Applications
边缘智能调度
支持智能边缘设备自主调节资源和服务质量,提升响应速度和系统鲁棒性,适用于工业自动化和智慧城市。
远程医疗监控
实现医疗设备的自主调节和数据质量控制,确保高可靠性和隐私保护,适合远程健康管理。
Long-term Vision
智能基础设施
推动智能城市和工业互联网的自主调度平台,减少人工干预,提升系统弹性和效率,未来可实现全自动化管理。
Abstract
The Computing Continuum (CC) is an emerging Internet-based computing paradigm that spans from local Internet of Things sensors and constrained edge devices to large-scale cloud data centers. Its goal is to orchestrate a vast array of diverse and distributed computing resources to support the next generation of Internet-based applications. However, the distributed, heterogeneous, and dynamic nature of CC platforms demands distributed intelligence for adaptive and resilient service management. This article introduces a distributed stream processing pipeline as a CC use case, where each service is managed by an Active Inference (AIF) agent. These agents collaborate to fulfill service needs specified by SLOiDs, a term we introduce to denote Service Level Objectives that are aware of its deployed devices, meaning that non-functional requirements must consider the characteristics of the hosting device. We demonstrate how AIF agents can be modeled and deployed alongside distributed services to manage them autonomously. Our experiments show that AIF agents achieve over 90% SLOiD fulfillment when using tested transition models, and around 80% when learning the models during deployment. We compare their performance to a multi-agent reinforcement learning algorithm, finding that while both approaches yield similar results, MARL requires extensive training, whereas AIF agents can operate effectively from the start. Additionally, we evaluate the behavior of AIF agents in offloading scenarios, observing a strong capacity for adaptation. Finally, we outline key research directions to advance AIF integration in CC platforms.