AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

TL;DR

AutoDesign employs a meta-harness optimization framework to recursively improve long-horizon multimodal design quality, achieving a top score of 78.32.

cs.CV 🔴 Advanced 2026-08-14 115 views
Yaxin Luo Haobin Jiang Jialv Zou Xu Huang Wenhao Yan Haodong Li Zhengrong Yue Jing Li Xiaofu Chen Xiaohan Zhao Jiacheng Liu Jiacheng Cui Zhiqiang Shen Xiaotong Li
AI Meta-Learning Multimodal Design Automation Long-Horizon Optimization

Key Findings

Methodology

AutoDesign utilizes a nested feedback loop architecture, where the inner loop is a design system (DesignHarness) responsible for generating and iteratively refining outputs based on critic feedback, and the outer loop is a meta-harness (Meta-Harness) that analyzes multi-round trajectories and evaluation metrics to guide systematic component updates. The inner loop employs a collaborative mechanism between a large language model (e.g., GPT-4) and a critic module, producing candidate artifacts and applying rule-based and visual critic feedback for localized revisions. The outer loop leverages a comprehensive evaluator, combining rule-based checks and vision-language models, to score generated artifacts across multiple dimensions such as faithfulness, aesthetics, and layout. It then synthesizes failure patterns, proposes bounded component updates via Bayesian optimization or reinforcement learning, and accepts only those updates that improve performance on training data without degrading validation results. This recursive process enables the system to evolve a reusable, human-aligned design system capable of end-to-end academic paper-to-poster generation, covering source extraction, content organization, visual layout, and detail refinement.

Key Results

  • AutoDesign achieved a score of 78.32 on PosterBench main track, surpassing the closed-source commercial system Claude Design (71.87) by 7.45 points, demonstrating superior multimodal academic poster generation performance. Incorporating the learned DesignHarness across seven configurations increased the average PosterBench score from 54.99 to 67.39, a +12.4% improvement, validating the robustness and transferability of the approach.
  • In fully autonomous long-horizon runs, AutoDesign completed 253 tool calls and 11 editing turns within 40 minutes at a cost below $3, producing conference-quality posters as judged by human evaluators, confirming its practical efficiency and effectiveness.
  • A system-blind human preference study showed AutoDesign was favored in 64.0% of pairwise comparisons, especially excelling in visual coherence and content density, indicating strong user acceptance and real-world applicability.

Significance

This work advances the field of automated multimodal content creation by framing system design as a recursive meta-optimization problem, enabling continuous self-improvement grounded in human preferences. It addresses longstanding challenges in maintaining source fidelity, dense communication, and visual quality simultaneously, which are critical for practical deployment in scientific, educational, and creative domains. The framework's ability to produce high-quality, editable artifacts with minimal human intervention signifies a major step toward autonomous content generation systems that can adapt to complex, long-term tasks, reducing manual effort and enhancing productivity across industries.

Technical Contribution

The paper introduces a modular, hierarchical system architecture that separates the fixed model parameters from the surrounding system components, enabling targeted, bounded updates driven by multi-dimensional evaluation feedback. The design harness is decomposed into five functional modules, each optimized via a meta-harness that employs Bayesian optimization or reinforcement learning to identify recurrent failure modes and propose component-level updates. The framework incorporates a comprehensive evaluation protocol (PosterBench) that combines rule-based validation, vision-language model judgments, and human preferences. The recursive optimization process ensures the system's continual evolution without catastrophic forgetting, offering a new paradigm for self-improving AI systems in complex multimodal tasks.

Novelty

This research is pioneering in applying meta-harness optimization to multimodal content generation, particularly in the context of long-horizon, multi-step design tasks. Unlike prior work focusing on static models or single-pass optimization, AutoDesign employs a dual-loop recursive framework that iteratively refines the entire system architecture based on multi-dimensional feedback. Its modular decomposition and bounded component updates enable systematic, interpretable improvements, setting a new standard for autonomous, self-evolving AI systems in creative and scientific content production.

Limitations

  • The current system is primarily validated on academic paper-to-poster generation, and its generalization to other multimodal tasks such as video synthesis or interactive design remains untested, potentially limiting its broader applicability.
  • The recursive optimization process involves substantial computational overhead due to multiple trajectory sampling and evaluation steps, which may hinder scalability in larger or more complex tasks.
  • While human-in-the-loop guidance is incorporated, fully autonomous operation still faces challenges in robustness and interpretability, especially in scenarios with ambiguous or conflicting feedback, necessitating further research into explainability and safety.

Future Work

Future directions include extending the meta-harness framework to multi-task and multi-modal scenarios, improving sample efficiency through advanced Bayesian or reinforcement learning techniques, and exploring more scalable architectures. Enhancing interpretability and robustness, especially in real-world deployment, will be critical. Additionally, integrating user feedback more seamlessly and enabling real-time customization could further broaden its applicability. Long-term, the goal is to develop fully autonomous, adaptive content creation systems capable of personalized, high-quality multimodal generation across diverse domains.

AI Executive Summary

In an era where information overload challenges effective communication, the ability to automatically generate high-quality, multimodal content has become a critical research frontier. Traditional content creation tools rely heavily on manual effort, static models, or single-pass optimization, which limit their adaptability and scalability. Recognizing these limitations, the authors introduce AutoDesign, a pioneering framework that leverages a meta-harness optimization paradigm to enable recursive, self-improving content generation systems.

AutoDesign conceptualizes the content creation process as a long-horizon, agentic task, where a system (the design harness) transforms complex multimodal sources—such as scientific papers—into structured, visually appealing artifacts like posters. Unlike conventional systems, AutoDesign employs a hierarchical feedback mechanism: the inner loop involves a collaborative content generator and critic that iteratively refine outputs, while the outer loop analyzes multi-round trajectories to identify systemic failures and guide targeted component updates. This dual-loop architecture ensures the system not only produces high-quality artifacts but also continually improves its design capabilities over time.

The core innovation lies in treating system components—such as context management, layout tools, rendering engines, and evaluation modules—as modular entities that can be boundedly updated through a meta-optimization process. By grounding the evaluation in a comprehensive, multi-dimensional benchmark (PosterBench), which assesses fidelity, coverage, density, visual evidence, layout, readability, and aesthetics, AutoDesign aligns its improvements with human preferences and real-world standards. The framework employs Bayesian optimization and reinforcement learning strategies to propose and accept component updates only when they demonstrably enhance performance, ensuring stability and interpretability.

Experimental results demonstrate the effectiveness of AutoDesign: on PosterBench, it achieved a score of 78.32, surpassing commercial solutions like Claude Design by 7.45 points. It also outperformed baseline configurations across multiple metrics, with an average score increase from 54.99 to 67.39 across seven configurations. Notably, the system can operate fully autonomously, completing a high-quality poster generation cycle in under 40 minutes at a cost below $3, with minimal human intervention. Human preference studies further confirmed its superiority, with 64% of evaluators favoring AutoDesign outputs.

This research marks a significant step toward autonomous, self-improving AI systems capable of complex multimodal content creation. Its modular, hierarchical approach opens avenues for broader applications, including automated video production, personalized content generation, and real-time adaptive design. While current limitations include computational costs and task-specific validation, ongoing work aims to enhance efficiency, robustness, and generalization. Ultimately, AutoDesign paves the way for intelligent systems that can learn, adapt, and produce high-quality content with minimal human oversight, transforming industries from academia to entertainment.

Deep Analysis

Background

The evolution of multimodal content generation has seen rapid advancements with models like GPT-4, Perceiver, and CoCa, which integrate text, images, and other modalities to produce coherent outputs. Early systems focused on single-modality tasks, such as language modeling or image synthesis, but recent efforts aim to unify these capabilities for complex applications like automatic report generation, multimedia summarization, and scientific visualization. Despite progress, existing systems predominantly operate as static models, lacking mechanisms for continuous self-improvement. Prior works such as TextGrad, DSPy, and GEPA have explored component-level optimization, while systems like STOP, GPTSwarm, and AFlow have emphasized workflow search. However, these approaches do not address the recursive, long-term optimization of entire content generation systems grounded in human preferences. The challenge remains in designing systems that can adaptively refine their architecture and strategies over multiple tasks and iterations, especially in high-stakes domains like scientific publishing, where fidelity, clarity, and visual appeal are paramount. The paper situates itself within this landscape, proposing a novel meta-harness framework that enables persistent, recursive system evolution, validated through the academic paper-to-poster task, a representative multimodal design problem.

Core Problem

Traditional content generation pipelines are limited by their static nature, often producing outputs that are either faithful but unrefined or visually appealing but lacking in fidelity. In complex multimodal tasks, this trade-off is exacerbated by the inability to incorporate feedback effectively across multiple stages. The core problem addressed by this work is how to develop a system that can not only generate high-quality artifacts but also learn from its own outputs to improve over time. Specifically, in the context of academic paper-to-poster conversion, the challenge lies in balancing source fidelity, dense scientific communication, visual coherence, and usability of the final artifact. Existing methods lack a systematic mechanism for persistent knowledge accumulation and self-directed improvement, leading to suboptimal performance and manual intervention requirements. The key difficulty is designing a framework that can leverage multi-round feedback, analyze recurrent failures, and implement targeted system updates without overfitting or catastrophic forgetting, thereby enabling long-term autonomous evolution.

Innovation

The primary innovation of AutoDesign is the introduction of a meta-harness optimization framework that treats system design as a recursive, long-horizon problem. Unlike prior work focusing on static models or single-pass optimization, AutoDesign employs a hierarchical feedback loop architecture:

  • �� The inner loop involves a collaborative content generator (designer) and critic, which iteratively produce and refine artifacts based on rule-based and visual feedback.
  • �� The outer loop analyzes multi-round trajectories and evaluation scores to identify systemic failures and propose bounded component updates.
  • �� The updates are generated by a meta-optimizer, which employs Bayesian optimization or reinforcement learning, ensuring performance gains on training data without regressions on validation data.
  • �� The system decomposes the design harness into five modules—context, tools, runtime, orchestration, and evaluation—allowing targeted, interpretable updates.
  • �� The evaluation protocol (PosterBench) comprehensively assesses multiple quality dimensions, aligning system improvements with human preferences.

This approach enables the system to continually learn from its outputs, forming a self-improving loop that surpasses traditional static systems, especially in complex, multi-step design tasks.

Methodology

  • �� The design harness (H) is modular, comprising five functional components: context/memory, tools/specifications, execution runtime, orchestration, and evaluation/feedback.
  • �� The inner loop involves a designer model (e.g., GPT-4) generating an artifact y and a critic providing feedback f, iteratively refining the artifact until a satisfactory version is produced.
  • �� The outer loop executes multiple rounds of system deployment on diverse tasks, collecting trajectories τ and evaluation scores s based on a comprehensive evaluator Rmeta, which combines rule-based checks and vision-language model judgments.
  • �� The optimizer P analyzes these trajectories and scores, identifying recurrent failure modes and proposing bounded component updates H′ based on observed deficiencies.
  • �� Updates are accepted only if they improve performance on the training set without degrading validation results, using an acceptance gate.
  • �� The process repeats over multiple iterations, gradually evolving the design system into a self-improving, human-aligned content generator.
  • �� The entire framework is instantiated in the academic paper-to-poster task, demonstrating end-to-end automation from source ingestion to final visualization.

Experiments

  • �� The evaluation employs PosterBench, a comprehensive benchmark with 100 papers spanning five disciplines and a smaller 10-paper subset for rapid testing.
  • �� Baselines include static systems like Claude Design and various configurations of AutoDesign with different system modules.
  • �� Metrics encompass overall PosterBench score, fidelity, coverage, density, visual evidence, layout, readability, and aesthetics, evaluated through rule-based checks, vision-language models, and human judgments.
  • �� Hyperparameters such as trajectory sampling size, update thresholds, and module-specific learning rates are tuned for stability.
  • �� Ablation studies compare the impact of individual modules and feedback mechanisms, confirming the importance of recursive, bounded updates.
  • �� Results show that AutoDesign consistently outperforms baselines, with the best configuration reaching 78.32 points, a significant margin over existing solutions.

Results

  • �� AutoDesign achieved a PosterBench main track score of 78.32, outperforming Claude Design by 7.45 points, demonstrating its superior ability to generate faithful, dense, and visually appealing posters.
  • �� Across seven configurations, integrating the learned DesignHarness improved scores from 54.99 to 67.39, confirming the robustness of the recursive optimization process.
  • �� Fully autonomous runs completed within 40 minutes, executing 253 tool calls and 11 editing turns at a cost below $3, producing artifacts comparable to human-designed posters.
  • �� Human preference surveys indicated that 64% of evaluators favored AutoDesign outputs, especially in visual coherence and content density, validating its practical effectiveness.

Applications

  • �� Immediate: Researchers can use AutoDesign to automate poster creation, reducing manual effort and accelerating dissemination of scientific results.
  • �� Educational institutions can generate teaching materials, visual summaries, and interactive content with minimal human input.
  • �� Industry sectors like marketing and media can leverage the framework for rapid multimodal content production, enhancing creativity and productivity.
  • �� Long-term: The framework can evolve into a general-purpose autonomous content creation platform, capable of real-time, personalized multimedia generation, transforming creative industries and knowledge dissemination processes.

Limitations & Outlook

  • �� The current validation is limited to academic paper-to-poster tasks; adaptation to other domains like video synthesis or interactive design remains untested.
  • �� The recursive optimization process incurs high computational costs, especially for larger-scale or real-time applications.
  • �� Fully autonomous operation may still face robustness challenges, such as handling ambiguous feedback or conflicting evaluation signals, requiring further research into explainability and safety mechanisms.

Plain Language Accessible to non-experts

Imagine you’re in a kitchen trying to make a fancy dish. Normally, you follow a recipe once, cook, and hope it turns out well. But what if you had a smart chef who not only follows the recipe but also tastes the dish, thinks about what could be better, and then adjusts the ingredients or cooking time? This chef keeps trying different tweaks, remembers what worked, and gradually learns how to make the dish perfect. Over time, it gets better and better without needing you to tell it exactly what to do each time. That’s what AutoDesign does—it's like a clever chef that keeps improving its cooking skills by constantly tasting, adjusting, and learning from its own experience, making the best possible dish every time, all on its own.

ELI14 Explained like you're 14

Imagine you want to make a really cool poster for your school project, but you’re not sure how to start. Usually, you might look at some examples, pick your favorite parts, and then put everything together. Now, think of a robot friend who can do this for you. It looks at your paper, then tries to make a poster. If it doesn’t look quite right, it checks with a judge (like a teacher or a critic), gets feedback, and then changes parts of the poster—maybe moves some pictures around or changes the colors. It keeps doing this over and over, each time making the poster a little better. After a while, it learns what makes a good poster and can make one that looks really professional—all by itself! That’s what AutoDesign does: it’s like a smart robot that learns how to make perfect posters without needing much help from you.

Glossary

Meta-Harness (Meta-Harness)

A system architecture designed to guide the recursive optimization of content generation systems by analyzing multi-round outputs and feedback, enabling continuous self-improvement.

In the paper, Meta-Harness directs the overall system evolution beyond static models.

DesignHarness (Design Harness)

A modular, executable system responsible for transforming multimodal sources into human-facing artifacts through content generation, editing, and rendering modules.

Serves as the core content creation engine in the AutoDesign framework.

PosterBench

A comprehensive evaluation benchmark specifically designed for academic paper-to-poster tasks, covering multiple quality dimensions such as fidelity, aesthetics, and layout.

Used to measure the performance of AutoDesign and baseline systems.

Rollout (Trajectory)

The sequence of actions, states, and revisions produced during the content generation process, used for analysis and system improvement.

Collected during inner-loop iterations for feedback analysis.

Evaluator

A multi-dimensional scoring tool combining rule-based checks and vision-language models to assess the quality of generated artifacts.

Provides the quantitative basis for outer-loop system updates.

Feedback

Critic evaluations that guide localized revisions within the content generation process, based on rules and perceptual judgments.

Critical for iterative refinement in the inner loop.

Meta-Optimization

The process of systematically updating the system architecture or components based on feedback to improve overall performance.

Implemented in AutoDesign via bounded component updates.

Inner Loop

The iterative process of generating and refining content within a fixed system, driven by feedback.

Ensures local artifact quality before system-wide updates.

Outer Loop

The higher-level process analyzing multiple inner-loop trajectories to propose and accept system component updates.

Enables the system's recursive self-improvement.

Autonomous Loop

A feedback cycle where the system independently generates, evaluates, and improves content without human intervention.

Demonstrated in the paper through fully automated poster creation.

Open Questions Unanswered questions from this research

  • 1 未来如何将AutoDesign的元优化机制推广到其他多模态内容生成任务(如视频、交互式内容),仍需验证其泛化能力。当前方法在高复杂度任务中的效率和效果有待提升,特别是在多任务、多目标场景下的协同优化策略仍需探索。此外,系统的可解释性和鲁棒性也是未来研究的重要方向,以确保其在实际应用中的安全性和可靠性。

Applications

Immediate Applications

科研论文海报自动生成

研究人员可以利用AutoDesign快速生成符合会议要求的海报,节省设计时间,提升展示效果,特别适合多学科交叉研究团队。

多模态内容创作助手

内容创作者可以用其自动生成网页、幻灯片、短视频等多模态内容,降低制作难度,提升创作效率。

教育材料自动化制作

教师和教育机构可以借助AutoDesign快速生成教学演示、科普海报等,增强教学互动和趣味性。

Long-term Vision

智能内容创作平台

未来发展出全自动、多模态、个性化的内容生成平台,实现实时定制和优化,推动内容产业数字化转型。

跨领域知识融合系统

结合AutoDesign的递归优化能力,建立跨学科、多模态的知识融合平台,支持科学研究、艺术创作和商业创新的深度结合。

Abstract

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on rollout feedback. To instantiate and evaluate this framework, we focus on the academic paper-to-poster generation task and introduce PosterBench, comprising a 100-paper Main Track spanning five disciplines and PosterBench-mini, a shared 10-paper subset for controlled evaluation. On the PosterBench Main Track, AutoDesign achieves the highest score of 78.32, surpassing the closed-source commercial system Claude Design by 7.45 points. Across seven controlled code-agent-model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%). In a fully autonomous long-horizon loop, it executes 253 tool calls and 11 editing turns within 40 minutes for under $3, reaching average conference-poster quality in human evaluation. A system-blind human study further demonstrates that AutoDesign achieves the highest human preference among evaluated systems.

cs.CV cs.AI cs.CL

References (20)

P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark

Tao Sun, Enhao Pan, Zhengkai Yang et al.

2025 16 citations ⭐ Influential View Analysis →

Self-Improvements in Modern Agentic Systems: A Survey

Zhenjiang Ren, Yimeng Chen, Dandan Guo et al.

2026 4 citations ⭐ Influential View Analysis →

Deep Submodular Optimization and LLM for Multimodal Content Extraction and Automatic Poster Generation from Long Document

Vijay Jaisankar, Sambaran Bandyopadhyay, Kalp Vyas et al.

2025 2 citations ⭐ Influential

Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

Wei Pang, K. Lin, Xiangru Jian et al.

2025 51 citations ⭐ Influential View Analysis →

Workshops

Johanna Amaya

1993 81 citations

Recursive Harness Self-Improvement

Hyunin Lee, Jinglue Xu, Jeffrey Seely et al.

2026 5 citations View Analysis →

PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation

Jiho Choi, Seojeong Park, Seongjong Song et al.

2025 6 citations View Analysis →

Paper2Video: Automatic Video Generation from Scientific Papers

Zeyu Zhu, Kevin Qinghong Lin, M. Shou

2025 25 citations View Analysis →

Huxley-Gödel Machine: Human-Level Coding Agent Development by an Approximation of the Optimal Self-Improving Machine

Wenyi Wang, Piotr Piekos, Nanbo Li et al.

2025 30 citations View Analysis →

Igd: Instructional Graphic Design With Multimodal Layer Generatio

Yadong Qu, Shancheng Fang, Yuxin Wang et al.

2025 6 citations View Analysis →

UICoder: Finetuning Large Language Models to Generate User Interface Code through Automated Feedback

Jason Wu, E. Schoop, Alan Leung et al.

2024 53 citations View Analysis →

Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

Chun Xia, Zhe Wang, Yan Yang et al.

2025 86 citations View Analysis →

DesignAsCode: Bridging Structural Editability and Visual Fidelity in Graphic Design Generation

Ziyuan Liu, Shizhao Sun, Danqing Huang et al.

2026 1 citations View Analysis →

MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems

Qi Cai, Yonggang Zhang, Xianzhang Jia et al.

2026 5 citations View Analysis →

Evolutionary principles in self-referential learning, or on learning how to learn: The meta-meta-. hook

Jürgen Schmidhuber

1987 862 citations

Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering

Chenglei Si, Yanzhe Zhang, Ryan Li et al.

2024 152 citations View Analysis →

Any2Poster: Any-Source Poster Generation Across Modalities and Domains

Amogh Vinaykumar, A. Li, Suozhi Huang et al.

2026 3 citations View Analysis →

Voyager: An Open-Ended Embodied Agent with Large Language Models

Guanzhi Wang, Yuqi Xie, Yunfan Jiang et al.

2023 2128 citations View Analysis →

GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Lakshya A. Agrawal, Shangyin Tan, Dilara Soylu et al.

2025 313 citations View Analysis →

SciPostLayout: A Dataset for Layout Analysis and Layout Generation of Scientific Posters

Shohei Tanaka, Hao Wang, Yoshitaka Ushiku

2024 16 citations View Analysis →