VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation

TL;DR

VAGS adaptively adjusts guidance scale based on velocity direction similarity, improving image fidelity and semantic control without extra training.

cs.CV 🔴 Advanced 2026-05-15 40 views
Yan Luo Ahmadou Aidara Jingyi Lu Jeremy Moebel Kai Han Mengyu Wang
image generation image editing flow models guidance scale training-free

Key Findings

Methodology

VAGS introduces a training-free, per-step guidance scale modulation mechanism that combines temporal denoising signals with cosine similarity of velocity vectors. It leverages the model's returned velocity vectors during sampling to assess the alignment between current and target directions. This alignment, quantified via cosine similarity, guides the adaptive adjustment of guidance strength, ensuring stronger guidance when velocity directions agree and weaker when they conflict. The adjustment is multiplicative, based on an exponential function of the similarity and denoising progress, allowing smooth, stage-aware control. The approach applies to both image editing—by aligning source and target velocities—and text-to-image generation—by comparing unconditional and conditional velocities—without requiring additional training, auxiliary networks, or extra forward passes.

Key Results

  • On PIE-Bench and DIV2K datasets, VAGS significantly improves structural fidelity, reducing structure distance metrics by approximately 50%, increasing background PSNR by over 20%, and outperforming fixed CFG and recent guidance variants in FID and Inception Score.
  • In text-to-image tasks on COCO17, CUB-200, and Flickr30K, VAGS-Gen reduces FID by about 2.4 points, raises Inception Score by 1.5 points, and maintains semantic consistency, demonstrating robustness across diverse datasets.
  • Ablation studies confirm that step-wise velocity alignment is the key driver of performance gains, outperforming static or monotone schedules, especially in preserving details and preventing semantic drift.

Significance

This work addresses the fundamental limitation of fixed guidance scales in flow-based models, offering a simple yet effective method to adaptively modulate guidance strength based on internal velocity signals. It enhances the fidelity and semantic accuracy of generated images and edits, facilitating more reliable and controllable content creation without additional training or complex tuning. The approach bridges the gap between early noise-dominated and late structure-dominated sampling stages, providing a theoretically grounded, practical solution for high-quality image synthesis. Its simplicity and effectiveness make it highly promising for real-world applications in creative industries, virtual reality, and AI-assisted content generation.

Technical Contribution

VAGS pioneers the use of velocity direction cosine similarity as a real-time, step-wise guidance regulator, combined with temporal denoising signals, to produce a multiplicative, stage-aware guidance scale. This approach departs from traditional fixed or pre-scheduled guidance, offering a theoretically sound, model-internal signal-driven method that requires no training, auxiliary networks, or additional passes. It introduces a unified framework applicable to both editing and generation tasks, leveraging the internal velocity vectors to assess local semantic support and adapt guidance accordingly. This innovation significantly enhances the robustness and controllability of flow-based sampling processes.

Novelty

This study is the first to utilize velocity direction cosine similarity as a dynamic, per-step guidance modulation signal in flow-based image synthesis. Unlike prior methods relying on fixed or monotone schedules, VAGS adaptively responds to the model's internal velocity geometry, providing a principled, real-time adjustment mechanism. Its integration of denoising progress with velocity alignment offers a novel, unified approach to improving both image editing and generation without additional training, marking a significant step forward in guidance control strategies.

Limitations

  • The effectiveness of velocity-based adjustment depends on the accuracy of velocity vectors, which may degrade under high noise or model inaccuracies, potentially leading to suboptimal guidance modulation.
  • Selection of the modulation strength parameter κ requires task-specific tuning; an inappropriate choice can lead to over- or under-guidance, affecting quality.
  • Current implementation does not explicitly handle multi-objective or multi-modal guidance scenarios, which may limit its applicability in more complex tasks requiring diverse control signals.

Future Work

Future research could focus on learning adaptive parameters for κ, integrating multi-modal guidance signals, and extending the framework to multi-objective tasks. Combining VAGS with reinforcement learning or meta-learning could further enhance its robustness and generalization. Additionally, exploring its application in 3D content synthesis or video generation could open new avenues for high-fidelity, controllable content creation.

AI Executive Summary

Flow-based generative models have revolutionized image synthesis, offering high-quality, flexible content creation. However, a persistent challenge lies in the guidance scale used during sampling. Traditionally, a fixed scalar guides the entire process, but this approach neglects the evolving nature of the sampling trajectory. Early steps are dominated by noise with weak semantic signals, requiring gentle guidance to prevent drift, while later stages involve well-structured images demanding stronger, more precise guidance. This mismatch often results in semantic drift, structural distortion, or loss of fine details.

To address this, the authors propose Velocity-Adaptive Guidance Scale (VAGS), a simple yet powerful, training-free method that dynamically adjusts guidance based on internal velocity signals. By leveraging the velocity vectors returned by flow models, VAGS computes the cosine similarity between current and target velocity directions. This similarity, combined with the denoising progress, forms a multiplicative adjustment factor that modulates the guidance strength at each step. When velocity directions align, guidance is amplified; when they conflict, guidance is dampened. This adaptive mechanism ensures guidance is contextually appropriate, improving both fidelity and semantic accuracy.

Experimental results across multiple datasets demonstrate VAGS’s effectiveness. In image editing tasks on PIE-Bench and DIV2K, it significantly reduces structure distance metrics and enhances background preservation without sacrificing edit strength. In text-to-image generation on COCO17, CUB-200, and Flickr30K, VAGS-Gen improves FID scores by approximately 2.4 points and boosts Inception Scores, outperforming fixed guidance and recent training-free variants. Ablation studies confirm that step-wise velocity alignment is the key to these gains, offering a robust, generalizable solution.

This work advances the field by introducing a principled, signal-driven guidance regulation that adapts in real-time, eliminating the need for additional training or complex scheduling. Its simplicity, efficiency, and broad applicability promise to enhance practical deployment of flow-based models, enabling more precise, high-quality image synthesis and editing. Future directions include parameter learning, multi-modal guidance, and broader application in video and 3D content generation, paving the way for more intelligent, controllable AI content creation systems.

Deep Dive

Plain Language Accessible to non-experts

想象你在厨房里做一道菜,厨师需要不断调整火候和调料用量。刚开始,火还很小,调料也不要放太多,否则菜会糊掉。等到火候合适,调料放得恰到好处,菜就会变得香味十足。这个过程就像模型在生成图片时,早期噪声多,语义信号弱,需要轻微引导;而到后期,结构已定,需要更强引导来细化细节。VAGS就像厨师根据菜的状态不断调整火候,利用“菜的香味”——在这里是velocity方向——判断是否需要加大或减小火力,从而让菜既不糊也不生,做出完美的菜肴。它通过观察模型内部的“味道”变化,实时调节指导力度,确保每一步都恰到好处。这样,生成的图片既有丰富细节,又保持整体结构,就像一道色香味俱佳的菜肴。

ELI14 Explained like you're 14

想象你在玩一款游戏,里面的角色需要不断调整自己的动作来完成任务。刚开始,角色还在探索,动作不太准确,不能用太大力气,否则会偏离目标。等到熟悉了环境,动作变得精准,就需要用更大力气,才能完成细节。VAGS就像这个游戏里的教练,根据角色的动作方向,实时告诉它该用多大力气,确保动作既不太轻也不太猛。它通过观察角色的运动方向,判断是否在正确的轨道上,然后调整指导力度。这样,角色就能更快、更准确地完成任务,图片也能更清晰、更符合预期。这个方法不用额外训练,只需要观察模型内部的运动方向,就能让生成效果变得更棒,就像有个聪明的教练在旁边指导一样!

Abstract

Classifier-free guidance (CFG) is the primary control over how strongly text semantics move a flow-based sampler, yet standard practice holds its scale fixed across the entire ODE trajectory. This is a fundamental mismatch: early steps are noise-dominated and carry weak semantic signal, while late steps commit image structure and demand stronger directional commitment; more critically, the value of any guidance strength depends on whether the guided velocity is consistent with the model's current dynamics or working against them. We propose \textit{Velocity-Adaptive Guidance Scale} (VAGS), a training-free replacement that multiplies the nominal scale by a bounded factor combining a temporal signal-level term with the cosine similarity between task-relevant velocity fields. For inversion-free editing, VAGS measures the alignment between source- and target-guided velocities, so edit strength at each step reflects local compatibility between preservation and transformation. For generation, VAGS-Gen uses the alignment between unconditional and conditional velocities as the analogous signal. Neither variant requires fine-tuning, auxiliary networks, or extra forward passes, and fixed CFG is recovered as a special case. On PIE-Bench and DIV2K for editing, and COCO17, CUB-200, and Flickr30K for generation, VAGS consistently improves structural fidelity and generation quality over fixed CFG and recent training-free guidance variants. The code is publicly available at https://github.com/Harvard-AI-and-Robotics-Lab/Velocity_Adaptive_Guidance_Scale.

cs.CV cs.AI