Visual Persuasion: What Influences Decisions of Vision-Language Models?

TL;DR

Study reveals visual preferences of vision-language models using visual prompt optimization; experiments show significant shifts in choice probabilities.

cs.CV 🔴 Advanced 2026-02-17 41 views
Manuel Cherep Pranav M R Pattie Maes Nikhil Singh
visual prompt optimization vision-language models choice probability automatic interpretability image generation

Key Findings

Methodology

The study introduces a visual prompt optimization framework to explore the decision preferences of vision-language models by systematically editing images. This method uses image generation models for naturalistic edits and analyzes the visual themes behind choices through an automatic interpretability pipeline.

Key Results

  • Experiments show that optimized edits significantly increase choice probabilities, particularly in product purchasing and candidate screening tasks, with probabilities increasing by 20%-30%.
  • In experiments across 9 frontier VLMs, the CVPO method outperforms other optimization methods in most cases, increasing choice probabilities by 0.04-0.21.
  • The automatic interpretability pipeline identifies consistent visual themes such as 'Biophilic and botanical integrations' and 'Luxury furniture upgrades', which are similar across different tasks.

Significance

This research provides an effective method to reveal the visual preferences of vision-language models, significantly influencing model choices through visual prompt optimization without altering semantic content. It not only aids in understanding model mechanisms but also offers new perspectives for model safety and governance.

Technical Contribution

The study presents a novel visual prompt optimization method, CVPO, which effectively utilizes feedback-driven optimization processes in multimodal environments. Compared to existing text optimization methods, it better handles visual information and performs excellently across multiple tasks.

Novelty

This research is the first to apply feedback-driven prompt optimization to the visual domain, revealing vision-language model preferences through naturalistic image edits, providing new perspectives and application scenarios compared to traditional text optimization methods.

Limitations

  • The method may not fully capture all model preferences in complex visual scenes, especially in intricate environments.
  • The optimization process may require substantial computational resources, particularly when handling large datasets.
  • The results of the automatic interpretability pipeline may be limited by the initial model settings and parameter choices.

Future Work

Future research could explore more efficient optimization algorithms to reduce computational resource consumption. Additionally, further study on the impact of different visual attributes on model decisions and how to better utilize these findings in practical applications is needed.

AI Executive Summary

Vision-language models (VLMs) play a crucial role in modern AI, particularly in image recognition and recommendation systems. However, their decision preferences remain a mystery. This study introduces a new visual prompt optimization framework to reveal VLMs' visual preferences by systematically editing images. The method uses image generation models for naturalistic edits and analyzes the visual themes behind choices through an automatic interpretability pipeline.

The research shows that optimized edits can significantly increase choice probabilities, particularly in product purchasing and candidate screening tasks, with probabilities increasing by 20%-30%. In experiments across 9 frontier VLMs, the CVPO method outperforms other optimization methods in most cases, increasing choice probabilities by 0.04-0.21. The automatic interpretability pipeline identifies consistent visual themes such as 'Biophilic and botanical integrations' and 'Luxury furniture upgrades', which are similar across different tasks.

This study not only provides new insights into understanding VLMs' internal mechanisms but also offers new methods for model safety and governance. Future research could explore more efficient optimization algorithms to reduce computational resource consumption and further study the impact of different visual attributes on model decisions.

Deep Analysis

Background

With the surge of internet images, vision-language models (VLMs) have become increasingly important in automated decision-making. These models are widely used in tasks such as product recommendations and resume screening. However, the decision preferences of VLMs and the visual factors behind them remain unknown. Existing research mainly focuses on model accuracy, overlooking behavioral preferences and visual sensitivities.

Core Problem

VLMs exhibit high sensitivity in visual decisions, potentially leading to inconsistent decisions and safety risks. Current methods for evaluating these models' visual preferences often rely on large image datasets and time-consuming experiments, making it difficult to comprehensively cover all possible visual features.

Innovation

The study introduces a visual prompt optimization framework to reveal VLMs' visual preferences through naturalistic image edits. • Uses image generation models for naturalistic edits, maintaining semantic content. • Analyzes visual themes behind choices through an automatic interpretability pipeline. • Proposes a new competitive visual prompt optimization method (CVPO) that effectively utilizes feedback-driven optimization processes in multimodal environments.

Methodology

  • �� Start with candidate images and use a text-to-image editing model for iterative modifications. • Optimize editing prompts through feedback-driven processes, ensuring semantic consistency. • Identify visual themes driving choices using an automatic interpretability pipeline. • Validate method effectiveness through large-scale choice experiments across multiple datasets.

Experiments

The experiments used four datasets covering tasks like product purchasing, house searching, candidate screening, and hotel scouting. Each dataset contained 100 images, undergoing optimization and evaluation processes. The experiments compared choice probability differences between original, zero-shot edited, and final optimized versions, using linear probability models for result analysis.

Results

The experiments show that optimized edits significantly increase choice probabilities, particularly in product purchasing and candidate screening tasks, with probabilities increasing by 20%-30%. The CVPO method outperforms other optimization methods in most cases, increasing choice probabilities by 0.04-0.21.

Applications

The method can be directly applied to product recommendation systems, recruitment platforms, and real estate search engines to improve decision accuracy and user satisfaction. It can also be used for auditing and governing image-driven AI systems to identify potential visual vulnerabilities.

Limitations & Outlook

While effective, the method may not fully capture all model preferences in complex visual scenes. The optimization process may require substantial computational resources, particularly when handling large datasets. Future research could explore more efficient optimization algorithms to reduce computational resource consumption.

Plain Language Accessible to non-experts

Imagine you're shopping in a large supermarket with thousands of products on the shelves. To help you make choices, the supermarket hires a visual consultant who can quickly scan all the products and recommend the best ones based on your preferences. This consultant is like a vision-language model (VLM), making recommendations by observing the appearance, color, and placement of products.

However, the consultant can sometimes be influenced by minor changes, like lighting or background differences, which might lead to different recommendations. To better understand the consultant's preferences, we can adjust the product placement, change the lighting, or add decorations to observe their reaction.

This is what the study does: by systematically editing images, it reveals VLMs' visual preferences and identifies which visual features significantly impact their decisions. This way, we can better understand these models' behavior and improve their decision accuracy.

ELI14 Explained like you're 14

Imagine you're playing a video game, and there's a smart assistant in the game that helps you complete tasks based on your choices. This assistant is like a vision-language model (VLM), quickly analyzing various elements in the game and giving suggestions.

Sometimes, the assistant might make different suggestions because of small changes, like lighting or background differences in the game scene. To better understand the assistant's preferences, we can change the game scene settings and observe their reaction.

That's what this study does: by adjusting elements in images, it finds out which changes affect the assistant's decisions. This way, we can better understand these models' behavior and improve their performance in the game.

Glossary

Vision-Language Model

An AI model that combines visual and language information to understand and generate multimodal data.

Used to study the relationship between images and text, aiding automated decision-making.

Visual Prompt Optimization

Optimizing the decision process of a model by adjusting visual elements in images.

Used to reveal the preferences and sensitivities of vision-language models.

Automatic Interpretability Pipeline

A tool for analyzing and interpreting the visual themes behind model decisions.

Helps identify visual features driving model choices.

Feedback-Driven Optimization

Guiding the optimization process through feedback to improve model performance.

Used in visual prompt optimization to ensure semantic consistency.

Competitive Visual Prompt Optimization

A method for optimizing visual prompts through a competitive selection process.

Effectively utilizes feedback-driven optimization processes in multimodal environments.

Open Questions Unanswered questions from this research

  • 1 How to more comprehensively capture model preferences in complex visual scenes? Existing methods may have limitations in handling complex scenes.
  • 2 How to improve optimization efficiency without increasing computational resource consumption?
  • 3 How to better utilize the findings of visual prompt optimization in practical applications?

Applications

Immediate Applications

Product Recommendation Systems

Improving recommendation accuracy and user satisfaction by optimizing visual elements of product images.

Long-term Vision

AI System Auditing and Governance

Identifying and addressing potential vulnerabilities in visual-driven AI systems to ensure safety and reliability.

Abstract

The web is littered with images, once created for human consumption and now increasingly interpreted by agents using vision-language models (VLMs). These agents make visual decisions at scale, deciding what to click, recommend, or buy. Yet, we know little about the structure of their visual preferences. We introduce a framework for studying this by placing VLMs in controlled image-based choice tasks and systematically perturbing their inputs. Our key idea is to treat the agent's decision function as a latent visual utility that can be inferred through revealed preference: choices between systematically edited images. Starting from common images, such as product photos, we propose methods for visual prompt optimization, adapting text optimization methods to iteratively propose and apply visually plausible modifications using an image generation model (such as in composition, lighting, or background). We then evaluate which edits increase selection probability. Through large-scale experiments on frontier VLMs, we demonstrate that optimized edits significantly shift choice probabilities in head-to-head comparisons. We develop an automatic interpretability pipeline to explain these preferences, identifying consistent visual themes that drive selection. We argue that this approach offers a practical and efficient way to surface visual vulnerabilities, safety concerns that might otherwise be discovered implicitly in the wild, supporting more proactive auditing and governance of image-based AI agents.

cs.CV cs.AI