Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

TL;DR

Diffusion TV offers a tangible experience of diffusion models using a modified CRT TV, letting audiences simulate the denoising process.

cs.HC 🟡 Intermediate 2026-09-05 92 views
Sihwa Park
interactive installation AI art generative AI diffusion models explainable AI

Key Findings

Methodology

Diffusion TV uses a modified CRT TV to let audiences experience the denoising process of diffusion models through physical manipulation of antennas and knobs. The installation features three channels displaying AI-generated animals from the past, present, and future, allowing audiences to explore intermediate states.

Key Results

  • Audiences control the clarity of AI-generated images by rotating the antenna, simulating the denoising process and enhancing understanding of the generative process.
  • Three channels present animals from different times, allowing reflection on humanity's relationship with the environment.
  • Experiments show that familiarity with CRT TVs affects interaction experience, with older participants navigating more intuitively.

Significance

This research presents a new mode of explainable AI through an art installation, emphasizing embodied interaction and sensory engagement, breaking the limitations of traditional technical explanations. It offers new perspectives on the application of generative AI in cultural production, promoting public understanding and reflection on AI technologies.

Technical Contribution

Diffusion TV treats intermediate states of diffusion models as experiential material, providing a new mode of explainable AI. Through physical interaction, audiences directly engage with the generative process, enhancing perceptual and emotional understanding of AI systems.

Novelty

This project is the first to transform the denoising process of diffusion models into an experiential art installation, combining physical interaction and ecological narratives to offer a novel AI explanation method.

Limitations

  • The installation relies on audiences' familiarity with CRT TVs, which may affect the universality of the interaction experience.
  • The system's generated content is pre-generated, lacking the flexibility of real-time generation.

Future Work

Future research will include more systematic audience studies to assess how embodied interaction shapes public understanding of generative AI systems. Additionally, plans to extend the system to support real-time generation will enhance interaction flexibility.

AI Executive Summary

Diffusion TV is an interactive AI art installation that provides a tangible experience of diffusion models through a modified CRT TV. By manipulating the TV's antenna, audiences control the clarity of AI-generated images, simulating the denoising process. The knob switches between three channels, showcasing animals from the past, present, and future, forming a temporal and ecological narrative.

The installation emphasizes the generative process over final outputs, allowing audiences to explore intermediate states as experiential material. Through continuous audiovisual feedback and physical interaction, Diffusion TV offers a new mode of explainable AI, inviting audiences to explore and reflect on generative technologies.

While the installation does not provide explicit technical explanations, it demonstrates the potential impact of AI technologies through artistic practice. Future research will include more systematic audience studies to assess how embodied interaction shapes public understanding of generative AI systems. Additionally, plans to extend the system to support real-time generation will enhance interaction flexibility.

Deep Analysis

Background

Since the introduction of Denoising Diffusion Probabilistic Models by Ho et al., diffusion-based approaches have become a dominant paradigm for generative modeling. These models frame generation as an iterative transformation from noise to structured data, making the process temporally observable. However, outside research contexts, these intermediate states are rarely exposed to broader audiences.

Core Problem

Traditional explainable AI often relies on 2D screen-based technical explanations, which are difficult for non-experts to understand the complex processes of generative AI. How to make these processes more intuitive and sensory for the public is a challenge.

Innovation

Diffusion TV transforms the denoising process of diffusion models into an experiential art installation, combining physical interaction and ecological narratives to offer a novel AI explanation method. Audiences engage directly with the generative process through antenna and knob manipulation, enhancing perceptual and emotional understanding of AI systems.

Methodology

  • �� Uses a modified CRT TV as an interface, with audiences controlling image clarity through antenna manipulation.
  • �� The knob switches between three channels, showcasing animals from different times.
  • �� Generated content is pre-generated using Stable Diffusion XL and Stable Audio Open, with user interaction captured by a sensor system.

Experiments

The experiment used data from the IUCN Red List and the World Wildlife Fund website to generate images and sounds of extinct and endangered species. Future species were generated using ChatGPT. User interaction was captured through rotary and magnetic encoders, with the system dynamically updating audiovisual output.

Results

Experiments show that familiarity with CRT TVs affects interaction experience, with older participants navigating more intuitively. The antenna-based denoising interaction was generally perceived as intuitive and engaging.

Applications

The installation is suitable for art exhibitions and educational settings, helping the public understand the generative process of AI through interactive experiences. It offers new perspectives on the application of generative AI in cultural production.

Limitations & Outlook

The installation relies on audiences' familiarity with CRT TVs, which may affect the universality of the interaction experience. The generated content is pre-generated, lacking the flexibility of real-time generation. Future research will include more systematic audience studies to assess how embodied interaction shapes public understanding of generative AI systems.

Plain Language Accessible to non-experts

Imagine playing an old-school TV game where you rotate an antenna and adjust a knob to see different animal images. These images go from blurry to clear, like solving a puzzle. Each channel shows animals from different times: extinct animals from the past, endangered animals from the present, and imagined creatures from the future. This process is like watching a documentary about Earth's history, where physical interaction helps you better understand the stories of these animals and the AI-generated process.

ELI14 Explained like you're 14

Imagine playing an old-school TV game where you rotate an antenna and adjust a knob to see different animal images. These images go from blurry to clear, like solving a puzzle. Each channel shows animals from different times: extinct animals from the past, endangered animals from the present, and imagined creatures from the future. This process is like watching a documentary about Earth's history, where physical interaction helps you better understand the stories of these animals and the AI-generated process.

Glossary

Diffusion Model

A generative model that transforms noise into structured data through an iterative denoising process.

Used in the paper to generate animal images and sounds.

Denoising

The process of extracting useful information from noisy data.

Simulated by audiences through antenna manipulation.

CRT TV

An old television technology using cathode-ray tubes to display images.

Used as the interface and conceptual framework for the installation.

Explainable AI

Techniques that make AI systems' decision processes transparent and understandable.

Demonstrated through the art installation to show AI generative processes.

Generative AI

Technology that uses AI to generate new data.

Used to generate animal images and sounds.

Open Questions Unanswered questions from this research

  • 1 How to enhance the universality of the interaction experience without relying on audiences' familiarity with CRT TVs?
  • 2 How to achieve real-time generation to increase the system's flexibility and interactivity?

Applications

Immediate Applications

Art Exhibitions

Showcases AI generative processes through interactive installations, engaging audiences in participation and reflection.

Long-term Vision

Educational Tool

As an educational tool, it helps students understand AI generative processes and ecological issues through interactive experiences.

Abstract

Diffusion TV is an interactive AI art installation that offers a tangible and embodied experience of diffusion models through a modified CRT TV. By physically manipulating the TV's antenna, audiences control the clarity of AI-generated images and sounds, metaphorically enacting the denoising process that underlies diffusion-based generation. Using the tuning knob, participants switch between three channels featuring AI-generated animals from the Past (extinct species), Present (endangered species), and Future (speculative creatures), situating the interaction within a temporal and ecological narrative. Through continuous audiovisual feedback and physical interaction, Diffusion TV foregrounds the generative process over final outputs, allowing audiences to explore intermediate states as experiential material. Rather than providing explicit technical explanation, the work presents an alternative, embodied mode of explainable AI that invites exploratory engagement with and reflection on generative technologies.

cs.HC cs.AI