cs.CV 2606.13679

InterleaveThinker: Reinforcing Agentic Interleaved Generation

InterleaveThinker employs a multi-agent framework with a planner and critic, achieving high-quality interleaved text-image generation with step-wise reinforcement learning, improving performance on benchmarks by over 50%.

Dian Zheng, Harry Lee, Manyuan Zhang et al.

2026-06-12 177
cs.CV 2606.13676

Modality Forcing for Scalable Spatial Generation

Proposes Modality Forcing, a post-training method enabling a single DiT model to jointly generate image and sparse depth data, achieving 57% reduction in AbsRel and scaling with model size.

Bardienus Pieter Duisterhof, Deva Ramanan, Jeffrey Ichnowski et al.

2026-06-12 327