cs.CV 2608.05149

CoCo-IR: Contextual Composed Image Retrieval

Proposes CoCo-IR and TIE model, leveraging large multimodal models for multi-turn contextual image retrieval; achieves 39.4 mAP@5 (single-turn) and 44.1 R@1 (4-turn).

Shengcao Cao, Tanmaya Shekhar Dabral, Zhongli Ding et al.

2026-08-06 120
cs.CV 2608.05145

Objects as Audio-Visual Modal Sound Fields

AV-MSF combines multi-view images and few impact recordings to physically reconstruct object impact sounds, outperforming physics-based and data-driven baselines.

Zisen Shao, Zihao Wei, Derong Jin et al.

2026-08-06 126
cs.CV 2608.04995

Promptable Animal Pose Tracking Across Species

Leveraging foundation vision models for cross-species animal pose tracking, achieving high accuracy with limited labels.

Le Li, Daniela Ivanova, Nicolas Pugeault

2026-08-06 37