cs.CV 2607.00578

Caption Bottleneck Models

CaBM replaces fixed concept sets with natural language captions, enabling leakage-free, interpretable image recognition with competitive accuracy.

Seref Baris Cagliyan, Umut Ozdemir, Merve Tapli et al.

2026-07-01 45
cs.CV 2606.31734

MemLearner: Learning to Query Context memory for Video World Models

MemLearner employs a learned query mechanism leveraging pre-trained visual priors to enhance scene memory in video world models, significantly improving scene consistency under occlusion and dynamic scenarios.

Jiwen Yu, Jianxiong Gao, Jianhong Bai et al.

2026-06-30 38