Caption Bottleneck Models
CaBM replaces fixed concept sets with natural language captions, enabling leakage-free, interpretable image recognition with competitive accuracy.
Seref Baris Cagliyan, Umut Ozdemir, Merve Tapli et al.
CaBM replaces fixed concept sets with natural language captions, enabling leakage-free, interpretable image recognition with competitive accuracy.
Seref Baris Cagliyan, Umut Ozdemir, Merve Tapli et al.
DCCD combines document and token confidence to improve conflict resolution in multi-document QA, outperforming baselines on DRQA.
Raymond Li, Md Tawkat Islam Khondaker, Amirhossein Abaskohi et al.
Active-GRPO combines active imitate-reinforce and dynamic referencing, boosting SR×Sim from 0.0959 to 0.1773 in molecular optimization.
Xuefeng Liu, Mingxuan Cao, Qinan Huang et al.
PlanRAG uses logical query trees to solve exploratory reasoning problems, improving accuracy.
Ganlin Xu, Linghao Zhang, Zhitao Yin et al.
HOMER employs hierarchical memory and agentic reasoning, achieving +10.8 performance on long video benchmarks.
Yixin Ji, Fanghua Ye, Juntao Li et al.
AMVL framework resolves train-inference mismatch via bidirectional calibration, improving BLINK benchmark average score by +10.83.
Shijie Li, Yilin Gao, Siyuan Yang et al.
DriveVer is a lightweight trajectory verifier that enhances planning models on the NAVSIM benchmark.
Chong He, Yuechen Luo, Fang Li et al.
Vitality-guided compression reduces 3D Diffusion Transformer models by 66%, maintaining geometric fidelity while significantly decreasing size.
Jaeah Lee, Hyunjin Kim, Jaewoong Cho et al.
MEPA introduces a scale-aware MoE framework with semantic guidance, achieving faster training and superior image quality (FID 2.32 vs. 4.10).
Nuoyan Zhou, Zhijun Tu, Lei Yu et al.
CORGI employs CDOG, CA-3DGS, and DCGR to reconstruct high-fidelity 3D dogs from a single image without supervision.
Yuxiao Wu, Weile Li, Boyi Zhu et al.
Proposes adaptive perturbation selection via structured audio transformations, boosting contrastive decoding accuracy to 81.4% on temporal tasks.
Aaron Isidore Grace, Zhouyuan Huo, Weiran Wang
ALEE framework uses AMR-based minimal pairs to evaluate 275+ languages' embedding models, revealing performance gaps related to resources and linguistic phenomena.
Andrianos Michail, Stylianos Psychias, Michelle Wastl et al.
PointSplat predicts Gaussian primitives directly from point clouds, enabling efficient real-time 3D human reconstruction with high quality.
Yujie Guo, Yudong Jin, Lingteng Qiu et al.
ERA introduces entropy-guided visual token pruning with bias rectification, effectively addressing attention logit collapse and boosting inference efficiency in multimodal large models.
Yuhao Wang, Mu Qiao, Haiwen Diao et al.
Introduces signed-permutation coordinate transport for RMSNorm transformers, recovering 91.1% of coordinates.
John Sweeney
InstanceControl generates complex images without instance labeling, enhancing precision and control.
Xiaoyu Liu, Huan Wang, Fan Li et al.
CoDex framework achieves complex dexterous functional manipulation without demonstrations, with a 73% success rate.
Bowen Jiang, William Painter Reger, Roberto Martin-Martin
MemLearner employs a learned query mechanism leveraging pre-trained visual priors to enhance scene memory in video world models, significantly improving scene consistency under occlusion and dynamic scenarios.
Jiwen Yu, Jianxiong Gao, Jianhong Bai et al.
Introduces WorldRoamBench, a benchmark for long-horizon stability of interactive world models, evaluating per-frame action accuracy, visual drift, physics, and memory over 600+ cases.
Ting-Bing Xu, Jiacheng Sui, Zhe Gao et al.
ECHO achieves 43.4% accuracy in Agentic RL using selective turn memory, outperforming GRPO and SUPO.
Zijun Xie, Binbin Zheng, Enlei Gong et al.