COMIX: Compositional Explanations using Prototypes
COMIX method explains ML model decisions by decomposing images into prototypes, achieving a 48.82% improvement in C-insertion score.
Sarath Sivaprasad, Dmitry Kangin, Plamen Angelov et al.
COMIX method explains ML model decisions by decomposing images into prototypes, achieving a 48.82% improvement in C-insertion score.
Sarath Sivaprasad, Dmitry Kangin, Plamen Angelov et al.
VideoRAG leverages LVLM for dynamic video retrieval and multimodal response generation, outperforming baselines with significant improvements in ROUGE-L and BLEU scores.
Soyeong Jeong, Kangsan Kim, Jinheon Baek et al.
OVO-Bench evaluates Video-LLMs' temporal awareness, highlighting gaps with human understanding.
Yifei Li, Junbo Niu, Ziyang Miao et al.
LongProc benchmarks long-context language models on long procedural generation, revealing current model limitations.
Xi Ye, Fangcong Yin, Yinghui He et al.
Proposes a RT-based search testing framework for Simulink models, achieving 85% fault detection in 60 model-RT combinations.
Federico Formica, Chris George, Shayda Rahmatyan et al.
Search-o1 integrates agentic retrieval and document reasoning, boosting large reasoning models' performance on complex tasks.
Xiaoxi Li, Guanting Dong, Jiajie Jin et al.
Proposes Deep Feature IV (DFIV) achieving minimax optimal rates in nonparametric IV regression, adaptive to complex functions.
Juno Kim, Dimitri Meunier, Arthur Gretton et al.
Introduces URSA framework with large-scale datasets MMathCoT-1M and DualMath-1.1M, combined with PS-GRPO algorithm, significantly enhancing multimodal mathematical reasoning, outperforming GPT-4o by 8.4%.
Ruilin Luo, Zhuofan Zheng, Yifan Wang et al.
Proposed interval-based tokenization for symbolic music, improving model performance and interpretability.
Dinh-Viet-Toan Le, Louis Bigo, Mikaela Keller
Partition cover-based tokenization algorithm GREEDTOK outperforms BPE and Unigram, achieving ~3% better compression on real-world corpora and lower bits per byte in large-scale pretraining.
Jia Peng Lim, Shawn Tan, Davin Choo et al.
Generative AI-based user simulation using large language models enhances behavior modeling, synthetic data generation, and system evaluation, advancing personalization and safety.
Krisztian Balog, ChengXiang Zhai
Proposes a Token-level shuffling and mixing framework for unbiased deepfake detection, significantly improving cross-dataset generalization with state-of-the-art AUC scores.
Xinghe Fu, Zhiyuan Yan, Taiping Yao et al.
RoRA optimizes LoRA's scaling factor to α/√r, significantly improving fine-tuning accuracy on large and pruned models.
Jun Liu, Zhenglun Kong, Peiyan Dong et al.
Agent Laboratory employs multi-agent LLM framework for end-to-end autonomous research, reducing costs by 84% and achieving SOTA ML performance.
Samuel Schmidgall, Yusheng Su, Ze Wang et al.
Sa2VA unifies SAM-2 and MLLM for dense, multi-task image/video understanding, achieving over 15% improvement in complex scene segmentation.
Haobo Yuan, Xiangtai Li, Tao Zhang et al.
This study compares noise injection, multi-task learning, and layer swapping for Norwegian dialect slot and intent detection, achieving up to 97.6% accuracy.
Verena Blaschke, Felicia Körner, Barbara Plank
Using Vaidya's plane cutting method, the paper proposes an optimal accuracy-communication-privacy trade-off algorithm for distributed DP convex optimization.
Sudeep Salgia, Nikola Pavlovic, Yuejie Chi et al.
STAR method enhances video super-resolution using text-to-video models, improving spatio-temporal consistency.
Rui Xie, Yinhong Liu, Penghao Zhou et al.
SceneVTG++ combines TLCG and CLTD to generate realistic, controllable multilingual scene text, surpassing SOTA in fidelity and utility.
Jiawei Liu, Yuanzhi Zhu, Feiyu Gao et al.
This survey links Tent, CoT, PRMs, and MCTS to explain how test-time compute moves models from intuition toward deliberate reasoning.
Yixin Ji, Juntao Li, Yang Xiang et al.