Conformal Risk Training: End-to-End Optimization of Conformal Risk Control
Proposes end-to-end conformal risk training extending CRC to OCE risks, improving model performance and guarantees.
Christopher Yeh, Nicolas Christianson, Adam Wierman et al.
Proposes end-to-end conformal risk training extending CRC to OCE risks, improving model performance and guarantees.
Christopher Yeh, Nicolas Christianson, Adam Wierman et al.
FlexTraj introduces a point-based trajectory control framework enabling multi-granularity, alignment-agnostic image-to-video synthesis with improved efficiency.
Zhiyuan Zhang, Can Wang, Dongdong Chen et al.
AutoMLGen integrates knowledge base and MCGS to optimize end-to-end ML pipelines, achieving 36.4% medal rate within 12 hours, outperforming baselines.
Shangheng Du, Xiangchao Yan, Dengyang Jiang et al.
ARES uses difficulty-aware window entropy shaping for adaptive multimodal reasoning, achieving superior performance across benchmarks.
Shuang Chen, Yue Guo, Yimeng Ye et al.
Proposed a guided diffusion model with Transformer for hyperspectral data augmentation, boosting forest classification accuracy by 4.8%.
Mattia Ferrari, Lorenzo Bruzzone
Proposes a dynamic optimal transport-based framework for multivariate counterfactual identification, ensuring uniqueness and consistency.
Fabio De Sousa Ribeiro, Ainkaran Santhirasekaram, Ben Glocker
VirT-Lab simulates customizable LLM teams in 2D worlds, validated through rescue alignment, ablations, a 12-person user study, and scaling tests.
Mohammed Almutairi, Charles Chiang, Haoze Guo et al.
Leveraging Whisper decoder embeddings for audio lyrics matching, achieving performance comparable to state-of-the-art methods.
Eleonora Mancini, Joan Serrà, Paolo Torroni et al.
HALCON method reduces insertion hallucination in video-to-audio generation by over 50%.
Liyang Chen, Hongkai Chen, Yujun Cai et al.
GRADE framework uses group-relative policy optimization and Dirichlet exploration for personalized multi-task fusion, improving CTR by 0.595% and CVR by 1.193%.
Tingfeng Hong, Pingye Ren, Xinlong Xiao et al.
Proposes Translation Tangles framework, integrating multi-metric, multi-domain, bias detection for 24 language pairs, analyzing translation quality and biases.
Md. Faiyaz Abdullah Sayeedi, Md. Mahbub Alam, Subhey Sadi Rahman et al.
Repainter uses spatial-matting reinforcement learning to significantly enhance e-commerce image removal.
Zipeng Guo, Lichen Ma, Xiaolong Fu et al.
Proposed ISMIE framework models information seeking via components, variables, activities; validated in misinformation and AI content trust scenarios.
Shuoqi Sun, Danula Hettiachchi, Damiano Spina
Proposes PA-Tool, using peakedness to align tool schemas with pretrained models, boosting accuracy by 17% without retraining.
Jonggeun Lee, Woojung Song, Jongwook Han et al.
Customer-R1 employs RL with explicit user personas, boosting next-action prediction accuracy from 7.32% to 39.58%, outperforming baselines.
Ziyi Wang, Yuxuan Lu, Yimeng Zhang et al.
MV-Performer employs depth-guided video diffusion to synthesize 360° synchronized multi-view human videos from monocular input.
Yihao Zhi, Chenghong Li, Hongjie Liao et al.
Proposes ETD, training models to iterate over key layers during mid-training, boosting reasoning accuracy by up to 36% on benchmarks.
Yeskendir Koishekenov, Aldo Lipani, Nicola Cancedda
TrackVLA++ enhances visual tracking with spatial reasoning and memory modules, achieving 5.1% and 12% improvements.
Jiahang Liu, Yunpeng Qi, Jiazhao Zhang et al.
VRPAgent uses LLMs to generate heuristic operators, optimized via genetic algorithms, outperforming handcrafted methods on multiple VRP variants.
André Hottung, Federico Berto, Chuanbo Hua et al.
SUPO algorithm scales LLM multi-turn RL via summarization-based context management, enhancing success rate and reducing context length.
Miao Lu, Weiwei Sun, Weihua Du et al.