Training Language Model to Critique for Better Refinement
RCO framework trains critic models via refinement signals, improving critique quality and response refinement by 15% on average.
Tianshu Yu, Chao Xiang, Mingchuan Yang et al.
RCO framework trains critic models via refinement signals, improving critique quality and response refinement by 15% on average.
Tianshu Yu, Chao Xiang, Mingchuan Yang et al.
A Vicsek-like model with random non-reciprocal interactions reveals flocking, chiral, and oscillatory states, with implications for multi-species collective behavior.
Jiwon Choi, Jae Dong Noh, Heiko Rieger
StableCodec employs single-step diffusion for ultra-low bitrate image compression, achieving FID of 7.2 and real-time inference speed.
Tianyu Zhang, Xin Luo, Li Li et al.
Proposed Image2Net framework achieves 80.77% success rate in converting complex analog circuit diagrams to netlists, with a novel dataset and NED evaluation, advancing circuit recognition.
Haohang Xu, Chengjie Liu, Qihang Wang et al.
ARAG employs multi-agent reasoning—user understanding, semantic inference, context summarization, and ranking—achieving up to 42.1% NDCG@5 improvement in personalized recommendations.
Reza Yousefi Maragheh, Pratheek Vadla, Priyank Gupta et al.
DIVE employs iterative reasoning with semantic decomposition, achieving 81.44% accuracy on CVRR-ES, winning the challenge.
Umihiro Kamoto, Tatsuya Ishibashi, Noriyuki Kugo
Proposed IHRUT network leverages degradation modeling to reconstruct IHI images, achieving 36.04dB PSNR.
Yuansheng Li, Yunhao Zou, Linwei Chen et al.
Introduces Mind2Web 2 benchmark with Agent-as-a-Judge, evaluating 130 long-horizon tasks achieving 50-70% of human performance.
Boyu Gou, Zanming Huang, Yuting Ning et al.
TR²C combines MCR² and temporal continuity, reaching 97.96% ACC and 98.96% NMI across five HMS benchmarks.
Xianghan Meng, Zhengyu Tong, Zhiyuan Huang et al.
DFVEdit uses Conditional Delta Flow Vectors for zero-shot video editing, achieving 20× speed-up and 85% memory reduction, maintaining state-of-the-art quality.
Lingling Cai, Kang Zhao, Hangjie Yuan et al.
This study introduces a post-training framework for VLA models based on human motor learning constraints, improving perception, embodiment, and task understanding, with 15% success rate gains.
Tian-Yu Xiang, Ao-Qun Jin, Xiao-Hu Zhou et al.
FineWeb2 pipeline enables automatic filtering and deduplication supporting over 1000 languages, improving multilingual model performance.
Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec et al.
Cognitive models reveal value trade-offs in language models using RSA for polite speech analysis.
Sonia K. Murthy, Rosie Zhao, Jennifer Hu et al.
Analyzes two-tower model identifiability and bias amplification; proposes sample weighting to mitigate bias effects.
Philipp Hager, Onno Zoeter, Maarten de Rijke
Mobile-R1 enhances VLM-based mobile agents' interactive capabilities through systematic training, significantly improving exploration and self-correction.
Jihao Gu, Qihang Ai, Yingyao Wang et al.
Proposes multi-perspective soft label models to improve inclusivity and performance in subjective NLP tasks.
Benedetta Muscato, Lucia Passaro, Gizem Gezici et al.
AdaDeDup combines density clustering and model feedback for adaptive data pruning, reducing 20% data with minimal performance loss.
Feiyang Kang, Nadine Chang, Maying Shen et al.
This study analyzes the behavioral diversity limitations of LLM social simulations, emphasizing the importance of boundary setting.
Zengqing Wu, Run Peng, Takayuki Ito et al.
Utilizing virtual memory for real-time rendering of large-scale 3D Gaussian Splatting scenes, significantly reducing memory usage.
Jonathan Haberl, Philipp Fleck, Clemens Arth
Proposes Surfel-Indexed View Memory (VMem) for long-term scene consistency, reducing computational cost while maintaining scene coherence.
Runjia Li, Philip Torr, Andrea Vedaldi et al.