MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
MathCoder-VL uses executable code for vision–math alignment, reaching 73.6% on MathVista GPS.
Ke Wang, Junting Pan, Linda Wei et al.
MathCoder-VL uses executable code for vision–math alignment, reaching 73.6% on MathVista GPS.
Ke Wang, Junting Pan, Linda Wei et al.
MIRAGE benchmark evaluates models' spatial perception in Counting, Relation, and Counting with Relation tasks.
Chonghan Liu, Haoran Wang, Felix Henry et al.
L2T framework optimizes LLM reasoning efficiency and effectiveness using information-theoretic reinforcement learning, reducing unnecessary token usage.
Jingyao Wang, Wenwen Qiang, Zeen Song et al.
Integrating the ATACOM safety layer with geometric inductive biases ensures formal safety guarantees for robot foundation models without extensive safe demonstrations.
Maximilian Tölle, Theo Gruner, Daniel Palenicek et al.
Proposes a comprehensive safety assessment framework for ADS deployment, integrating multi-layer risk analysis and validation, ensuring no unreasonable residual risk.
Francesca Favaro, Scott Schnelle, Laura Fraade-Blanar et al.
Proposes EnerVerse-AC, a multi-level action-conditioned generative model for robotic environment simulation, reducing reliance on physical robots.
Yuxin Jiang, Shengcong Chen, Siyuan Huang et al.
MetaSPO optimizes system prompts via meta-learning, enhancing performance across 14 datasets.
Yumin Choi, Jinheon Baek, Sung Ju Hwang
Qwen3 integrates thinking and non-thinking modes, with 235B parameters, enhancing multilingual and multi-task performance.
An Yang, Anfeng Li, Baosong Yang et al.
InForage, a reinforcement learning framework based on Information Foraging Theory, optimizes dynamic retrieval and reasoning, boosting complex task performance.
Hongjin Qian, Zheng Liu
CodePDE framework leverages LLMs for PDE solver code generation, integrating reasoning, debugging, self-refinement, and scaling to surpass traditional methods.
Shanda Li, Tanya Marwah, Junhong Shen et al.
TRAIL framework evaluates multi-agent systems using 148 human-annotated traces; best model achieves only 11%.
Darshan Deshpande, Varun Gangal, Hersh Mehta et al.
Utilizing Fourier Neural Operators and Kernel Operator Learning to predict cardiac activation and repolarization times, enhancing computational efficiency.
Edoardo Centofanti, Giovanni Ziarelli, Nicola Parolini et al.
ResULIC combines Semantic Residual Coding and Compression-aware Diffusion for ultra-low bitrate image compression, saving over 66% in FID and LPIPS BD-rate.
Anle Ke, Xu Zhang, Tong Chen et al.
This survey reviews LLMs like GPT-4 and LLaMA in CAD, covering methods, datasets, experiments, and future directions, highlighting their industrial impact.
Licheng Zhang, Bach Le, Naveed Akhtar et al.
This paper introduces DanceGRPO, leveraging stable Group Relative Policy Optimization (GRPO) to significantly improve reinforcement learning for visual content generation, with up to 181% performance gains on benchmarks.
Zeyue Xue, Jie Wu, Yu Gao et al.
Step1X-3D combines VAE-DiT and diffusion models to generate high-quality, controllable 3D assets with a curated 2M dataset.
Weiyu Li, Xuanyang Zhang, Zheng Sun et al.
This study evaluates large language models' chemical reasoning via ChemIQ, with top models achieving 57% accuracy in structure understanding tasks.
Nicholas T. Runcie, Charlotte M. Deane, Fergus Imrie
This study introduces ChemRAG-Bench and ChemRAG-Toolkit, boosting chemistry task performance by 17.4% through multi-source retrieval-augmented generation.
Xianrui Zhong, Bowen Jin, Siru Ouyang et al.
MiMo-7B employs multi-stage data mixing and multi-token prediction to enhance reasoning, trained on 25 trillion tokens, with post-training on 130K problems, outperforming larger models.
LLM-Core Xiaomi, :, Bingquan Xia et al.
DARLR optimizes recommender systems with dynamic rewards, showing a 16% improvement on KuaiRand.
Yi Zhang, Ruihong Qiu, Xuwei Xu et al.