Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry
Euclid-Omni unifies deduction and algebra through Euclidea, solving 595/601 Geometry3K problems and 16 IMO-AG-30 problems.
Zhaoyu Li, Hangrui Bi, Youyuan Zhang et al.
Euclid-Omni unifies deduction and algebra through Euclidea, solving 595/601 Geometry3K problems and 16 IMO-AG-30 problems.
Zhaoyu Li, Hangrui Bi, Youyuan Zhang et al.
CHIEF is a creator-driven video loop that extends student-made films from 1 to 10 minutes.
Denis Savytski, Aiden Lei, Heding Liu et al.
Quantifies text-conditioned impact on grayscale-to-color models (U-Net, Stable Diffusion 1.5); PSNR +5.6%-5.8%, LPIPS -7.6%-11.3%.
Colten Reissmann, Hugo Garrido-Lestache Belinchon
Introduces CF-GRPO framework to enhance video reasoning performance with Consensus Frame Reward, significantly improving multiple benchmarks.
Chengwen Liu, Zhe Huang, Jisheng Dang et al.
UniAR introduces a unified autoregressive framework with a single discrete visual tokenizer, achieving state-of-the-art results in image generation and understanding.
Wujian Peng, Lingchen Meng, Yuxuan Cai et al.
VERITAS framework uses inference-time verification with visual models to improve robot policies by 10% success rate without additional training.
Mingtong Zhang, Dhruv Shah
EventDrive integrates event cameras with vision-language models, significantly improving perception, understanding, prediction, and planning in autonomous driving.
Dongyue Lu, Rong Li, Ao Liang et al.
Proposes LoopWM, a parameter-shared transformer with iterative latent refinement, achieving 100× parameter efficiency and stable long-horizon environment prediction.
Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang et al.
RubricsTree constructs a hierarchical Boolean rubric system guided by expert-curated clinical criteria, enabling scalable, expert-aligned evaluation with over 100 atomic metrics, surpassing industry baselines.
Weizhi Zhang, Zechen Li, Hamid Palangi et al.
Proposed DRFLOW benchmark with 7 metrics, evaluating personalized workflow prediction across 100 tasks and 1246 steps, using multi-source evidence integration.
Md Tawkat Islam Khondaker, Raymond Li, Muhammad Abdul-Mageed et al.
Constructed a multi-source log dataset with 870 sessions, 2.3 million events, labeled with ATT&CK techniques, fine-tuned three small language models (Qwen, Llama, Phi) using LoRA, achieving up to 97% accuracy in chunk classification.
Abir Ashab Niloy, Ahmed Ryan, Imamul Hossain Rafi et al.
Introduces Kolmogorov PDE-based diffusion policies with dimension-independent convergence, improving long-horizon control in robotics and manufacturing.
Lekan Molu
Proposes quality-aware self-distillation with correctness gating and confidence scaling, boosting GUI coordinate accuracy by 2.16%.
Jingyuan Huang, Zuming Huang, Yucheng Shi et al.
Qwen-RobotManip introduces a unified alignment framework enabling large-scale multi-robot pretraining, achieving zero-shot instruction following and cross-embodiment transfer, surpassing SOTA.
Haoqi Yuan, Zhixuan Liang, Anzhe Chen et al.
Proposes a lightweight neuromorphic trigger based on fully connected LIF SNN achieving 0.97 F1 on ASD and 42.6× FLOPs reduction in SED.
Benjamin Hatton, Oliver Rhodes, Luca Peres
ActWorld adds action-aware memory to a 100K-video world model, enabling both navigation and mid-rollout object interaction.
Zhexiao Xiong, Yizhi Song, Hao Kang et al.
This study introduces RecLoop, a closed-loop simulation framework, comparing generative and traditional recommenders; findings show generative models better preserve diversity but still face cocoon effects.
Jiyuan Yang, Gengxin Sun, Mengqi Zhang et al.
FllumaOne introduces a multimodal CAD dataset with executable Python programs validated via OpenCASCADE, enabling high-precision geometric and feature-based learning.
Jizong Zhan
OPD-Evolver uses slow-fast on-policy distillation to learn memory selection, use, writing, and maintenance; its 9B model challenges 397B-scale systems.
Guibin Zhang, Xun Xu, Yanwei Yue et al.
SEAGym creates a dynamic evaluation environment for self-evolving LLM agents, comparing ACE, TF-GRPO, and AHE across multiple metrics on Terminal-Bench 2.0 and HLE.
Congjie Zheng, Chuanyi Xue, Bin Liang et al.