Attractor-Keyed Memory
Attractor-Keyed Memory merges selection and memory access, reducing latency and energy in sparse routing architectures.
Natalia G. Berloff
Attractor-Keyed Memory merges selection and memory access, reducing latency and energy in sparse routing architectures.
Natalia G. Berloff
Video models exhibit reasoning via Chain-of-Steps mechanism during diffusion denoising steps.
Ruisi Wang, Zhongang Cai, Fanyi Pu et al.
MessyKitchens achieves high-precision monocular 3D scene reconstruction using the MOD algorithm, significantly enhancing the physical plausibility of inter-object contacts.
Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati et al.
Efficient reasoning in small LLMs using LoRA adapters and RL, significantly reducing response length.
Yelysei Bondarenko, Thomas Hehn, Rob Hesselink et al.
SegviGen repurposes 3D generative models for part segmentation, achieving a 40% improvement in interactive segmentation using only 0.32% labeled data.
Lin Li, Haoran Feng, Zehuan Huang et al.
ManiTwin generates 100K high-quality 3D digital assets from a single image for large-scale robotic manipulation data generation.
Kaixuan Wang, Tianxing Chen, Jiawei Liu et al.
BrickSim is a physics-based simulator for real-time simulation of brick assemblies, achieving 100% accuracy.
Haowei Wen, Ruixuan Liu, Weiyi Piao et al.
M^3 integrates multi-view foundation models with monocular Gaussian splatting SLAM, reducing ATE RMSE by 64.3%.
Kerui Ren, Guanghao Li, Changjian Jiang et al.
LEAFE framework internalizes recovery agency from reflective experience, enhancing Pass@k performance in long-horizon tasks.
Rui Ge, Yichao Fu, Yuyang Qian et al.
Proposes a compilation framework with persistent dimensional annotations for joint numeric representation and deterministic memory management.
Houston Haynes
PKINet-v2 combines anisotropic strip and square kernels to build multi-scale, geometry-adaptive features, achieving 80.46% mAP and 54.6 FPS on DOTA-v1.0.
Xinhao Cai, Liulei Li, Gensheng Pei et al.
Achieved superior drone interception using PPO-based competitive reinforcement learning with high catch rates.
Timothée Gavin, Simon Lacroix, Murat Bronz
MSRAMIE employs structured multimodal reasoning with Tree-of-States and Graph-of-References to improve multi-instruction image editing by over 15%.
Zhaoyuan Qiu, Ken Chen, Xiangwei Wang et al.
EVPV gates PRM rewards by visual-premise reliability, reaching 67.46% Macro-F1 and improving Best-of-8 reranking.
Junxin Wang, Dai Guan, Weijie Qiu et al.
Proposes data-dependent and post-hoc equivalence margins using e-values for model validation with uniform guarantees.
Stan Koobs, Nick W. Koning
Proposes SignNav and START model, using spatial-temporal Transformer for semantic indoor navigation, achieving 80% success rate.
Jian Sun, Yuming Huang, He Li et al.
DualPrim achieves compact 3D reconstruction using positive and negative superquadrics, enhancing structural expressiveness.
Xiaoxu Meng, Zhongmin Chen, Bo Yang et al.
Helium framework boosts LLM workflow efficiency via proactive caching and scheduling, achieving up to 1.56x speedup.
Noppanat Wadlom, Junyi Shen, Yao Lu
Challenges the intra-modal misalignment hypothesis in CLIP, finding task ambiguity, not misalignment, is key.
Jonas Herzog, Yue Wang
RAISE method ensures safety for VLA driving systems, enhancing system trust.
Gerhard Yu, Fuyuki Ishikawa, Oluwafemi Odu et al.