SpatialVAM:Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy
SpatialVAM achieves data-efficient robot policy learning via 3D video diffusion, improving Meta-World success rates by 22%.
Peiyan Li, Yixiang Chen, Yuan Xu et al.
SpatialVAM achieves data-efficient robot policy learning via 3D video diffusion, improving Meta-World success rates by 22%.
Peiyan Li, Yixiang Chen, Yuan Xu et al.
Advantage Reward Modeling (ARM) uses tri-state labels and a MIMO Transformer to estimate relative advantage, achieving 99.4% success in long-horizon towel-folding tasks.
Yiming Mao, Zixi Yu, Weixin Mao et al.
InfoSeeker uses Host–Manager–Worker parallelism, reaching 8.38% WideSearch success and 3–5× faster execution.
Ka Yiu Lee, Yuxuan Huang, Zhiyuan He et al.
RTT framework maps response-level scores to token rewards via a Token Relevance Discriminator, improving instruction-following accuracy by 2.5% over baselines.
Tianze Xu, Yanzhao Zheng, Pengrui Lu et al.
PolyBench combines 38,666 binary prediction markets, order book states, and news streams to evaluate 7 LLMs, with only MiMo-V2-Flash and Gemini-3-Flash achieving positive returns.
Pu Cheng, Juncheng Liu, Yunshen Long
Trivial vocabulary bans (e.g., 'very', 'just') outperform deep linguistic constraints like E-Prime in improving LLM reasoning, due to output regularization effects.
Rodney Jehu-Appiah
LLMimic uses role-playing LLM training to enhance AI literacy, reduce persuasion success, and improve social responsibility.
Qihui Fan, Min Ge, Chenyan Jia et al.
VERTIGO optimizes visual preference, reducing off-screen rate to 0% and enhancing shot quality.
Mengtian Li, Yuwei Lu, Feifei Li et al.
Multi-agent video recommenders with LLMs enhance precision and explainability.
Srivaths Ranganathan, Abhishek Dharmaratnakar, Anushree Sinha et al.
Using social media simulation, analysis reveals LLMs over-idealize disability, reinforcing biases and stereotypes.
Marco Bombieri, Simone Paolo Ponzetto, Marco Rospocher
RebusBench evaluates LVLMs' cognitive visual reasoning, performance below 10% exact match.
Seyed Amir Kasaei, Arash Marioriyad, Mahbod Khaleti et al.
SteerFlow introduces fixed-point and trajectory interpolation techniques to improve faithful inversion-based image editing, outperforming existing methods in source preservation.
Thinh Dao, Zhen Wang, Kien T. Pham et al.
Unified modular framework for agent memory; fusion model outperforms SOTA in long-term tasks with 8-15% improvements.
Yanchen Wu, Tenghui Lin, Yingli Zhou et al.
ReFlow employs self-correction flow matching for monocular 4D scene reconstruction, surpassing existing methods with no external motion guidance, achieving PSNR of 27.65dB.
Yanzhe Liang, Ruijie Zhu, Hanzhi Chang et al.
DBCooker leverages LLMs to automate database native function synthesis, achieving 34.55% higher accuracy than state-of-the-art methods.
Wei Zhou, Xuanhe Zhou, Qikang He et al.
Proposes the Magic-Madness-Heaven-Sin framework, categorizing LLM output diversity by task goals, revealing trade-offs across contexts.
Harnoor Dhingra
Introduced FoodGuardBench to evaluate LLMs' food safety, revealing three major vulnerabilities.
Weidi Luo, Xiaofei Wen, Tenghao Huang et al.
Enhancing latent generalization using test-time compute with reinforcement learning for long chain-of-thoughts.
Arslan Chaudhry, Sridhar Thiagarajan, Andrew Lampinen
Study reveals GRU-based RSSM's safety risks under adversarial attacks, with a 59.5% reward reduction.
Manoj Parmar
HippoCamp benchmarks multimodal file management agents, revealing limitations in user environments with top accuracy only 48.3%.
Zhe Yang, Shulin Tian, Kairui Hu et al.