Do You Remember? Toward Memory-Centric Multimodal AI
Proposes DoYouRemember architecture combining VQ-VAE, LoRA-tuned LLM, and diffusion decoder for image memory reconstruction.
Xuguang Yu, Weigang Zheng, Minyue Yu
Proposes DoYouRemember architecture combining VQ-VAE, LoRA-tuned LLM, and diffusion decoder for image memory reconstruction.
Xuguang Yu, Weigang Zheng, Minyue Yu
Introduces a stochastic quantum channel polynomial processing framework (QCPP) that enables polynomial approximations of Hamiltonian functions with reduced circuit complexity.
Tianhan Liu, Fedor Simkovic, Martin Leib
WristMimic uses wrist-guided whole-body control with reinforcement learning, achieving comparable performance to full finger supervision on manipulation tasks.
Wongyun Yu, Youngwoon Kim, Minsu Cho
LingBot-VLA 2.0 enhances task and embodiment generalization via 60,000-hour data pretraining.
Wei Wu, Fangjing Wang, Fan Lu et al.
Physically-grounded GPU apple-tree simulation using Euler-Bernoulli beams, rupture, and detachment mechanics, enabling autonomous harvesting research.
Humphrey Munn
Proposes a concept-based interpretability framework using unsupervised dictionary learning to improve transparency and performance of end-to-end autonomous driving models.
Franz Motzkus, Sebastian Bernhard
OTQL integrates advantage-weighted conditional optimal transport with flow models, achieving few-step inference and efficient policy fine-tuning, boosting success rates from 36% to 86%.
Andreas Sochopoulos, Esmeralda S. Whitammer, Nikolaos Tsagkas et al.
Comparative analysis of three major LLMs across 304 neighborhoods in five US cities reveals systematic fabrication and omission biases, influenced by socioeconomic factors, affecting urban information distribution.
Lin Chen, Guangyuan Weng, Esteban Moro
Pluralis v0.1 introduces a culture-first multimodal evaluation framework with 6,448 prompts, exposing VLM blind spots in cultural alignment.
Alicia Parrish, Rajat Shinde, Sanket Badhe et al.
MobileWan employs recursive distillation and structured pruning to deploy a 5B-parameter video diffusion model on mobile devices, enabling 5-second 480x832 videos at 16 FPS in 20 seconds.
Mohsen Ghafoorian, Denis Korzhenkov, Adil Karjauv et al.
Identifies reward hacking in reference-free LLM judges via hidden-anchor audit; “commit-answer-first” strategy effectively reduces false positives.
Chenyu Zhou
Using horizon scanning, identified security and privacy challenges in agentic AI, proposed future research directions.
Adam Jenkins, Agnieszka Kitkowska, Caterina Maidhof et al.
TurnOPD employs adaptive turn-depth control and progressive loss normalization to enhance long-horizon agent distillation efficiency, improving validation accuracy.
Yuhang Zhou, Kai Zheng, Haoling Li et al.
Proposes Heading-Specific Activation Steering to causally control tool invocation in five open-source models, validated through geometric and causal analysis.
Yuqi Chen, Vincent Siu, Yang Liu et al.
Introduced LanEvil++ benchmark, evaluating lane perception robustness; models drop 5.27% accuracy under environmental illusions.
Tianyuan Zhang, Xianglong Liu, Aishan Liu et al.
Image2Sim uses neural simulation to build scalable, high-fidelity interactive environments for embodied navigation training.
Zihan Wang, Seungjun Lee, Yinghao Xu et al.
Proposes TuneNNGen, leveraging source models and LLMs to boost CIFAR-10 accuracy from 23.98% to 50.49%.
Kabir Dev Paul Baghel, Radu Timofte, Dmitry Ignatov
This study systematically compares BPE and Unigram-LM tokenizers on fixed 165-token chemical bases, revealing near-disjoint vocabularies and emphasizing tokenizer choice as a key model design decision.
Hunter Heidenreich
Proposes in-loop memory for language agents, reducing latency to 100μs and improving task efficiency and accuracy.
Yusuf Khan, Carlo Lipizzi
This study introduces M3Bench, a benchmark for evaluating model editing in medical VLMs, measuring reliability, locality, and generalization across 16,276 clinical questions.
Guli Zhu, Chenwei Wu, Liyue Shen