MoReact: Generating Reactive Motion from Textual Descriptions
MoReact uses diffusion models to generate reactive motions based on textual descriptions, enhancing interaction realism.
Xiyan Xu, Sirui Xu, Yu-Xiong Wang et al.
MoReact uses diffusion models to generate reactive motions based on textual descriptions, enhancing interaction realism.
Xiyan Xu, Sirui Xu, Yu-Xiong Wang et al.
Proposed a multi-agent reinforcement learning framework with hybrid proximal policy optimization and action masking for energy-efficient UAV swarm communication and control, achieving 25% energy savings and 0.99 fairness index.
Tianjiao Sun, Ningyan Guo, Haozhe Gu et al.
DocPruner employs adaptive patch pruning based on global token attention, reducing storage by 50-60% with minimal performance loss.
Yibo Yan, Guangwei Xu, Xin Zou et al.
GUI-Shepherd improves long-sequence GUI task success rate by 7.7 points using a Process Reward Model.
Cong Chen, Kaixiang Ji, Hao Zhong et al.
Introduces ReWatch dataset and multi-stage synthesis with Multi-Agent ReAct, boosting large vision-language models' video reasoning accuracy by over 3%.
Congzhi Zhang, Zhibin Wang, Yinchao Ma et al.
Proposes CLAP framework combining task decomposition and 3D keypoint prediction, achieving 12% higher success rate on GemBench with only 1/5 training data.
Jianshu Hu, Lidi Wang, Shujia Li et al.
A dual-distribution study finds weak-to-strong judge transfer fails, while DPO refreshes deliver up to 7.6 percentage points.
Janvijay Singh, Austin Xu, Yilun Zhou et al.
SPEC-RL integrates speculative decoding with RL rollouts, reusing previous trajectory segments to accelerate training 2-3× while maintaining policy quality.
Bingshuai Liu, Ante Wang, Zijun Min et al.
Introduces a geometric credal set framework to quantify and decompose uncertainty in language models, analyzing 500 prompts with only 0.434 calibration at best.
Esteban Garces Arias, Julian Rodemann, Christian Heumann
ARMimic uses XR headsets for passive demonstrations, combining hand tracking and virtual robots to improve data collection efficiency and generalization.
Rohan Walia, Yusheng Wang, Ralf Römer et al.
WoW: a 14B-parameter generative world model trained on 2 million robot trajectories, demonstrating physical intuition and causal reasoning.
Xiaowei Chi, Peidong Jia, Chun-Kai Fan et al.
Using n-gram novelty as a sole metric for textual creativity is inadequate; models and experts show high correlation but also significant divergence, emphasizing multi-dimensional evaluation.
Arkadiy Saakyan, Najoung Kim, Smaranda Muresan et al.
LongLive uses a causal autoregressive model with KV recaching, achieving 20.7FPS for real-time long video generation up to 240 seconds.
Shuai Yang, Wei Huang, Ruihang Chu et al.
JanusVLN employs dual implicit neural memory to separate spatial geometry and semantic info, boosting vision-language navigation.
Shuang Zeng, Dekang Qi, Xinyuan Chang et al.
Proposed AMBS method improves HHH performance to 56.5% on LLaMA-2-7B while maintaining efficiency.
Gautam Siddharth Kashyap, Mark Dras, Usman Naseem
Proposes EELMA, an information-theoretic method to estimate language model agents' empowerment, strongly correlating with task performance.
Jinyeop Song, Jeff Gore, Max Kleiman-Weiner
Proposes AURA, an auction-based multi-robot task allocation algorithm that models task requirement uncertainty, improving expected mission value by up to 15%.
Ben Rossano, Jaein Lim, Jonathan P. How
Aurora is a multimodal time series foundation model supporting zero-shot cross-domain forecasting, using flow matching and prototype-guided mechanisms.
Xingjian Wu, Jianxin Jin, Wanghui Qiu et al.
RepT framework uses representation gradients to accurately trace undesirable LLM behaviors at sample and token levels, outperforming gradient-based baselines.
Zhe Li, Wei Zhao, Yige Li et al.
MimicDreamer aligns human videos with robot actions via diffusion-based video synthesis, stabilizes egocentric views, and maps human trajectories to robot joints, boosting VLA performance.
Haoyun Li, Ivan Zhang, Runqi Ouyang et al.