DHAGrasp: Synthesizing Affordance-Aware Dual-Hand Grasps with Text Instructions
DHAGrasp synthesizes dual-hand grasps guided by text, improving grasp quality on unseen objects.
Quanzhou Li, Zhonghua Wu, Jingbo Wang et al.
DHAGrasp synthesizes dual-hand grasps guided by text, improving grasp quality on unseen objects.
Quanzhou Li, Zhonghua Wu, Jingbo Wang et al.
Proposes a multi-level MOWI framework analyzing hallucinations, revealing data biases and model mechanisms affecting output fidelity.
Zhengyi Ho, Siyuan Liang, Dacheng Tao
Efficient trial-and-error GPT-2 combines rules, DFS backtracking, and verification, reaching 99% on Sudoku and 1-in-3 SAT.
Panagiotis Giannoulis, Yorgos Pantis, Christos Tzamos
Bilinear representation mitigates reversal curse, enabling consistent model editing.
Dong-Kyum Kim, Minsung Kim, Jea Kwon et al.
Using EgoScaler to automatically extract 6DoF object trajectories from unlabeled egocentric videos significantly improves VLA pre-training, achieving over 20% success rate gains.
Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura et al.
LDAR introduces a distraction-aware retrieval method using similarity distribution bands, improving knowledge grounding with 5%+ accuracy gains and reduced token usage.
Seongwoong Shim, Myunsoo Kim, Jae Hyeon Cho et al.
D-Artemis framework achieves 75.8% success rate on AndroidWorld, enhancing MLLMs for GUI tasks.
Hongze Mi, Yibo Feng, Wenjie Lu et al.
Proposed ASIMOV-2.0 benchmark integrates real injury reports and generative models to evaluate AI's physical risk perception and intervention.
Abhishek Jindal, Dmitry Kalashnikov, R. Alex Hofer et al.
Semantic F1 scores incorporate label semantic similarity, enabling fair evaluation of fuzzy multi-label predictions, outperforming traditional F1.
Georgios Chochlakis, Jackson Trager, Vedant Jhaveri et al.
Introduces Tree-GRPO, a tree search-based policy optimization, boosting sample efficiency and process supervision in multi-turn LLM agent tasks.
Yuxiang Ji, Ziyu Ma, Yong Wang et al.
Study reveals syntactic-domain spurious correlations in language models, affecting performance.
Chantal Shaib, Vinith M. Suriyakumar, Levent Sagun et al.
Introduces Visual Test-Time Scaling (VTTS) with iterative perception, boosting multimodal reasoning by over 5% on benchmarks.
Ziang Yan, Xinhao Li, Yinan He et al.
MTRDrive combines memory retrieval and tool interaction, boosting autonomous driving robustness in corner cases by 20% on key metrics.
Ziang Luo, Kangan Qian, Jiahua Wang et al.
Meta-Memory employs joint semantic-spatial reasoning with LLMs to significantly improve robot spatial question-answering, achieving 67.8% success on SpaceLocQA, outperforming SOTA.
Yufan Mao, Hanjing Ye, Wenlong Dong et al.
RobotDancing uses residual-action reinforcement learning for robust long-horizon humanoid motion tracking.
Zhenguo Sun, Yibo Peng, Yuan Meng et al.
RuN framework uses residual policy for natural humanoid locomotion, achieving 0-2.5m/s velocity range with improved training efficiency.
Qingpeng Li, Chengrui Zhu, Yanming Wu et al.
SIL-C framework employs bilateral lazy learning to maintain skill-policy compatibility, improving incremental learning with 42.5% FWT.
Daehee Lee, Dongsu Lee, TaeYoon Kwack et al.
Proposes an on-device AI framework with adaptive memory and tool management, reducing context size by over 6x while maintaining performance.
Sanidhya Vijayvargiya, Rahul Lokesh
Proposes CR-PPO, replacing entropy regularization with complexity measure for robustness and exploration balance.
Luca Serfilippi, Giorgio Franceschelli, Antonio Corradi et al.
Seedream 4.0 uses an efficient diffusion transformer and VAE to enhance high-res multimodal image generation.
Team Seedream, :, Yunpeng Chen et al.