Attribute-Based Robotic Grasping with Data-Efficient Adaptation
Using attribute learning and data-efficient adaptation, robotic grasping achieves over 81% success rate.
Yang Yang, Houjian Yu, Xibai Lou et al.
Using attribute learning and data-efficient adaptation, robotic grasping achieves over 81% success rate.
Yang Yang, Houjian Yu, Xibai Lou et al.
AVTrustBench assesses AVLLM reliability; CAVPref improves performance by 30.19%.
Sanjoy Chowdhury, Sayan Nag, Subhrajyoti Dasgupta et al.
Proposes BeliN corpus and MultiGen model, integrating category, aspect, sentiment features, achieving BLEU 18.61 in Bengali religious news headline generation.
Md Osama, Ashim Dey, Kawsar Ahmed et al.
FlashInfer employs block-sparse formats and JIT compilation to optimize large language model attention inference.
Zihao Ye, Lequn Chen, Ruihang Lai et al.
Introduces a multimodal textbook dataset to enhance VLMs' performance in knowledge reasoning tasks.
Wenqi Zhang, Hang Zhang, Xin Li et al.
VinT-6D dataset integrates vision, touch, and proprioception to enhance robotic manipulation precision.
Zhaoliang Wan, Yonggen Ling, Senlin Yi et al.
Introduces Sui Generis score to quantify plot diversity in LLM stories; experiments with GPT-4 and LLaMA-3 show high plot element repetition, indicating lack of creativity.
Weijia Xu, Nebojsa Jojic, Sudha Rao et al.
This paper introduces LS-GAN, a latent-space GAN for human motion synthesis, achieving FID 0.482 with 91% FLOPs reduction compared to diffusion models.
Avinash Amballa, Gayathri Akkinapalli, Vinitra Muralikrishnan
Using Bentkus and Pinelis techniques, the paper improves martingale concentration inequalities, recovering missing factors and approaching near-optimal bounds.
Arun Kumar Kuchibhotla
Survey on visual grounding, analyzing new concepts like grounded pre-training and multimodal LLMs, providing comprehensive research directions.
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan et al.
Proposes InfAlign, optimizing inference win rate via reward transformation, achieving 3-8% improvements.
Ananth Balashankar, Ziteng Sun, Jonathan Berant et al.
OS-Genesis employs reverse task synthesis via interaction-driven exploration, significantly improving GUI trajectory quality and diversity, boosting model performance by nearly 80% on benchmarks.
Qiushi Sun, Kanzhi Cheng, Zichen Ding et al.
BeSplat recovers high-quality radiance fields from a single blurry image and event stream.
Gopi Raju Matta, Reddypalli Trisha, Kaushik Mitra
ETTA model optimizes text-to-audio conversion through large-scale experiments, enhancing generation quality and speed.
Sang-gil Lee, Zhifeng Kong, Arushi Goel et al.
TravelAgent combines generative agents with 3D environments, achieving 76% task completion over 1898 steps, enhancing urban behavior simulation.
Ariel Noyman, Kai Hu, Kent Larson
DrivingGPT employs a multimodal autoregressive Transformer to unify world modeling and path planning, outperforming baselines with superior video quality and planning scores.
Yuntao Chen, Yuqi Wang, Zhaoxiang Zhang
Introduces LongDocURL, a comprehensive multimodal benchmark for long document understanding, with 20 sub-tasks, 2,325 high-quality QA pairs, revealing significant performance gaps.
Chao Deng, Jiale Yuan, Pi Bu et al.
Proposes Multimodal Chain-of-Thought Co-Navigation (MCoCoNav) for multi-robot semantic navigation, achieving over 87% success rate and 46% SPL on HM3D.
Zhixuan Shen, Haonan Luo, Kexun Chen et al.
VLABench benchmark evaluates language-conditioned robotic manipulation with 100 long-horizon tasks, testing generalization and reasoning capabilities of VLAs.
Shiduo Zhang, Zhe Xu, Peiju Liu et al.
HSfM combines deep learning and SfM to jointly reconstruct multiple human meshes, scene point clouds, and camera parameters with high accuracy.
Lea Müller, Hongsuk Choi, Anthony Zhang et al.