Tac-Man: Tactile-Informed Prior-Free Manipulation of Articulated Objects
Tac-Man leverages tactile feedback for prior-free manipulation of articulated objects, achieving over 98% success in complex scenarios.
Zihang Zhao, Yuyang Li, Wanlin Li et al.
Tac-Man leverages tactile feedback for prior-free manipulation of articulated objects, achieving over 98% success in complex scenarios.
Zihang Zhao, Yuyang Li, Wanlin Li et al.
SceneCraft uses LLMs to convert text into Blender scripts, enabling the creation of complex scenes with up to 100 assets.
Ziniu Hu, Ahmet Iscen, Aashi Jain et al.
LLaMoCo fine-tunes large language models with contrastive learning for expert-level optimization code generation, outperforming GPT-4 Turbo in benchmarks.
Zeyuan Ma, Hongshu Guo, Jiacheng Chen et al.
This study systematically compares multiple scoring methods for LLMs in multiple-choice tasks, revealing high sensitivity and variability across methods and models.
Polina Tsvilodub, Hening Wang, Sharon Grosch et al.
Proposes Selective Recurrent Unit (SRU) with adaptive frequency fusion, achieving top accuracy in stereo matching, ranked first on KITTI and other benchmarks.
Xianqi Wang, Gangwei Xu, Hao Jia et al.
TempCompass benchmark evaluates 8 SOTA Video LLMs across five temporal dimensions, revealing their poor temporal perception abilities, with average accuracy around 33.9%.
Yuanxin Liu, Shicheng Li, Yi Liu et al.
Proposes artwork explanation generation task; LVLMs face challenges in integrating language and visual information.
Kazuki Hayashi, Yusuke Sakai, Hidetaka Kamigaito et al.
RAG combines retrieval with generation, boosting accuracy and robustness in AI-generated content.
Penghao Zhao, Hailin Zhang, Qinhan Yu et al.
Curiosity-driven red teaming (CRT) leverages exploration rewards to enhance test coverage, successfully eliciting toxic responses from LLaMA2, with a 19.6% toxicity rate compared to 10.2% by baseline methods.
Zhang-Wei Hong, Idan Shenfeld, Tsun-Hsuan Wang et al.
Introduced MATH() benchmark revealing reasoning gaps of 58.35% to 80.31%.
Saurabh Srivastava, Annarose M B, Anto P et al.
ArCHer employs hierarchical RL with high-level value functions guiding token-level policies, achieving 100x sample efficiency in multi-turn tasks.
Yifei Zhou, Andrea Zanette, Jiayi Pan et al.
Proposes Lower-Left Partial AUC (LLPAUC) as an efficient, top-K aligned metric for recommendation, validated through theory and experiments, improving ranking performance.
Wentao Shi, Chenxu Wang, Fuli Feng et al.
ToolNet connects large language models with massive tools via a tool graph, significantly improving performance on multi-hop tool learning datasets.
Xukun Liu, Zhiyuan Peng, Xiaoyuan Yi et al.
Proposed a coarse-to-fine affordance learning method to significantly reduce the impact of point cloud noise.
Suhan Ling, Yian Wang, Shiguang Wu et al.
CLLMs improve Jacobi decoding to achieve 2.4x to 3.4x speedup while maintaining quality.
Siqi Kou, Lanxiang Hu, Zhezhi He et al.
The study reveals cultural bias in XAI research, highlighting most studies ignore cultural differences.
Uwe Peters, Mary Carman
Proposed a multimodal handover failure detection dataset using video, force-torque, and gripper data; baseline methods include 3D CNN and action segmentation, achieving 67.9% accuracy.
Santosh Thoduka, Nico Hochgeschwender, Juergen Gall et al.
Proposes TedRec, a Fourier transform-based sequence-level semantic fusion method that improves recommendation accuracy by 14%-38% across five datasets.
Lanling Xu, Zhen Tian, Bingqian Li et al.
Developed JAMA Clinical Challenge and Medbullets datasets to evaluate 7 LLMs on complex medical QA.
Hanjie Chen, Zhouxiang Fang, Yash Singla et al.
ZOD-MC offers gradient-free sampling for non-log-concave distributions with polynomial convergence guarantees in low dimensions.
Ye He, Kevin Rojas, Molei Tao