Speculative Decoding: Exploiting Speculative Execution for Accelerating Seq2seq Generation
SpecDec leverages speculative execution to accelerate seq2seq generation by 5x with quality comparable to beam search.
Heming Xia, Tao Ge, Peiyi Wang et al.
SpecDec leverages speculative execution to accelerate seq2seq generation by 5x with quality comparable to beam search.
Heming Xia, Tao Ge, Peiyi Wang et al.
TubeDETR uses Transformers for spatio-temporal video grounding, improving performance on VidSTG and HC-STVG benchmarks.
Antoine Yang, Antoine Miech, Josef Sivic et al.
Proposes a step-function-based univariate direct sampling method that reduces manual tuning and rejection, enabling exact sampling from complex distributions.
Andrew M. Raim
OakInk combines object affordance and human interaction data, with 50,000 instances, advancing pose estimation and interaction generation.
Lixin Yang, Kailin Li, Xinyu Zhan et al.
Proposed a primal-dual kernelized bandit algorithm with sublinear regret and soft constraint violation guarantees, compatible with UCB, TS, and random exploration.
Xingyu Zhou, Bo Ji
This study introduces compute-optimal training for large language models, showing model size and data should scale together; trained 70B Chinchilla surpasses larger models.
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch et al.
Proposes adversarial motion priors (AMP) using limited motion capture data to replace complex reward functions for natural, energy-efficient robot locomotion.
Alejandro Escontrela, Xue Bin Peng, Wenhao Yu et al.
OpenDet method enhances open-set object detection by expanding low-density latent regions, reducing errors by 25%-35%.
Jiaming Han, Yuqiang Ren, Jian Ding et al.
MedMCQA is a large-scale, multi-subject medical QA dataset with over 194k questions, designed to advance deep reasoning in medical AI.
Ankit Pal, Logesh Kumar Umapathi, Malaikannan Sankarasubbu
BARCOR: A BART-based unified framework for conversational recommendation systems achieving state-of-the-art performance in the movie domain.
Ting-Chun Wang, Shang-Yu Su, Yun-Nung Chen
Introduced a spatio-temporal grounded audio-visual network, excelling on the MUSIC-AVQA dataset.
Guangyao Li, Yake Wei, Yapeng Tian et al.
SolidGen directly generates B-reps autoregressively, enabling conditional CAD synthesis without sequence supervision.
Pradeep Kumar Jayaraman, Joseph G. Lambourne, Nishkrit Desai et al.
CODEGEN uses multi-turn prompts to improve synthesis, reaching 47.34% on MTPB with its 16.1B model.
Erik Nijkamp, Bo Pang, Hiroaki Hayashi et al.
UMT framework significantly improves video moment retrieval and highlight detection, excelling on the QVHighlights dataset.
Ye Liu, Siyuan Li, Yang Wu et al.
Proposes rhythmic precision-modulation active inference to enhance robotic perception-action coupling, improving state estimation and path planning.
Ajith Anil Meera, Filip Novicky, Thomas Parr et al.
VideoMAE employs extremely high masking ratios (90%-95%) for self-supervised pretraining, significantly boosting video recognition accuracy.
Zhan Tong, Yibing Song, Jue Wang et al.
This paper introduces Visual Prompt Tuning (VPT), which inserts less than 1% trainable prompts into the input space of frozen vision transformers, outperforming full fine-tuning on 20 out of 24 tasks with significantly fewer parameters.
Menglin Jia, Luming Tang, Bor-Chun Chen et al.
Survey of Vision-and-Language Navigation tasks, methods, and future directions, emphasizing datasets and evaluation metrics.
Jing Gu, Eliana Stefani, Qi Wu et al.
CREStereo employs a hierarchical recurrent network with adaptive correlation, achieving top performance on Middlebury and ETH3D benchmarks with significant detail preservation.
Jiankun Li, Peisen Wang, Pengfei Xiong et al.
Proposes oscillation dampening and iterative weight freezing to improve low-bit quantization accuracy of models like MobileNetV2, achieving state-of-the-art results.
Markus Nagel, Marios Fournarakis, Yelysei Bondarenko et al.