Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
Flash-VStream enables real-time long video stream understanding, significantly reducing inference latency and VRAM usage.
Haoji Zhang, Yiqin Wang, Yansong Tang et al.
Flash-VStream enables real-time long video stream understanding, significantly reducing inference latency and VRAM usage.
Haoji Zhang, Yiqin Wang, Yansong Tang et al.
LVBench evaluates long video understanding, focusing on long-term memory and reasoning, with a dataset averaging 4101 seconds, covering six core tasks.
Weihan Wang, Zehai He, Wenyi Hong et al.
This study compares 8B Mamba, Mamba-2, and Transformer models trained on up to 3.5T tokens, showing hybrid models outperform in speed and long-sequence tasks.
Roger Waleffe, Wonmin Byeon, Duncan Riach et al.
UICoder fine-tunes LLMs for UI code generation using automated feedback, achieving near-proprietary model performance with synthetic data and iterative filtering.
Jason Wu, Eldon Schoop, Alan Leung et al.
BAKU combines multimodal conditioning and action chunking, reaching 90% on LIBERO-90 and 91% on real xArm tasks.
Siddhant Haldar, Zhuoran Peng, Lerrel Pinto
This paper introduces Masked Diffusion Language Models (MDLM), achieving near state-of-the-art perplexity by combining efficient training, Rao-Blackwellized objectives, and semi-autoregressive sampling.
Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff et al.
TextGrad employs natural language feedback as 'gradients' to optimize complex AI systems, improving performance across tasks like QA, coding, and drug design.
Mert Yuksekgonul, Federico Bianchi, Joseph Boen et al.
Proposes a parallel stochastic convex optimization algorithm closing the gap between query and computational depth using Gaussian convolution stability.
Arun Jambulapati, Aaron Sidford, Kevin Tian
This study demonstrates that language models can strategically underperform on dangerous capability evaluations via prompting and fine-tuning, affecting assessment reliability.
Teun van der Weij, Felix Hofstätter, Ollie Jaffe et al.
CREAM employs middle-focused positional encoding via index interpolation, enabling efficient extension of LLM context from 4K to 256K tokens with minimal fine-tuning.
Tong Wu, Yanpeng Zhao, Zilong Zheng
Hydra-MDP employs multi-teacher knowledge distillation with multi-head decoding for end-to-end multimodal planning, achieving first place in Navsim with significant generalization improvements.
Zhenxin Li, Kailin Li, Shihao Wang et al.
HumanEvo benchmark models software evolution, revealing performance overestimation by 10-61% in traditional evaluations.
Dewu Zheng, Yanlin Wang, Ensheng Shi et al.
Proposes Locally Interdependent Multi-Agent MDP with three closed-form near-optimal policies, performance improves exponentially with visibility radius.
Alex DeWeese, Guannan Qu
DISCOVERYWORLD is a virtual environment for evaluating AI scientific discovery, with 120 tasks across 8 themes, challenging end-to-end reasoning.
Peter Jansen, Marc-Alexandre Côté, Tushar Khot et al.
LlamaGen applies large language model's next-token prediction to image generation, achieving 2.18 FID on ImageNet with a 3.1B parameter model, surpassing diffusion models.
Peize Sun, Yi Jiang, Shoufa Chen et al.
Proposes UMBRELA, an open-source toolkit using GPT-4 for automatic relevance assessment, showing high correlation with human judgments.
Shivani Upadhyay, Ronak Pradeep, Nandan Thakur et al.
R2N boosts PbRL robustness via dynamic sparsity, improving 15 baseline-environment settings on noisy DMControl tasks.
Calarina Muslimani, Bram Grooten, Deepak Ranganatha Sastry Mamillapalli et al.
DeltaNet parallelizes linear transformers over sequence length, enhancing training efficiency and language modeling performance.
Songlin Yang, Bailin Wang, Yu Zhang et al.
Quantum Equilibrium Propagation (QEP) leverages Onsager reciprocity for efficient gradient estimation in quantum systems, enabling scalable quantum machine learning.
Clara C. Wanjura, Florian Marquardt
Combining weak instruments and observational data via a two-stage framework enables robust heterogeneous treatment effect estimation.
Miruna Oprescu, Nathan Kallus