cs.CV 2406.08035

LVBench: An Extreme Long Video Understanding Benchmark

LVBench evaluates long video understanding, focusing on long-term memory and reasoning, with a dataset averaging 4101 seconds, covering six core tasks.

Weihan Wang, Zehai He, Wenyi Hong et al.

2024-06-12 45
cs.LG 2406.07887

An Empirical Study of Mamba-based Language Models

This study compares 8B Mamba, Mamba-2, and Transformer models trained on up to 3.5T tokens, showing hybrid models outperform in speed and long-sequence tasks.

Roger Waleffe, Wonmin Byeon, Duncan Riach et al.

2024-06-12 41
cs.CL 2406.07524

Simple and Effective Masked Diffusion Language Models

This paper introduces Masked Diffusion Language Models (MDLM), achieving near state-of-the-art perplexity by combining efficient training, Rao-Blackwellized objectives, and semi-autoregressive sampling.

Subham Sekhar Sahoo, Marianne Arriola, Yair Schiff et al.

2024-06-12 837 citations 27
cs.CL 2406.07496

TextGrad: Automatic "Differentiation" via Text

TextGrad employs natural language feedback as 'gradients' to optimize complex AI systems, improving performance across tasks like QA, coding, and drug design.

Mert Yuksekgonul, Federico Bianchi, Joseph Boen et al.

2024-06-12 51