cs.CV 2406.09414

Depth Anything V2

Depth Anything V2 leverages synthetic images, larger teacher models, and large-scale pseudo-labels to achieve finer, more robust monocular depth estimation, over 10x faster than prior models.

Lihe Yang, Bingyi Kang, Zilong Huang et al.

2024-06-14 37
cs.CL 2406.09393

Improving Autoregressive Training with Dynamic Oracles

Proposes dynamic oracle-guided autoregressive training, improving NER and summarization, with exact and approximate algorithms for specific metrics.

Jianing Yang, Harshine Visvanathan, Yilin Wang et al.

2024-06-14 43
cs.RO 2406.08545

RVT-2: Learning Precise Manipulation from Few Demonstrations

RVT-2 employs multi-stage virtual view reasoning and system-level optimizations to achieve 82% success in high-precision multi-task manipulation with only 10 demonstrations, training 6× faster than RVT.

Ankit Goyal, Valts Blukis, Jie Xu et al.

2024-06-13 195 citations 49
cs.CL 2406.08446

OLMES: A Standard for Language Model Evaluations

OLMES standardizes large language model evaluation, covering prompt formatting, example selection, probability normalization, ensuring reproducibility and fairness.

Yuling Gu, Oyvind Tafjord, Bailey Kuehl et al.

2024-06-13 100 citations 33