Large Language Models Are State-of-the-Art Evaluators of Translation Quality
GEMBA leverages GPT-3.5+ for state-of-the-art translation quality evaluation, outperforming existing metrics on WMT22 data.
Tom Kocmi, Christian Federmann
GEMBA leverages GPT-3.5+ for state-of-the-art translation quality evaluation, outperforming existing metrics on WMT22 data.
Tom Kocmi, Christian Federmann
Proposed MV-DTSA framework maps time series to binary images, enhancing forecasting accuracy.
Luoxiao Yang, Xinqi Fan, Zijun Zhang
Localizing moments in long videos using multimodal guidance improves performance on MAD and Ego4D datasets.
Wayner Barrios, Mattia Soldan, Alberto Mario Ceballos-Arroyo et al.
Proposes a probabilistic roadmap-based method for contact-rich manipulation, enabling heavy object movement with environmental support, without explicit contact mode analysis.
Kento Nakatsuru, Weiwei Wan, Kensuke Harada
AugGPT leverages ChatGPT for data augmentation, significantly boosting few-shot text classification accuracy by generating diverse, semantically consistent samples.
Haixing Dai, Zhengliang Liu, Wenxiong Liao et al.
Voltron framework enhances robot learning via language-driven visual representation, excelling in high-level feature tasks.
Siddharth Karamcheti, Suraj Nair, Annie S. Chen et al.
SAN leverages a lightweight side network attached to frozen CLIP, achieving open-vocabulary semantic segmentation with only 8.4M parameters and 19× faster inference.
Mengde Xu, Zheng Zhang, Fangyun Wei et al.
Introduces Indirect Prompt Injection (IPI) attacks exploiting retrieval data to remotely control LLMs, enabling data theft, content manipulation, and API abuse.
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra et al.
FPT freezes GPT-2 blocks for seven time-series tasks, reaching 0.516 average long-horizon MSE and 74.0% classification accuracy.
Tian Zhou, PeiSong Niu, Xue Wang et al.
DeepSORVF employs deep learning for asynchronous trajectory matching, achieving 95% accuracy in vessel data fusion under complex conditions.
Yu Guo, Ryan Wen Liu, Jingxiang Qu et al.
Introduces the 'leap' measure to quantify hierarchical complexity, analyzing SGD training time on low-dimensional functions with neural networks.
Emmanuel Abbe, Enric Boix-Adsera, Theodor Misiakiewicz
Hyena, a subquadratic operator combining implicit long convolution and gating, achieves Transformer-level performance on long sequences with 20% less compute.
Michael Poli, Stefano Massaroli, Eric Nguyen et al.
Unifies multicalibration via game dynamics, reducing complexity and improving guarantees in multi-objective learning.
Nika Haghtalab, Michael I. Jordan, Eric Zhao
Hint-ReLIC method enhances neural algorithmic reasoning's OOD generalization with causal regularization, achieving 3x improvement on CLRS benchmark.
Beatrice Bevilacqua, Kyriacos Nikiforou, Borja Ibarz et al.
Survey on large-scale multi-modal pre-trained models, analyzing data, architecture, and experimental results.
Xiao Wang, Guangyao Chen, Guangwu Qian et al.
Proposed a general FTRL/OMD algorithm achieving best performance in adversarial and stochastic settings for linear bandits.
Christoph Dann, Chen-Yu Wei, Julian Zimmert
FLE algorithm uses MLE with probabilistic models to estimate return distributions in offline RL, providing error guarantees.
Runzhe Wu, Masatoshi Uehara, Wen Sun
Proposes a comprehensive safety evaluation framework combining preference testing, adversarial attacks, and risk detection to enhance large model safety.
Jiawen Deng, Jiale Cheng, Hao Sun et al.
This study systematically evaluates GPT models' performance in multilingual machine translation, revealing strong results in high-resource languages but limitations in low-resource scenarios.
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf et al.
Proposes three pacing algorithms; min-pacing achieves near-optimal guarantees with low constraint violations.
Santiago R. Balseiro, Kshipra Bhawalkar, Zhe Feng et al.