RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
RLAIF uses AI feedback instead of human feedback, achieving performance comparable to RLHF.
Harrison Lee, Samrat Phatale, Hassan Mansoor et al.
RLAIF uses AI feedback instead of human feedback, achieving performance comparable to RLHF.
Harrison Lee, Samrat Phatale, Hassan Mansoor et al.
PointLLM integrates point cloud encoders with LLMs, achieving 53%+ zero-shot classification accuracy and surpassing human annotations in detailed object captioning.
Runsen Xu, Xiaolong Wang, Tai Wang et al.
Proposed InterDiff combines diffusion models with physics priors for long-term 3D human-object interaction prediction.
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang et al.
EMDB leverages electromagnetic sensors and deep neural models to create a high-precision 3D human pose and shape dataset in the wild, with global trajectories.
Manuel Kaufmann, Jie Song, Chen Guo et al.
TouchStone evaluates LVLMs using strong LLMs like GPT-4, covering five abilities and 27 subtasks.
Shuai Bai, Shusheng Yang, Jinze Bai et al.
InteRecAgent combines large language models with recommender tools for interactive recommendations, enhancing conversational systems.
Xu Huang, Jianxun Lian, Yuxuan Lei et al.
Introduced Epsilon Scaling to reduce exposure bias in diffusion models, achieving 2.17 FID on CIFAR-10.
Mang Ning, Mingxiao Li, Jianlin Su et al.
Using DCMM model to derive finite-sample expansion of node mixing probabilities for uncertainty quantification and ranking inference.
Sohom Bhattacharya, Jianqing Fan, Jikai Hou
LongBench is the first bilingual, multi-task benchmark for long context understanding, with 21 datasets averaging 6,711 words, advancing long-sequence model evaluation.
Yushi Bai, Xin Lv, Jiajie Zhang et al.
Confucius framework enhances LLM tool usage via easy-to-difficult curriculum and introspective feedback, surpassing existing baselines.
Shen Gao, Zhengliang Shi, Minghang Zhu et al.
Residual Denoising Diffusion Model (RDDM) introduces dual diffusion processes for unified image generation and restoration, leveraging residuals and noise with independent scheduling.
Jiawei Liu, Qiang Wang, Huijie Fan et al.
Proposes RL-assisted evolutionary algorithms (RL-EA), leveraging deep RL (DQN, PPO) to enhance optimization, outperforming traditional EA on benchmarks with 15% average improvement.
Yanjie Song, Yutong Wu, Yangyang Guo et al.
Nougat uses a Visual Transformer to convert academic PDFs into lightweight markup, significantly improving semantic retention of mathematical expressions.
Lukas Blecher, Guillem Cucurull, Thomas Scialom et al.
OmniQuant employs learnable clipping and transformation to enable high-performance low-bit quantization of LLMs, achieving superior results with minimal training.
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang et al.
Project Aria device uses multi-modal sensors to advance personalized AI research.
Jakob Engel, Kiran Somasundaram, Michael Goesele et al.
This paper critically evaluates multivariate time series anomaly detection, exposes flaws in point-adjust evaluation protocol, and demonstrates PCA-based baseline surpassing deep learning models on benchmarks.
Mohamed El Amine Sehili, Zonghua Zhang
Introduces SWIE and OVERMISS to improve translation faithfulness, achieving significant BLEU and trustworthiness gains.
Yijie Chen, Yijin Liu, Fandong Meng et al.
Introduces the Instruction-Following Difficulty (IFD) metric for self-guided data selection, boosting LLM instruction tuning with only 10% data, outperforming full-data models.
Ming Li, Yong Zhang, Zhitao Li et al.
This study analyzes LLMs' sensitivity to option order in MCQs, revealing up to 75% performance variation, and proposes calibration strategies to improve robustness.
Pouya Pezeshkpour, Estevam Hruschka
Unified LLM autonomous agent framework integrates perception, memory, planning, and action modules, significantly improving task performance in complex environments.
Lei Wang, Chen Ma, Xueyang Feng et al.