Language Model Decoding as Direct Metrics Optimization
DAEMON formulates decoding as a multi-metric energy optimization, guaranteeing perplexity improvement and better alignment with human texts.
Haozhe Ji, Pei Ke, Hongning Wang et al.
DAEMON formulates decoding as a multi-metric energy optimization, guaranteeing perplexity improvement and better alignment with human texts.
Haozhe Ji, Pei Ke, Hongning Wang et al.
RoleLLM enhances role-playing in LLMs via a four-stage framework, creating the RoleBench dataset.
Zekun Moore Wang, Zhongyuan Peng, Haoran Que et al.
Gradient Refined Proposal enhances rejection sampling, boosting acceptance rate up to 7.3× with minimal assumptions.
Edward Raff, Mark McLean, James Holt
Corex employs multi-model collaboration (Discuss, Review, Retrieve) to significantly improve complex reasoning, outperforming existing baselines with detailed experimental validation.
Qiushi Sun, Zhangyue Yin, Xiang Li et al.
LVD leverages large language models to generate scene layouts guiding diffusion-based video synthesis, achieving high fidelity without training.
Long Lian, Baifeng Shi, Adam Yala et al.
Analyzed GPT-4V's multimodal capabilities, demonstrating superior performance in multi-domain tasks with visual marker understanding and prompting strategies.
Zhengyuan Yang, Linjie Li, Kevin Lin et al.
Universal speech enhancement model USES, independent of channels, length, and sampling rate, achieves strong multi-condition performance.
Wangyou Zhang, Kohei Saijo, Zhong-Qiu Wang et al.
Introduces RIS-CQ benchmark and DUMOGA model, boosting complex semantic understanding and target localization by over 200%.
Wei Ji, Li Li, Hao Fei et al.
DeltaXplainer dynamically explains model differences via decision rules, enhancing model selection and monitoring efficiency.
Adam Rida, Marie-Jeanne Lesot, Xavier Renard et al.
HoloAssist is a large-scale egocentric dataset combining multimodal data for training interactive AI assistants in real-world collaborative tasks.
Xin Wang, Taein Kwon, Mahdi Rad et al.
Introduces COBBLER benchmark to evaluate biases in 16 large language models, revealing prevalent cognitive biases affecting evaluation robustness.
Ryan Koo, Minhwa Lee, Vipul Raheja et al.
DreamGaussian uses generative Gaussian splatting to produce textured 3D meshes in 2 minutes, achieving 10x faster than prior methods.
Jiaxiang Tang, Jiawei Ren, Hang Zhou et al.
ConceptGraphs combines multi-view geometric fusion and large vision-language models to build open-vocabulary 3D scene graphs for robot perception and planning.
Qiao Gu, Alihusein Kuwajerwala, Sacha Morin et al.
Qwen series models excel in multiple tasks, notably Qwen-Chat with RLHF technology.
Jinze Bai, Shuai Bai, Yunfei Chu et al.
Proposes a doubly robust kernel Stein discrepancy estimator for unnormalized density-based counterfactual distribution modeling.
Diego Martinez-Taboada, Edward H. Kennedy
Systematic analysis of activation patching metrics and methods; STR, logit difference, and sliding window outperform alternatives.
Fred Zhang, Neel Nanda
Introduces RBRP algorithm for fast and accurate adaptive interpolative decomposition.
Yijun Dong, Chao Chen, Per-Gunnar Martinsson et al.
Show-1 combines pixel and latent diffusion models for efficient text-to-video generation, using only 15G GPU memory.
David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu et al.
Proposes a benchmark-based model routing framework using binary classifiers to improve large language model selection, achieving significant performance gains across 29 datasets.
Tal Shnitzer, Anthony Ou, Mírian Silva et al.
PolarNet leverages 3D point clouds and multimodal transformers for language-guided robotic manipulation, outperforming 2D methods with 92.1% success in single-task and 60% on real robots.
Shizhe Chen, Ricardo Garcia, Cordelia Schmid et al.