Gen-L-Video: Multi-Text to Long Video Generation via Temporal Co-Denoising
Gen-L-Video extends short-video diffusion through temporal co-denoising, achieving 93.18 frame consistency.
Fu-Yun Wang, Wenshuo Chen, Guanglu Song et al.
Gen-L-Video extends short-video diffusion through temporal co-denoising, achieving 93.18 frame consistency.
Fu-Yun Wang, Wenshuo Chen, Guanglu Song et al.
This study evaluates LFQA with human experts and automatic metrics, revealing no single metric reliably predicts human preferences.
Fangyuan Xu, Yixiao Song, Mohit Iyyer et al.
Marked Personas employs natural language prompts to unsupervisedly quantify intersectional stereotypes in LLM outputs, revealing higher bias levels in GPT-3.5 and GPT-4 compared to human descriptions.
Myra Cheng, Esin Durmus, Dan Jurafsky
VAST model uses the VAST-27M dataset to achieve omni-modality video understanding, setting 22 new SOTA results.
Sihan Chen, Handong Li, Qunbo Wang et al.
Proposes TTT-NN, a test-time fine-tuning method using nearest neighbor retrieval, improving over 20 tasks with only one gradient step on 20 neighbors.
Moritz Hardt, Yu Sun
Proposed a calibration framework to address positional bias in LLM evaluations, improving human alignment by 14.3%.
Peiyi Wang, Lei Li, Liang Chen et al.
LLM-QAT achieves 4-bit quantization via data-free distillation, enhancing large language model performance.
Zechun Liu, Barlas Oguz, Changsheng Zhao et al.
Proposes Multi-Task Diffusion Model (MTDiff) with Transformer and prompt learning for multi-task offline RL, outperforming state-of-the-art on Meta-World and Maze2D.
Haoran He, Chenjia Bai, Kang Xu et al.
AdaSketch-Newton combines exact augmented-Lagrangian line search with randomized sketching, proving almost-sure global and local linear/superlinear convergence.
Ilgee Hong, Sen Na, Michael W. Mahoney et al.
Proposes GRIT, a message-passing-free graph Transformer leveraging inductive biases, achieving state-of-the-art results on multiple benchmarks.
Liheng Ma, Chen Lin, Derek Lim et al.
Proposed FactFormer uses axial factorized kernel integral in Transformer for scalable high-dimensional PDE surrogate modeling.
Zijie Li, Dule Shu, Amir Barati Farimani
MeZO, a memory-efficient zeroth-order optimizer, enables fine-tuning of 30B models with 12× less memory, matching performance of backpropagation.
Sadhika Malladi, Tianyu Gao, Eshaan Nichani et al.
LATM framework uses GPT-4 to generate reusable Python tools, reducing inference costs significantly.
Tianle Cai, Xuezhi Wang, Tengyu Ma et al.
NavGPT leverages large language models for explicit reasoning in vision-and-language navigation, demonstrating zero-shot planning with path decomposition and landmark recognition.
Gengze Zhou, Yicong Hong, Qi Wu
Proposed Supervised Attention MIL (SAMIL) for multi-view ultrasound heart disease detection, achieving over 70% accuracy.
Zhe Huang, Benjamin S. Wessler, Michael C. Hughes
Banana network uses Banach fixed-point for pointcloud segmentation with inter-part equivariance, enhancing segmentation accuracy.
Congyue Deng, Jiahui Lei, Bokui Shen et al.
Proposes DPOK, a reinforcement learning framework with KL regularization for online fine-tuning of diffusion models, improving text-image alignment and image quality.
Ying Fan, Olivia Watkins, Yuqing Du et al.
Proposes an end-to-end differentiable Transformer Neural Process framework for meta Bayesian optimization, leveraging reinforcement learning to improve sample efficiency.
Alexandre Maraval, Matthieu Zimmer, Antoine Grosnit et al.
Proposes MeLoDy, an efficient diffusion-based neural music generator reducing inference steps by 95.7%/99.6%, maintaining high quality.
Max W. Y. Lam, Qiao Tian, Tang Li et al.
Plasticity injection enhances neural network adaptability in deep RL, boosting performance by 20% on Atari without increasing trainable parameters.
Evgenii Nikishin, Junhyuk Oh, Georg Ostrovski et al.