LLM4AD: A Platform for Algorithm Design with Large Language Model
LLM4AD unifies LLM-driven algorithm search; EoH, FunSearch, and (1+1)-EPS beat random sampling on most of nine tasks.
Fei Liu, Rui Zhang, Zhuoliang Xie et al.
LLM4AD unifies LLM-driven algorithm search; EoH, FunSearch, and (1+1)-EPS beat random sampling on most of nine tasks.
Fei Liu, Rui Zhang, Zhuoliang Xie et al.
TG-OT employs topology-guided optimal transport for fully automatic CCTA-IVUS registration, achieving high accuracy without prior segmentation.
R. L. M. van Herten, José P. Henriques, R. Nils Planken et al.
Proposes SympFlow, a time-dependent symplectic neural network using parameterized Hamiltonian flows, improving energy conservation and long-term stability.
Priscilla Canizares, Davide Murari, Carola-Bibiane Schönlieb et al.
Proposed an input perturbation method 'forget vector' for machine unlearning without altering model weights.
Changchang Sun, Ren Wang, Yihua Zhang et al.
Task Shield reduces attack success rate to 2.07% while maintaining 69.79% task utility.
Feiran Jia, Tong Wu, Xin Qin et al.
Proposes a kernel-based method to approximate Koopman eigenfunctions without explicit operator computation, decomposing into linear and nonlinear parts for enhanced system analysis.
Jonghyeon Lee, Boumediene Hamzi, Boya Hou et al.
BODex uses bilevel optimization for efficient dexterous grasp synthesis, achieving over 75% success rate.
Jiayi Chen, Yubin Ke, He Wang
Introduces Causally Regularized Tokenization (CRT) to optimize visual token compression, enabling parameter and token reduction by half while maintaining state-of-the-art generation quality.
Vivek Ramanujan, Kushal Tirumala, Armen Aghajanyan et al.
Combining DreamBooth and contrastive learning, this method uses only 3 real images to generate synthetic data, significantly outperforming pretrained models across tasks.
Shobhita Sundaram, Julia Chae, Yonglong Tian et al.
OREO method enhances LLM multi-step reasoning, outperforming on GSM8K and MATH datasets.
Huaijie Wang, Shibo Hao, Hanze Dong et al.
LE-MCTS uses PRM-guided MCTS to ensemble reasoning steps, reaching 45.2% on MATH and 71.1% on MQA.
Sungjin Park, Xiao Liu, Yeyun Gong et al.
Aria-UI achieves GUI instruction grounding using a pure vision approach, improving accuracy significantly.
Yuhao Yang, Yue Wang, Dongxu Li et al.
Proposes Cache-Augmented Generation (CAG) leveraging long-context LLMs to preload knowledge, eliminating retrieval latency, and outperforming RAG in efficiency and accuracy.
Brian J Chan, Chao-Ting Chen, Jui-Hung Cheng et al.
LeviTor integrates depth estimation with K-means clustered control points to enable precise 3D trajectory control in image-to-video synthesis, outperforming prior 2D methods.
Hanlin Wang, Hao Ouyang, Qiuyu Wang et al.
Generative Multiview Relighting combines diffusion harmonization with NeRF-Casting, reaching 31.34 PSNR on Objaverse under extreme lighting variation.
Hadi Alzayer, Philipp Henzler, Jonathan T. Barron et al.
LMFusion freezes language modules and trains dedicated image modules, achieving 20% better image understanding and 3.6% improved image generation with 50% FLOPs.
Weijia Shi, Xiaochuang Han, Chunting Zhou et al.
STRAP combines pre-trained vision models and subsequence dynamic time warping for sub-trajectory retrieval, significantly enhancing robot policy generalization and robustness.
Marius Memmel, Jacob Berg, Bingqing Chen et al.
Prompt-A-Video leverages preference-aligned LLMs with reward-guided evolution, boosting video quality metrics by over 0.2 on average across models.
Yatai Ji, Jiacheng Zhang, Jie Wu et al.
SMORE model reduces noise through frequency domain fusion, enhancing multimodal recommendation accuracy.
Rongqing Kenneth Ong, Andy W. H. Khong
Efficient knowledge injection via self-distillation surpasses RAG and supervised fine-tuning.
Kalle Kujanpää, Pekka Marttinen, Harri Valpola et al.