nach0-pc: Multi-task Language Model with Molecular Point Cloud Encoder
nach0-pc integrates point cloud encoder with T5, enabling multi-task 3D molecular generation with high efficiency.
Maksim Kuznetsov, Airat Valiev, Alex Aliper et al.
nach0-pc integrates point cloud encoder with T5, enabling multi-task 3D molecular generation with high efficiency.
Maksim Kuznetsov, Airat Valiev, Alex Aliper et al.
EGAD enhances Longformer with keyword-based global attention, boosting long-range dependency modeling for abstractive summarization.
Evan Lucas, Dylan Kangas, Timothy C Havens
HyperPG outperforms Euclidean prototypes on CUB-200-2011 with simplified training.
Maximilian Xiling Li, Korbinian Franz Rudolf, Paul Mattes et al.
This paper systematically reviews the roles of large language models in algorithm design, including optimizer, predictor, extractor, and designer, across three stages.
Fei Liu, Yiming Yao, Ping Guo et al.
Mycroft combines feature distance and gradient matching to efficiently select informative data subsets from private sources, boosting model performance with minimal data sharing.
Zain Sarwar, Van Tran, Arjun Nitin Bhagoji et al.
Mono-InternVL embeds visual experts into a pre-trained LLM with EViP, achieving superior multi-modal performance, surpassing 13 benchmarks.
Gen Luo, Xue Yang, Wenhan Dou et al.
DRAFT employs self-driven feedback to iteratively refine tool documentation, enhancing LLM understanding and utilization.
Changle Qu, Sunhao Dai, Xiaochi Wei et al.
SG-Nav uses online 3D scene graphs and hierarchical reasoning to boost zero-shot object navigation SR by over 10%.
Hang Yin, Xiuwei Xu, Zhenyu Wu et al.
RDT-1B model uses diffusion and Transformer to learn bimanual manipulation with multimodal data.
Songming Liu, Lingxuan Wu, Bangguo Li et al.
Proposes MinorityPrompt, an online prompt optimization method enhancing low-likelihood sample generation in diffusion models.
Soobin Um, Jong Chul Ye
IterComp leverages multi-model preferences and iterative feedback to enhance compositional text-to-image generation, outperforming SOTA methods.
Xinchen Zhang, Ling Yang, Guohao Li et al.
ReinDiffuse combines diffusion models with reinforcement learning, achieving 29% and 34% improvements in FID on HumanML3D and KIT-ML, respectively, for physically plausible motion.
Gaoge Han, Mingjiang Liang, Jinglei Tang et al.
DreamMesh4D combines mesh and Gaussian points for high-quality video-to-4D generation, achieving superior spatial-temporal consistency.
Zhiqi Li, Yiming Chen, Peidong Liu
Herald: Translates Mathlib4 to natural language using dual augmentation, enhancing LLM performance in mathematical reasoning.
Guoxiong Gao, Yutong Wang, Jiedong Jiang et al.
Using irreducible representations and Schur’s lemma, this paper systematically characterizes permutation-equivariant layers, simplifying derivations for DeepSets, graph networks, and weight spaces.
Yonatan Sverdlov, Ido Springer, Nadav Dym
TorchTitan supports 4D parallelism, boosting Llama 3.1 training by 65.08% with modular design and hardware co-optimization.
Wanchao Liang, Tianyu Liu, Less Wright et al.
Proposes MoLAS and AgentSquare, using module evolution and recombination to optimize LLM agents with a 17.2% performance boost.
Yu Shang, Yu Li, Keyu Zhao et al.
RepLDM achieves high-quality, efficient high-resolution image generation via attention guidance and progressive upsampling.
Boyuan Cao, Jiaxin Ye, Yujie Wei et al.
ACSSM combines multi-marginal Doob transform and stochastic optimal control for irregular time series modeling.
Byoungwoo Park, Hyungi Lee, Juho Lee
Proposes Gen-Drive, a diffusion-based generative driving policy with reward modeling and RL fine-tuning, achieving state-of-the-art planning scores on nuPlan.
Zhiyu Huang, Xinshuo Weng, Maximilian Igl et al.