Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
Proposes DS-MoE: dense training with sparse inference, achieving high parameter and computational efficiency.
Bowen Pan, Yikang Shen, Haokun Liu et al.
Proposes DS-MoE: dense training with sparse inference, achieving high parameter and computational efficiency.
Bowen Pan, Yikang Shen, Haokun Liu et al.
AutoCodeRover combines AST search, LLM agents, and SBFL to solve 19% of SWE-bench-lite at $0.43 per issue.
Yuntong Zhang, Haifeng Ruan, Zhiyu Fan et al.
CGDF: A diffusion-based method for dense, constrained 6-DoF grasping on complex shapes, achieving over 60% success in challenging scenarios.
Gaurav Singh, Sanket Kalwar, Md Faizal Karim et al.
SiD distills pretrained diffusion models into one-step generators with near-exponential FID drops and teacher-level quality.
Mingyuan Zhou, Huangjie Zheng, Zhendong Wang et al.
AutoWebGLM leverages HTML simplification and reinforcement learning to create a web navigation agent surpassing GPT-4 in performance.
Hanyu Lai, Xiao Liu, Iat Long Iong et al.
Proposes a multi-level irrelevant information framework using Wikipedia, evaluates LLM robustness; finds models are easily misled by semantically related distractors.
Siye Wu, Jian Xie, Jiangjie Chen et al.
VAR employs multi-scale autoregressive prediction, reducing FID to 1.73 and increasing speed 20×, surpassing diffusion models.
Keyu Tian, Yi Jiang, Zehuan Yuan et al.
This paper derives tight stability bounds for entropic Brenier maps, optimizing Lipschitz constants under support constraints, with applications to semi-discrete optimal transport.
Vincent Divol, Jonathan Niles-Weed, Aram-Alexandre Pooladian
I-Design employs multi-agent LLMs to generate personalized 3D interior scenes, integrating scene graph layout and asset retrieval for high-quality results.
Ata Çelen, Guo Han, Konrad Schindler et al.
Mixture-of-Depths method dynamically allocates compute, boosting Transformer inference speed by 50%.
David Raposo, Sam Ritter, Blake Richards et al.
Eurus models, finetuned with UltraInteract preference trees, achieve state-of-the-art open-source reasoning performance, surpassing GPT-3.5 Turbo in complex tasks.
Lifan Yuan, Ganqu Cui, Hanbin Wang et al.
Argues scenario concept completeness using Goal Structured Notation, applied to inD dataset.
Christoph Glasmacher, Hendrik Weber, Lutz Eckstein
Enhancing 6-DoF grasp detection generalization via domain prior knowledge, achieving significant improvement on GraspNet-1billion.
Haoxiang Ma, Modi Shi, Boyang Gao et al.
Introduces a triangular modulation-based Doppler velocity extraction method for spinning FMCW radar, significantly improving odometry robustness in degenerate environments.
Daniil Lisus, Keenan Burnett, David J. Yoon et al.
QuAD enhances autonomous driving planning by querying occupancy at relevant spatio-temporal points for improved safety and interpretability.
Sourav Biswas, Sergio Casas, Quinlan Sykora et al.
Proposes neural implicit-based method for reconstructing digital twins of unknown multi-part articulated objects from two RGB-D scans, outperforming prior approaches.
Yijia Weng, Bowen Wen, Jonathan Tremblay et al.
Proposes a multi-label contrastive learning framework for style descriptors, achieving state-of-the-art style retrieval accuracy.
Gowthami Somepalli, Anubhav Gupta, Kamal Gupta et al.
VQAScore leverages VQA models to evaluate text-image alignment, surpassing CLIPScore with state-of-the-art results on 8 benchmarks.
Zhiqiu Lin, Deepak Pathak, Baiqi Li et al.
Using Liang et al.'s (2024) distributional GPT framework, analyzed 950,965 papers (2020-2024), revealing rapid growth of LLM-modified content, especially in CS (up to 17.5%).
Weixin Liang, Yaohui Zhang, Zhengxuan Wu et al.
EvLight leverages multi-scale fusion and SNR-guided feature selection, achieving +1.14dB PSNR over frame-based methods on large real-world event-image datasets.
Guoqiang Liang, Kanghao Chen, Hangyu Li et al.