Top-Down Synthesis for Library Learning
Proposes Stitch, a top-down corpus-guided synthesis algorithm, achieving 3-4 orders faster and 20x less memory than prior methods.
Matthew Bowers, Theo X. Olausson, Lionel Wong et al.
Proposes Stitch, a top-down corpus-guided synthesis algorithm, achieving 3-4 orders faster and 20x less memory than prior methods.
Matthew Bowers, Theo X. Olausson, Lionel Wong et al.
Proposes point-language association via multi-view captioning and contrastive learning, boosting open-vocabulary 3D scene understanding by 25.8%-44.7%.
Runyu Ding, Jihan Yang, Chuhui Xue et al.
Probabilistic time series forecasting using quantile regression and ensemble methods, evaluated on weather and financial data.
Johannes Bracher, Nils Koster, Fabian Krüger et al.
Applying conditional diffusion models (Decision Diffuser) for offline decision-making surpasses traditional RL, enabling direct trajectory sampling with multi-condition support.
Anurag Ajay, Yilun Du, Abhi Gupta et al.
OpenScene employs CLIP embedding for zero-shot 3D scene understanding, enabling open-vocabulary queries with dense point features.
Songyou Peng, Kyle Genova, Chiyu "Max" Jiang et al.
UniD3 model achieves multimodal generation via unified discrete diffusion, matching SOTA performance.
Minghui Hu, Chuanxia Zheng, Heliang Zheng et al.
Proposes RbA, a region-level outlier scoring method based on 'rejected by all,' improving unknown object segmentation with minimal supervision.
Nazir Nayal, Mısra Yavuz, João F. Henriques et al.
CoMFormer leverages transformer architecture with adaptive distillation and pseudo-labeling to enable continual semantic and panoptic segmentation, outperforming existing methods with less forgetting.
Fabio Cermelli, Matthieu Cord, Arthur Douillard
Proposes Semantic Completion Learning (SCL) to enhance global-to-local cross-modal alignment, achieving SOTA results on vision-language benchmarks.
Yatai Ji, Rongcheng Tu, Jie Jiang et al.
LaCAM is a search-based multi-agent pathfinding algorithm capable of solving hundreds of agents quickly with high success rates.
Keisuke Okumura
Proposes LVDM, a latent diffusion model enabling high-fidelity long video synthesis, outperforming pixel-space models.
Yingqing He, Tianyu Yang, Yong Zhang et al.
Proposes a training-free 'Plug-and-Play' diffusion feature injection for text-guided image translation, preserving semantic layout.
Narek Tumanyan, Michal Geyer, Shai Bagon et al.
EDICT employs coupled affine transformations for exact diffusion inversion, reducing reconstruction error by over 50% compared to DDIM.
Bram Wallace, Akash Gokul, Nikhil Naik
Optimization-based control using contact modeling and model simplification enables real-time dynamic legged robot motion.
Patrick M. Wensing, Michael Posa, Yue Hu et al.
InstructPix2Pix leverages GPT-3 and Stable Diffusion to generate 450K+ training pairs for instruction-based image editing.
Tim Brooks, Aleksander Holynski, Alexei A. Efros
Proposes Null-text inversion for high-fidelity real image editing using guided diffusion, achieving PSNR >30dB without model fine-tuning.
Ron Mokady, Amir Hertz, Kfir Aberman et al.
PromptInject framework reveals GPT-3's vulnerability to goal hijacking (58.6%) and prompt leaking (23.6%) via adversarial prompts.
Fábio Perez, Ian Ribeiro
DexPoint achieves cross-object sim-to-real transfer using point clouds and dexterous hands, with success rates over 80%.
Yuzhe Qin, Binghao Huang, Zhao-Heng Yin et al.
Proposes TART, a task-aware retrieval system using multi-task instruction tuning, excelling in zero-shot and cross-domain scenarios.
Akari Asai, Timo Schick, Patrick Lewis et al.
Proposes GO-UCB, a parametric model-based method achieving \~O(√T) regret for high-dimensional global optimization.
Chong Liu, Yu-Xiang Wang