DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation
DiffusionCLIP uses diffusion models for text-guided image manipulation, enhancing performance on the ImageNet dataset.
Gwanghyun Kim, Taesung Kwon, Jong Chul Ye
DiffusionCLIP uses diffusion models for text-guided image manipulation, enhancing performance on the ImageNet dataset.
Gwanghyun Kim, Taesung Kwon, Jong Chul Ye
Proposes Anomaly Transformer using association discrepancy for unsupervised time series anomaly detection, achieving SOTA results.
Jiehui Xu, Haixu Wu, Jianmin Wang et al.
CLIP-Forge employs a two-stage training framework to generate 3D shapes from text without paired data, achieving high diversity and accuracy.
Aditya Sanghi, Hang Chu, Joseph G. Lambourne et al.
Proposes a language-conditioned waypoint prediction network, boosting instruction-guided navigation success by 4% in continuous environments.
Jacob Krantz, Aaron Gokaslan, Dhruv Batra et al.
Uses statistical physics models to analyze urban, traffic, financial, and social network phenomena, revealing underlying dynamics.
Marko Jusup, Petter Holme, Kiyoshi Kanazawa et al.
Proposes generalized kernel thinning (TARGET KT) with dimension-free error bounds, improving high-dimensional distribution compression.
Raaz Dwivedi, Lester Mackey
Proposes a unified framework integrating dense and sparse retrieval via logical scoring and physical retrieval models.
Jimmy Lin
XR-Transformer employs multi-resolution recursive fine-tuning, reducing training time from 23 days to 29 hours on Amazon-3M, with Prec@1 reaching 54%.
Jiong Zhang, Wei-cheng Chang, Hsiang-fu Yu et al.
Proposed a deep learning-based multi-layered architecture integrating perception, understanding, and prediction for robotic situational awareness, improving target detection and relation inference by 15%.
Hriday Bavle, Jose Luis Sanchez-Lopez, Claudio Cimarelli et al.
CrossCLR introduces influence-based sample filtering and weighting, significantly improving multi-modal video-text embeddings with up to 4% higher retrieval accuracy.
Mohammadreza Zolfaghari, Yi Zhu, Peter Gehler et al.
ALIP and MPC-based gait controller enhances bipedal robot agility on complex terrains.
Grant Gibson, Oluwami Dosunmu-Ogunbi, Yukai Gong et al.
Introduces ViPTL method for learning periodic tasks from visual demonstrations using rDMPs and Bayesian optimization.
Jingyun Yang, Junwu Zhang, Connor Settle et al.
PDC-Net+ uses a probabilistic model to achieve accurate dense correspondence and confidence estimation, improving geometric matching performance.
Prune Truong, Martin Danelljan, Radu Timofte et al.
Multiwavelet neural operator compresses kernels for high-accuracy PDE solutions, achieving 2-10x error reduction.
Gaurav Gupta, Xiongye Xiao, Paul Bogdan
Proposes verification error as a key metric for machine unlearning; analyzes SGD to design low-error training objectives.
Anvith Thudi, Gabriel Deza, Varun Chandrasekaran et al.
Proposes a controllable neural dialogue summarization framework with personal entity planning, improving factual consistency and diversity; achieves ROUGE-2 of 58.7 on SAMSum.
Zhengyuan Liu, Nancy F. Chen
MetaDrive employs procedural generation and real data import to create diverse driving scenarios, enhancing RL generalization in autonomous driving.
Quanyi Li, Zhenghao Peng, Lan Feng et al.
Linear policies enable robust bipedal walking on challenging terrains with only 13 learnable parameters.
Lokesh Krishna, Guillermo A. Castillo, Utkarsh A. Mishra et al.
Proposes neural network-based model reduction for MPM, enabling continuous deformation mapping and real-time large-scale simulation.
Peter Yichen Chen, Maurizio M. Chiaramonte, Eitan Grinspun et al.
Formal taxonomy of resilience approaches in multi-robot systems emphasizing system reconfiguration, adaptation, and growth.
Amanda Prorok, Matthew Malencia, Luca Carlone et al.