Stable Flow: Vital Layers for Training-Free Image Editing
Stable Flow finds vital layers in DiT to enable training-free, stable image editing.
Omri Avrahami, Or Patashnik, Ohad Fried et al.
Stable Flow finds vital layers in DiT to enable training-free, stable image editing.
Omri Avrahami, Or Patashnik, Ohad Fried et al.
DINO-X achieves state-of-the-art open-world object detection with 56.0 AP on COCO.
Tianhe Ren, Yihao Chen, Qing Jiang et al.
Introduces MorphScore to evaluate morphological alignment across 22 languages; finds dataset size and encoding efficiency outweigh morphological complexity in influencing model performance.
Catherine Arnett, Benjamin K. Bergen
Developed constant-factor approximation algorithms for Nash social welfare with capacities: 6+ε for submodular preferences, 1.33 for subadditive in two-sided models.
Salil Gokhale, Harshul Sagar, Rohit Vaish et al.
Hymba employs a hybrid-head architecture combining transformer attention and state space models, achieving superior efficiency and performance with fewer parameters, surpassing comparable small models.
Xin Dong, Yonggan Fu, Shizhe Diao et al.
EAST combines dynamic ReLU, weight sharing, and cyclic sparsity to achieve 99.99% sparsity with competitive accuracy.
Andy Li, Aiden Durrant, Milan Markovic et al.
Introduced CAFE dataset with spontaneous Algerian dialect, French, English speech; improved Whisper-based ASR with decoding strategies achieving MER 0.310.
Houssam Eddine-Othman Lachemat, Akli Abbas, Nourredine Oukas et al.
Proposes offline R/R* metric and scaling law for online ad retrieval models, enabling cost-effective performance prediction.
Yunli Wang, Zhen Zhang, Zixuan Yang et al.
Video-RAG integrates external visual-aligned texts via retrieval, no training needed, boosting long video comprehension by 2.8% on average.
Yongdong Luo, Xiawu Zheng, Guilin Li et al.
RONAR system leverages large language models to generate natural language descriptions of robot multi-modal experiences, improving failure analysis by 11%.
Zihan Wang, Brian Liang, Varad Dhat et al.
Proposed Geo-Time Re-ranking (GT-R) model using LLMs, achieving 15%-20% improvement on LEO dataset for event recommendation.
Yuanyuan Tian, Wenwen Li, Lei Hu et al.
Post-hoc regression-based λ optimization reduces variance, improving PPI in few-label settings.
Benjamin Eyre, David Madras
Extending LRNNs' eigenvalues to [-1,1] enables robust state-tracking, significantly improving performance on parity and modular tasks.
Riccardo Grazzi, Julien Siems, Arber Zela et al.
Enhancing LLaMA-3.1-70B's reasoning via Additional Logic Training (ALT) with FLD×2 corpus, achieving 30-point boost in logic benchmarks.
Terufumi Morishita, Gaku Morio, Atsuki Yamaguchi et al.
Analyzed 100,000 user logs to reveal domain preferences and translation directions in Tetun low-resource MT, emphasizing community needs.
Raphael Merx, Adérito José Guterres Correia, Hanna Suominen et al.
Proposed spatial-aware adversarial patches drastically reduce robot task success, up to 100%, revealing vulnerabilities in VLA models.
Taowen Wang, Cheng Han, James Chenhao Liang et al.
Proposes RandSymKL, combining symmetric KL and cross-entropy, to mitigate extrinsic gender bias in Bangla classification tasks, achieving 90.66% accuracy and bias reduction.
Sajib Kumar Saha Joy, Arman Hassan Mahy, Meherin Sultana et al.
USP-Gaussian unifies spike-based image reconstruction, pose correction, and Gaussian splatting for enhanced 3D reconstruction accuracy.
Kang Chen, Jiyuan Zhang, Zecheng Hao et al.
This study evaluates pre-trained visual representations (PVR) in model-based reinforcement learning (MBRL), revealing limited improvements in sample efficiency and out-of-distribution generalization.
Moritz Schneider, Robert Krug, Narunas Vaskevicius et al.
This study uses self-report grounded LLM agents, combining interviews and surveys, to predict individual responses across multiple outcomes with 83-86% accuracy without task-specific training.
Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst et al.