Gecko: Versatile Text Embeddings Distilled from Large Language Models
Gecko employs a two-step distillation from LLMs, achieving 66.31 on MTEB with only 256 dimensions, outperforming larger models.
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren et al.
Gecko employs a two-step distillation from LLMs, achieving 66.31 on MTEB with only 256 dimensions, outperforming larger models.
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren et al.
ECLIPSE leverages Visual Prompt Tuning, freezing backbone parameters and fine-tuning prompts, achieving state-of-the-art continual panoptic segmentation with only 1.3% trainable parameters.
Beomyoung Kim, Joonsang Yu, Sung Ju Hwang
MagicLens uses self-supervised learning on 36.7M web triplets to support open-ended image retrieval, outperforming SOTA.
Kai Zhang, Yi Luan, Hexiang Hu et al.
Propose GraspXL, a reinforcement learning framework that generates multi-object grasping motions for over 500k unseen objects with 82.2% success, without relying on 3D hand-object data.
Hui Zhang, Sammy Christen, Zicong Fan et al.
Proposes data-adaptive risk control via RRR, supporting monotone risks with theoretical guarantees.
Drew T. Nguyen, Reese Pathak, Anastasios N. Angelopoulos et al.
Introduced inconsistency-guided detail regularization, enhancing mask-guided matting accuracy, surpassing SOTA methods.
Weihao Jiang, Zhaozhi Xie, Yuxiang Lu et al.
Proposed a model predictive control framework with discrete high-order control barrier functions for collision-free trajectory planning in dynamic environments.
Shuo Liu, Yihui Mao, Calin A. Belta
ObjectDrop uses counterfactual datasets and diffusion model fine-tuning to achieve photorealistic object removal and insertion, outperforming prior methods.
Daniel Winter, Matan Cohen, Shlomi Fruchter et al.
Mini-Gemini employs dual visual encoders and patch info mining to enhance high-res visual understanding, surpassing private models in zero-shot benchmarks.
Yanwei Li, Yuechen Zhang, Chengyao Wang et al.
Proposes a risk-aware branch MPC using CVaR and augmented iLQR for robust autonomous driving under uncertain vehicle behaviors.
Luyao Zhang, George Pantazis, Shaohang Han et al.
Proposed a two-stage diffusion-based framework utilizing scene affordance for language-guided human motion synthesis, outperforming baselines on benchmark datasets.
Zan Wang, Yixin Chen, Baoxiong Jia et al.
AgentStudio offers a unified multi-modal environment with tools and benchmarks to evaluate general virtual agents' core abilities, including GUI grounding, video learning, and success detection.
Longtao Zheng, Zhiyuan Huang, Zhenghai Xue et al.
HOV-SG constructs hierarchical open-vocabulary 3D scene graphs, achieving 12-15% accuracy improvements and 75% storage reduction for language-grounded robot navigation.
Abdelrhman Werby, Chenguang Huang, Martin Büchner et al.
Systematic analysis of 448 jailbreak prompts, proposing two effective strategies and an automated generation system.
Zhiyuan Yu, Xiaogeng Liu, Shunning Liang et al.
QKFormer introduces a linear-complexity Q-K attention for hierarchical SNNs, achieving 85.65% top-1 accuracy on ImageNet with 64.96M parameters.
Chenlin Zhou, Han Zhang, Zhaokun Zhou et al.
Proposed GoodSAM framework combines DAR and MKA modules to transfer SAM's zero-shot instance segmentation for panoramic semantic segmentation, achieving +3.75% mIoU improvement.
Weiming Zhang, Yexin Liu, Xu Zheng et al.
CoverUp combines coverage analysis and LLM feedback, boosting Python test coverage from 47% to 89%.
Juan Altmayer Pizzorno, Emery D. Berger
PruMerge adaptively prunes and merges visual tokens, achieving 14× compression while maintaining performance.
Yuzhang Shang, Mu Cai, Bingxin Xu et al.
End-to-end onboard gesture-based UAV formation control integrating detection, tracking, and human-swarm interaction, enabling real-time adaptive safety monitoring.
Vít Krátký, Giuseppe Silano, Matouš Vrba et al.
CRPlace fuses multi-view camera images and radar point clouds using BEV representation, achieving 91.2% recall@1 for place recognition.
Shaowei Fu, Yifan Duan, Yao Li et al.