GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
GroundingSuite leverages multi-modal models for automatic pixel grounding annotation, achieving 68.9 cIoU on gRefCOCO with 955万样本。
Rui Hu, Lianghui Zhu, Yuxuan Zhang et al.
GroundingSuite leverages multi-modal models for automatic pixel grounding annotation, achieving 68.9 cIoU on gRefCOCO with 955万样本。
Rui Hu, Lianghui Zhu, Yuxuan Zhang et al.
CameraCtrl II employs camera-conditioned video diffusion with sequential generation to enable large-scale dynamic scene exploration, expanding viewpoint range and scene continuity.
Hao He, Ceyuan Yang, Shanchuan Lin et al.
Long Context Tuning (LCT) extends pre-trained video diffusion models' context window, enabling scene-level multi-shot generation with high consistency.
Yuwei Guo, Ceyuan Yang, Ziyan Yang et al.
CINEMA uses MLLM for coherent multi-subject video generation, enhancing video coherence and subject consistency.
Yufan Deng, Xun Guo, Yizhi Wang et al.
VicaSplat achieves 3D Gaussian reconstruction and camera estimation from unposed frames in one run, outperforming baselines.
Zhiqi Li, Chengrui Dong, Yiming Chen et al.
IMPACT framework uses Vision-Language Models for contact-rich motion planning, achieving a 73.75% success rate.
Yiyang Ling, Karan Owalekar, Oluwatobiloba Adesanya et al.
Proposes MoC framework with Boundary Clarity and Chunk Stickiness metrics, significantly improving text chunking and retrieval-augmented generation performance.
Jihao Zhao, Zhiyuan Ji, Zhaoxin Fan et al.
Proposes Plan-and-Act framework with synthetic data augmentation, achieving 57.58% success in long-horizon web tasks.
Lutfi Eren Erdogan, Nicholas Lee, Sehoon Kim et al.
Search-R1 employs reinforcement learning enabling LLMs to generate multi-turn search queries, boosting QA performance by 41%.
Bowen Jin, Hansi Zeng, Zhenrui Yue et al.
Týr-Pruner employs end-to-end global sparsity distribution optimization via supernet construction and evolutionary search, retaining 97% performance with 50% parameter removal.
Guanchen Li, Yixing Xu, Zeping Li et al.
UniCombine employs a Diffusion Transformer with Conditional MMDiT Attention and LoRA modules, enabling flexible multi-conditional image generation with SOTA performance.
Haoxuan Wang, Jinlong Peng, Qingdong He et al.
Reangle-A-Video generates multi-view videos via video-to-video translation, outperforming existing methods.
Hyeonho Jeong, Suhyeon Lee, Jong Chul Ye
LocAgent employs graph-based representation and multi-hop reasoning, achieving 92.7% accuracy in code localization with 86% cost reduction using fine-tuned Qwen-2.5-Coder-Instruct-32B.
Zhaoling Chen, Xiangru Tang, Gangda Deng et al.
SeqMultiGrasp system uses diffusion model for multi-object grasping with dexterous hand, achieving 65.8% success in simulation.
Sicheng He, Zeyu Shangguan, Kuanning Wang et al.
Proposes CycleVAE and causal Transformer for unsupervised cross-embodiment robotic manipulation, generating smooth trajectories with 85% success rate.
Apan Dastider, Hao Fang, Mingjie Lin
Proposes GoAI to build AI research knowledge graphs, enabling personalized learning paths and innovation support.
Xian Gao, Zongyun Zhang, Ting Liu et al.
This paper introduces RexSeek, a model combining multimodal large language models and object detection, achieving superior multi-person referring understanding on the HumanRef dataset.
Qing Jiang, Lin Wu, Zhaoyang Zeng et al.
OpenRAG end-to-end tuning improves retrieval relevance by 4.0%, surpassing SOTA retrievers by 2.1%, using contrastive learning and approximate labels.
Jiawei Zhou, Lei Chen
ArticulatedGS employs self-supervised learning with multi-view RGB images and 3D Gaussian splatting for part-level articulated object reconstruction and motion estimation.
Junfu Guo, Yu Xin, Gaoyi Liu et al.
Proposes a knowledge-centric RAG framework emphasizing knowledge lifecycle management for improved NLP tasks.
Mingyue Cheng, Yucong Luo, Jie Ouyang et al.