GroundingSuite: Measuring Complex Multi-Granular Pixel Grounding
GroundingSuite leverages multi-modal models for automatic pixel grounding annotation, achieving 68.9 cIoU on gRefCOCO with 955万样本。
Rui Hu, Lianghui Zhu, Yuxuan Zhang et al.
GroundingSuite leverages multi-modal models for automatic pixel grounding annotation, achieving 68.9 cIoU on gRefCOCO with 955万样本。
Rui Hu, Lianghui Zhu, Yuxuan Zhang et al.
CameraCtrl II employs camera-conditioned video diffusion with sequential generation to enable large-scale dynamic scene exploration, expanding viewpoint range and scene continuity.
Hao He, Ceyuan Yang, Shanchuan Lin et al.
Long Context Tuning (LCT) extends pre-trained video diffusion models' context window, enabling scene-level multi-shot generation with high consistency.
Yuwei Guo, Ceyuan Yang, Ziyan Yang et al.
VicaSplat achieves 3D Gaussian reconstruction and camera estimation from unposed frames in one run, outperforming baselines.
Zhiqi Li, Chengrui Dong, Yiming Chen et al.
UniCombine employs a Diffusion Transformer with Conditional MMDiT Attention and LoRA modules, enabling flexible multi-conditional image generation with SOTA performance.
Haoxuan Wang, Jinlong Peng, Qingdong He et al.
This paper introduces RexSeek, a model combining multimodal large language models and object detection, achieving superior multi-person referring understanding on the HumanRef dataset.
Qing Jiang, Lin Wu, Zhaoyang Zeng et al.
ArticulatedGS employs self-supervised learning with multi-view RGB images and 3D Gaussian splatting for part-level articulated object reconstruction and motion estimation.
Junfu Guo, Yu Xin, Gaoyi Liu et al.
SANDRO combines IRLS and splitting strategy, achieving 20% higher success in point cloud registration under high outlier rates, outperforming state-of-the-art methods.
Michael Adlerstein, João Carlos Virgolino Soares, Angelo Bratta et al.
Proposes MovieAgent, a multi-agent hierarchical CoT framework for automated movie generation, achieving superior narrative coherence and character consistency.
Weijia Wu, Zeyu Zhu, Mike Zheng Shou
COMODO enhances IMU recognition efficiency through cross-modal distillation, matching or surpassing supervised models.
Baiyu Chen, Wilson Wongso, Zechen Li et al.
Seg-Zero uses cognitive reinforcement to generate explicit reasoning chains, achieving 57.5% zero-shot performance on ReasonSeg with a decoupled architecture.
Yuqi Liu, Bohao Peng, Zhisheng Zhong et al.
Proposes Bayesian Fields, a task-driven open-set semantic mapping method using probabilistic 3D Gaussian representations and Bayesian updating for multi-view fusion, outperforming traditional averaging.
Dominic Maggio, Luca Carlone
VideoPainter uses a dual-stream architecture for any-length video inpainting, enhancing semantic consistency.
Yuxuan Bian, Zhaoyang Zhang, Xuan Ju et al.
TrajectorCrafter employs dual-stream diffusion to redirect monocular video camera trajectories, integrating point cloud rendering and source video for high-fidelity results.
Mark YU, Wenbo Hu, Jinbo Xing et al.
LION-FS enhances real-time video assistant accuracy via Fast & Slow path strategy.
Wei Li, Bing Hu, Rui Shao et al.
CMMLoc employs Cauchy Mixture Models within a Transformer framework to model partial relevance, achieving state-of-the-art localization accuracy on KITTI360Pose with 83% within 5m.
Yanlong Xu, Haoxuan Qu, Jun Liu et al.
Object-aware video matting with cross-frame guidance achieves SOTA, reducing reliance on manual trimaps, with MAD 4.23 and MSE 0.31 on RVM.
Huayu Zhang, Dongyue Wu, Yuanjie Shao et al.
SEE-Net uses event data for adaptive brightness adjustment, enhancing image quality across broad light ranges.
Yunfan Lu, Xiaogang Xu, Hao Lu et al.
EndoPBR uses physically-based differentiable rendering to estimate materials and lighting, enabling photorealistic novel view synthesis in surgical scenes.
John J. Han, Jie Ying Wu
InterMimic employs teacher-student RL framework to learn diverse, physically plausible human-object interactions from imperfect MoCap data, achieving zero-shot generalization.
Sirui Xu, Hung Yu Ling, Yu-Xiong Wang et al.