3D Implicit Transporter for Temporally Consistent Keypoint Discovery
Proposes 3D Implicit Transporter for temporally consistent keypoint detection, improving non-rigid object understanding.
Chengliang Zhong, Yuhang Zheng, Yupeng Zheng et al.
Proposes 3D Implicit Transporter for temporally consistent keypoint detection, improving non-rigid object understanding.
Chengliang Zhong, Yuhang Zheng, Yupeng Zheng et al.
SyncDreamer employs a multiview-synchronized diffusion framework to generate consistent multi-view images from a single view, achieving superior 3D reconstruction.
Yuan Liu, Cheng Lin, Zijiao Zeng et al.
CIEM method evaluates VLM hallucination by generating contrastive Q&A pairs, significantly improving model accuracy.
Hongyu Hu, Jiyuan Zhang, Minyi Zhao et al.
PointLLM integrates point cloud encoders with LLMs, achieving 53%+ zero-shot classification accuracy and surpassing human annotations in detailed object captioning.
Runsen Xu, Xiaolong Wang, Tai Wang et al.
Proposed InterDiff combines diffusion models with physics priors for long-term 3D human-object interaction prediction.
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang et al.
EMDB leverages electromagnetic sensors and deep neural models to create a high-precision 3D human pose and shape dataset in the wild, with global trajectories.
Manuel Kaufmann, Jie Song, Chen Guo et al.
TouchStone evaluates LVLMs using strong LLMs like GPT-4, covering five abilities and 27 subtasks.
Shuai Bai, Shusheng Yang, Jinze Bai et al.
Residual Denoising Diffusion Model (RDDM) introduces dual diffusion processes for unified image generation and restoration, leveraging residuals and noise with independent scheduling.
Jiawei Liu, Qiang Wang, Huijie Fan et al.
EigenPlaces employs geolocation-guided clustering and SVD to create multi-view categories, boosting viewpoint robustness with 8% higher Recall@1 and 50% smaller descriptors.
Gabriele Berton, Gabriele Trivigno, Barbara Caputo et al.
Proposes Point Prompt Training (PPT) with Prompt-driven Normalization and Language-guided Categorical Alignment to mitigate negative transfer in multi-dataset 3D point cloud learning, achieving SOTA results.
Xiaoyang Wu, Zhuotao Tian, Xin Wen et al.
Dynamic Gaussian model enables real-time 6-DOF scene tracking and novel-view synthesis without correspondence inputs, achieving 28.7 PSNR and 850 FPS.
Jonathon Luiten, Georgios Kopanas, Bastian Leibe et al.
EgoSchema introduces a long-video QA benchmark using temporal certificate length; models achieve <33% accuracy, humans 76%.
Karttikeya Mangalam, Raiymbek Akshulakov, Jitendra Malik
Introduces MeViS dataset for motion-guided video segmentation; current RVOS methods perform poorly (~35.5% J&F).
Henghui Ding, Chang Liu, Shuting He et al.
A2Nav leverages foundation models for zero-shot robot navigation, surpassing supervised methods with 22.6% success rate on R2R-Habitat.
Peihao Chen, Xinyu Sun, Hongyan Zhi et al.
DREAMWALKER employs explicit world models with environment graphs and scene synthesizers, combined with Monte Carlo Tree Search, to enable strategic mental planning in continuous vision-language navigation, achieving over 78% success rate.
Hanqing Wang, Wei Liang, Luc Van Gool et al.
ModelScopeT2V uses diffusion with spatio-temporal blocks, 1.7B parameters, achieving superior text-to-video synthesis with high temporal coherence.
Jiuniu Wang, Hangjie Yuan, Dayou Chen et al.
Introduced VPG-C module, significantly enhancing zero-shot performance on DEMON benchmark.
Juncheng Li, Kaihang Pan, Zhiqi Ge et al.
FourLLIE enhances low-light images using Fourier frequency information, surpassing SOTA on four datasets with only 0.31% parameters.
Chenxi Wang, Hongjun Wu, Zhi Jin
MVFlow leverages compressed video motion vectors to improve optical flow estimation, reducing AEPE by 1.09 and saving 52% computation time.
Shili Zhou, Xuhao Jiang, Weimin Tan et al.
OpenFlamingo is an open-source framework for training autoregressive vision-language models with 3B to 9B parameters, achieving 80-89% of Flamingo's performance.
Anas Awadalla, Irena Gao, Josh Gardner et al.