V-Nutri: Dish-Level Nutrition Estimation from Egocentric Cooking Videos
V-Nutri fuses final-dish and cooking keyframes; HD-EPIC experiments show process cues can improve nutrition estimation.
Chengkun Yue, Chuanzhi Xu, Jiangpeng He
V-Nutri fuses final-dish and cooking keyframes; HD-EPIC experiments show process cues can improve nutrition estimation.
Chengkun Yue, Chuanzhi Xu, Jiangpeng He
PLOVIS leverages pre-trained open-vocabulary image segmentation (DeCLIP) to generate pseudo labels directly from 3D point clouds, enabling effective training under scarce data conditions.
Takahiko Furuya
Introduces Energy-oriented Diffusion Bridge (E-Bridge), reducing sampling steps and improving image restoration quality via low-energy geodesic trajectories.
Jinhui Hou, Zhiyu Zhu, Junhui Hou
DeepShapeMatchingKit accelerates functional map solving by 33x using vectorized reformulation, addressing 3D shape matching bottlenecks.
Yizheng Xie, Lennart Bastian, Congyue Deng et al.
SMFormer integrates foundation models and data augmentation to enhance self-supervised stereo matching.
Yun Wang, Zhengjie Yang, Jiahao Zheng et al.
GREATEN integrates surface normals with efficient sparse attention to enhance cross-domain stereo matching, reducing errors by over 30%.
Jiahao Li, Xinhong Chen, Zhengmin Jiang et al.
VLA-World integrates predictive imagination with reflective reasoning, reducing collision rate to 0.10% and outperforming SOTA in autonomous driving benchmarks.
Guoqing Wang, Pin Tang, Xiangxuan Ren et al.
OV-Stitcher employs a global attention mechanism without training, boosting mIoU from 48.7 to 50.7 by stitching sub-image features for open-vocabulary segmentation.
Seungjae Moon, Seunghyun Oh, Youngmin Ro
Proposes a training-free open-vocabulary semantic segmentation method using analytic solutions based on distribution discrepancy analysis, achieving SOTA on 8 datasets.
Jiahao Li, Yang Lu, Yachao Zhang et al.
Introduced PAMELA dataset and personalized reward model, improving individual preference prediction in T2I generation.
Anne-Sofie Maerten, Juliane Verwiebe, Shyamgopal Karthik et al.
PhyEdit combines explicit geometric simulation with Diffusion Transformer to achieve high-precision 3D-aware image manipulation, outperforming existing methods with significant improvements in geometric accuracy and physical consistency.
Ruihang Xu, Dewei Zhou, Xiaolong Shen et al.
ModuSeg decouples object discovery and semantic retrieval, enabling training-free weakly supervised segmentation with state-of-the-art performance.
Qingze He, Fagui Liu, Dengke Zhang et al.
Proposes DOC-GS framework combining depth-guided Dropout and Dark Channel Prior for reliable sparse-view 3D Gaussian reconstruction.
Hantang Li, Qiang Zhu, Xiandong Meng et al.
This paper systematically reviews video generation evolution from GANs to diffusion and autoregressive models, analyzing core principles and innovations.
Teng Hu, Jiangning Zhang, Hongrui Huang et al.
SymphoMotion jointly controls camera trajectories and object dynamics, achieving coherent video generation with improved fidelity and spatial consistency.
Guiyu Zhang, Yabo Chen, Xunzhi Xiang et al.
VERTIGO optimizes visual preference, reducing off-screen rate to 0% and enhancing shot quality.
Mengtian Li, Yuwei Lu, Feifei Li et al.
RebusBench evaluates LVLMs' cognitive visual reasoning, performance below 10% exact match.
Seyed Amir Kasaei, Arash Marioriyad, Mahbod Khaleti et al.
SteerFlow introduces fixed-point and trajectory interpolation techniques to improve faithful inversion-based image editing, outperforming existing methods in source preservation.
Thinh Dao, Zhen Wang, Kien T. Pham et al.
ReFlow employs self-correction flow matching for monocular 4D scene reconstruction, surpassing existing methods with no external motion guidance, achieving PSNR of 27.65dB.
Yanzhe Liang, Ruijie Zhu, Hanzhi Chang et al.
Proposes a granular ball-based square superpixel method enabling end-to-end deep learning integration, improving efficiency and structure.
Shuyin Xia, Meng Yang, Dawei Dai et al.