Spatial Calibration of Diffuse LiDARs
Proposes a passive background subtraction method to estimate each diffuse LiDAR pixel’s footprint and spatial sensitivity in RGB images.
Nikhil Behari, Ramesh Raskar
Proposes a passive background subtraction method to estimate each diffuse LiDAR pixel’s footprint and spatial sensitivity in RGB images.
Nikhil Behari, Ramesh Raskar
Proposes WanderDream, a world model-based emulative framework generating 15.8K videos and 158K QA pairs for scene reasoning without active exploration.
Ruiping Liu, Yufan Chen, Yuheng Zhang et al.
LATO uses topology-preserving sparse voxel VAE to efficiently generate structured 3D meshes.
Tianhao Zhao, Youjia Zhang, Hang Long et al.
Proposes auto aggregation and per-pixel rescaling to enhance training-free diffusion segmentors, leveraging stronger generative models for improved semantic segmentation.
Benyuan Meng, Qianqian Xu, Zitai Wang et al.
Proposes Skeleton-to-Image encoding (S2I) enabling self-supervised skeleton representation learning via vision-pretrained models, improving cross-format generalization.
Siyuan Yang, Jun Liu, Hao Cheng et al.
InnoAds-Composer employs single-stage tri-conditional diffusion with importance routing, achieving 54.39 FID and 0.857 sentence accuracy, greatly improving e-commerce poster generation.
Yuxin Qin, Ke Cao, Haowei Liu et al.
Built on Mip-NeRF with adaptive weighted MSE, this method reconstructs 3D gas plumes from LWIR hyperspectral images, reducing training data by 50%.
Scout Jarman, Zigfried Hampel-Arias, Adra Carr et al.
Proposes VideoHV-Agent, a hypothesis-verification multi-agent framework, achieving state-of-the-art accuracy in long video question answering.
Zheng Wang, Haoran Chen, Haoxuan Qin et al.
Pointer-CAD addresses B-rep entity selection issues via pointer-based commands, reducing segmentation error significantly.
Dacheng Qi, Chenyu Wang, Jingwei Xu et al.
VAS quantifies visual attention; AVAR improves multimodal reasoning by 7% without retraining.
Ruilin Luo, Chufan Shi, Yizhen Zhang et al.
DriveMVS leverages sparse LiDAR as geometric priors, fuses multi-source cues, and employs spatio-temporal decoding to achieve high-accuracy, consistent depth estimation for autonomous driving.
Qihao Sun, Jiarun Liu, Ziqian Ni et al.
InfinityStory combines world-consistent generation and multi-subject transitions, reaching VBench scores of 88.94 and 82.11.
Mohamed Elmoghany, Liangbing Zhao, Xiaoqian Shen et al.
Proposes PhyPrompt, an RL-based two-stage prompt refinement framework, achieving 40.8% joint success with 7B parameters.
Shang Wu, Chenwei Xu, Zhuofan Xia et al.
MoD-DPO reduces cross-modal hallucinations via modality decoupling, boosting accuracy to 88.19%.
Ashutosh Chaubey, Jiacheng Pang, Mohammad Soleymani
Proposes RL3DEdit, leveraging VGGT rewards in reinforcement learning for multi-view consistent 3D scene editing, achieving high quality and efficiency.
Jiyuan Wang, Chunyu Lin, Lei Sun et al.
Proposes BrandFusion, a multi-agent framework for seamless brand embedding in text-to-video, improving recognizability and naturalness.
Zihao Zhu, Ruotong Wang, Siwei Lyu et al.
UniTalking framework uses Multi-Modal Transformer Blocks to generate high-fidelity speech and synchronized video, surpassing existing open-source methods.
Hebeizi Li, Zihao Liang, Benyuan Sun et al.
FlexiMMT achieves multi-object multi-motion transfer using Motion Decoupled Mask Attention, improving motion fidelity by 23%.
Yuze Li, Dong Gong, Xiao Cao et al.
SOLACE enhances text-to-image generation quality using intrinsic self-confidence signals, reducing external supervision needs.
Seungwook Kim, Minsu Cho
Fourier Angle Alignment (FAA) improves rotated object detection accuracy in remote sensing, achieving 78.72% mAP on DOTA-v1.0 with frequency-based orientation estimation.
Changyu Gu, Linwei Chen, Lin Gu et al.