HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
HunyuanOCR-1.5 enhances OCR performance using DFlash acceleration and Agentic Data Flow, achieving 6.37x inference speedup.
Gengluo Li, Xingyu Wan, Shangpin Peng et al.
HunyuanOCR-1.5 enhances OCR performance using DFlash acceleration and Agentic Data Flow, achieving 6.37x inference speedup.
Gengluo Li, Xingyu Wan, Shangpin Peng et al.
Unsupervised detection of underground tunnels using depth-restricted reconstruction scoring, achieving AUC of 0.994.
Muhammad Junaid, Shoab A. Khan, Nisar Ahmed
FocusGS uses targeted structure completion to reduce Gaussian count by 74%, improving sparse-view 3D reconstruction efficiency and quality.
Guoqing Wang, Pin Tang, Xiangxuan Ren et al.
PixelPilot decouples 2D planning from 3D lifting, enabling scalable vision-language-driven autonomous driving with state-of-the-art accuracy.
Pin Tang, Guoqing Wang, Xiangxuan Ren et al.
G2VD integrates counterfactual intervention and causal disentanglement, achieving over 90% accuracy in cross-domain AI video forgery detection.
Meng Du, Hongchang Chen, Ran Li et al.
Flash-BoN combines timestep truncation, layer skipping, and activation proxies to optimize inference, achieving +8% AUC under fixed budgets.
Ruchit Rawal, Reza Shirkavand, Sayak Paul et al.
This study investigates cross-task transferability in unified multimodal models, showing fully shared transformer architectures (like Lumina-DiMOO) achieve up to 9% accuracy gains in counting tasks.
Jiwon Kang, Heeji Yoon, Jaewoo Jung et al.
Proposed RZDG dataset and a tracker-based geo-localization pipeline significantly improve roadwork zone detection and global positioning accuracy.
Zhiran Yan, Yutong Xin, S Shyam Shenoi et al.
AdaptiveSplat uses texture-aware pruning and adaptive Gaussian prediction to improve sparse 3D scene reconstruction without fine-tuning.
Badrinath Singhal, Srihari K G, Sreehari Iyer et al.
MRAC method significantly improves depth estimation robustness without adding parameters.
Sohag Roy, Rajesh Misra, Swami Shastravidyananda et al.
GLLS introduces a training-free dual-stream framework combining global logical verification and active local search, achieving 84% accuracy on MMAD-QA.
Runzhi Deng, Yundi Hu, Yiming Zhong et al.
Proposes a Gaussian Splatting-based scene optimization with normal-guided depth propagation, outperforming SOTA in sparse-view 3D surface reconstruction.
Liang Han, Bangcai Wei, Junsheng Zhou et al.
Decomposable Probe separates prompt, latent, and score responses, detecting rectified flow through latent selectivity across 23 models.
Patrick Mu Haojie
BVS combines Bayesian optimization and multimodal large language models for fine-grained perception in ultra-high-resolution images, significantly improving accuracy and efficiency.
Geng Li, Yuxin Peng
SafeGuard employs multi-agent perception and reasoning, achieving 18.7% accuracy improvement in social-risk video detection on SafeVid dataset.
Wenlin Wu, Sheng Zhou, Peipei Song et al.
CONFLUX combines 3D latent flow generation with RL control, reaching tri-planar FID 32.3 versus MAISI’s 74.6.
Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge et al.
Proposed cross-device collaborative test-time adaptation using zeroth-order optimization and model merging, reducing error rate by 18.3% on CIFAR10-C.
Yu Mitsuzumi, Akisato Kimura, Yasuhiro Fujiwara et al.
EMA uses DiT Massive Activations: DG improves details and MREP improves dense features; MA disruption leaves BLIP/CLIP win rates at 0.462/0.512.
Chaofan Gan, Zicheng Zhao, Yuanpeng Tu et al.
R3D constructs 3D scenes from egocentric RGB-D video and uses tool calls to achieve 73.5% accuracy in quantitative 3D reasoning tasks.
Maxwell Horton, Wei Lu, Quan Tran et al.
UniCaMo constructs a 3D-grounded noise space for joint control of object and camera motion, improving video coherence and quality.
Long Vu, Tan Ngo, Animesh Karnewar et al.