Can 4D Foundation Models Remember?
PersistBench evaluates 4D models' visual memory using 360° videos, revealing short-term consistency issues.
Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma
PersistBench evaluates 4D models' visual memory using 360° videos, revealing short-term consistency issues.
Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma
SplashSplat reconstructs splashing liquids from multi-view videos, offering lower training costs and more plausible physical motion.
Peiyu Liu, Dingxi Zhang, Federico Tombari et al.
FAMOS predicts movable-part segmentation and joint parameters from sparse point clouds, improving segmentation performance by 64.5%.
Kevin Qu, Tao Sun, Massimiliano Viola et al.
Paint-Anything achieves any-color control for image generation/editing via shared hex prompts, improving ACBench-T2I by 85.3%.
Ji Xie, Dewei Zhou, Xinyu Huang et al.
ERCPMP-Gx dataset integrates endoscopic images, histopathology, and genomics for colorectal polyposis AI research.
Zahra Ghaffari, Massih Bahar, Mojgan Forootan et al.
FlowSGS enhances inverse imaging accuracy using Split Gibbs Sampling and Stochastic Interpolants.
Tianao Li, Xinhui Qian, Emma Alexander
Using prediction fragmentation to control test-time adaptation, significantly reducing harmful accepted area.
Lili Wang, Jing Li, Xiaowen Sun et al.
VideoResearcher enhances long-video understanding via automated tool design, achieving 74.5% accuracy.
Dingqiang Ye, Dongdi Zhao, Kaishen Wang et al.
ReconPlusGen injects reconstruction priors into multi-view 3D generation via noise inversion, significantly enhancing reconstruction accuracy.
Jiarui Liu, Heng Li, Weiyu Li et al.
CamPilot uses a multi-agent framework to enhance cinematographic control and quality in movie generation.
Yang Wu, Stefano Petrangeli, Ishita Dasgupta et al.
Introduced Programmable World Model, achieving 94% Count Accuracy and 98% State Accuracy by decoupling state evolution from visual generation.
Zheng-Hui Huang, Guixu Lin, Jiacheng Lin et al.
Field Converter reduces player pose estimation error in soccer broadcasts from 49cm to 10cm using geometry initialization and temporal residual refinement.
Simon Khan, Laurent Gajny, Jennyfer Lecompte et al.
Proposed RBQE framework uses cross-model agreement to estimate polyp segmentation reliability, achieving ROC-AUC of 0.960.
Siddharth Gupta, Jitin Singla
PACE reduces P95 PTFR from 0.53s to 0.29s via perceived-latency-aware dialogue routing.
Lin Huang, Yujuan Tan, Weisheng Li et al.
Survey of inference-efficiency mechanisms in VideoLLMs, analyzing frame sampling, modality encoding, and token compression techniques.
Killian Steunou, Yannis Tevissen, Mounîm A. El Yacoubi
Motar achieves 15.4 FPS high-fidelity streaming talking head generation by fusing conditions in a low-dimensional motion space.
Yanru An, Ruiyan Wang, Wenwu Wei et al.
ScopeMamba-YOLO enhances small object detection in remote sensing imagery by widening perceptual scope, achieving a 10.8 pp mAP50 improvement.
Junjie Fan, Yijun Mai, Linduo Wei et al.
An SDT extension combining geometric vision, class-wise fusion, and affective priors reaches 75.93% and 74.11% weighted F1 on MELD and IEMOCAP.
Oriol Marín, Roger Marí, Gloria Haro et al.
MotionBlind reveals Video-LLMs' inability to accurately understand motion in videos; only Gemini3.1 Pro passes.
Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche et al.
VANTAGE-Bench evaluates the Infrastructure AI gap in VLMs, revealing deficiencies in event verification and temporal localization tasks.
Zaid Pervaiz Bhat, Nimra Nayyar, Arihant Jain et al.