Think in Sets for Streaming Video Token Compression
NovaCov uses set-wise selection with a recency-weighted reference bank, retaining 99.6% accuracy and reducing latency by 46% in streaming video token compression.
Moxu Duan, Jingwen Fu, Yuwang Wang
NovaCov uses set-wise selection with a recency-weighted reference bank, retaining 99.6% accuracy and reducing latency by 46% in streaming video token compression.
Moxu Duan, Jingwen Fu, Yuwang Wang
AIMold automates complex mold design using the MoldCAD dataset with 4,934 CAD models.
Pengyun Qiu, Shuo Wang, Zeyuan Chen et al.
Strict institution-held-out testing found Beta-Binomial EB and target-only logistic regression best, with Brier scores of 0.0853 and 0.0855.
Pengyang Yu, Yiou Wang, Zhongping Dong et al.
Proposed ID-Constraint enhances monocular depth estimation robustness against camera roll variations, improving accuracy by 0.02-0.03 in AbsRel across datasets.
Kaihua Tang, Ziqing Xia, Xiaoxu Zheng et al.
P2Voxel achieves efficient 3D mesh reconstruction via pyramid pivot voxelization, reducing storage significantly.
Zhenhong Sun, Haozhe Liu, Yifu Wang et al.
Video generators struggle with irreversible processes; proposed progress and stasis rate protocol for evaluation.
Jian Xu, Yanning Wu, Delu Zeng et al.
Introduces FD-loss post-training using Fréchet distance in feature space to improve autoregressive image generation, reducing FID by 41.4%.
Jinhua Zhang, Yisong Lin, Wei Long et al.
AOCT-MSQR enables training-free video attribution by retrieval, reaching 84.6% Rank-1/78.3% mAP on GenVidBench at 100-shot.
Renxi Cheng, Chaolei Han, Jie Gui et al.
ShadowDancer learns unified dynamic representations from video-shadow pairs, enabling precise, label-free action transfer across diverse scenarios.
Jin Cao, Zian Meng, Kaipeng Zhang
TARS learns camera motion at high-noise diffusion steps, reaching 83.18% R-Prec, 72.76% T-Prec, and 10.34° V-MPGE.
Jiwen Liu, Shujuan Li, Xiaohan Li et al.
Large-scale evaluation of 194 VLMs shows model size correlates with overall accuracy (ρ=0.68), but not with robustness to multi-attribute bias; data quality is more critical.
Ioannis Sarridis, Ioannis Kompatsiaris, Symeon Papadopoulos
CineWeaver uses inference-time positional encoding, attention, routing, and memory to generate long, reference-controlled multi-shot videos without retraining; supplied text reports no numeric metrics.
Yuyang Huang, Yabo Chen, Wenrui Dai et al.
Introduces cross-branch semantic steering to test if understanding and generation share a unified semantic space; transfer from understanding to generation is effective, reverse is limited.
Yu Wang, Sharon Li
MoSAIC achieves part-local motion style transfer using aligned intervention supervision, reducing errors and enhancing response.
Nazanin Amini, Kevin Desai
MODUS is a decoder-only multimodal model supporting arbitrary input-output combinations, enabling multi-task and cross-modal applications.
Mingqiao Ye, Zhaochong An, Zhitong Gao et al.
Argus-Unified leverages pretrained VLM with hybrid tokens, achieving SOTA multimodal understanding and competitive generation at $2000 and 15.6M data.
Weiming Zhuang, Jiabo Huang, Jingtao Li et al.
I2VShield uses generative adversarial attacks to protect privacy, reducing computational costs and enhancing DiT model defenses.
Yimao Guo, Zuomin Qu, Wei Lu
CLBench-V evaluates multimodal context learning; best model scores only 0.2847.
Lai Wei, Chengqi Li, Jiapeng Li et al.
PerceptionBench evaluates atomic visual perception in MLLMs; highest accuracy among 16 models is only 59.7%.
Zichao Lin, Yifeng Xie, Bowen Qu et al.
Mage-VL employs Codec-native sparse token encoding, reducing 75% tokens and achieving 3.5× inference speedup for real-time multimodal streaming.
Senqiao Yang, Kaichen Zhang, Zhaoyang Jia et al.