T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation
T-FunS3D employs task-driven hierarchical open-vocabulary 3D segmentation, boosting speed and resource efficiency.
Jingkun Feng, Reza Sabzevari
T-FunS3D employs task-driven hierarchical open-vocabulary 3D segmentation, boosting speed and resource efficiency.
Jingkun Feng, Reza Sabzevari
Proposes a geometry-aware dataset condensation method to enhance fidelity and distribution coverage in diffusion model training.
Xiao Cui, Yulei Qin, Mo Zhu et al.
Parallel Jacobi Decoding achieves 4.8x to 6.4x speedup in autoregressive image generation.
Boya Liao, Ying Li, Siyong Jian et al.
WorldBench constructs a multi-domain visual concept taxonomy, evaluating multimodal models; top accuracy is only 64%, revealing significant gaps.
Yida Yin, Harish Krishnakumar, Chung Peng Lee et al.
BloomBench evaluates VLM cognition across six Bloom levels using 7,747 bilingual items and reports 98.45% audited quality.
Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor et al.
ZipSplat uses scene tokens to decouple Gaussian placement from pixels, achieving high-quality scene reconstruction with 6× fewer Gaussians, and adjustable quality-efficiency trade-off.
Alexander Veicht, Sunghwan Hong, Dániel Baráth et al.
M³Eval, grounded in cognitive psychology, assesses multi-modal video memory across four dimensions with controlled tasks and quantitative metrics.
Jie Huang, Ruixun Liu, Sirui Sun et al.
Empirical analysis shows increasing data scale improves generalization; model complexity has unstable effects; removing color info degrades performance.
Yidi Zhouluo
UniCanvas employs diffusion models on a shared pixel canvas for joint text-image generation, significantly enhancing multi-step multimodal reasoning.
Zeyuan Yang, Hao-Wei Chen, Xueyang Yu et al.
VidMsg benchmark evaluates implicit message understanding in short videos, enhancing retrieval with VidVec-Msg.
Issar Tzachor, Michael Green, Rami Ben-Ari
EvoMemNav constructs a self-evolving visual-semantic memory graph with budgeted coarse-to-fine policy, achieving 59.6% SR and 38.9 SPL on GOAT-Bench.
Zuhao Ge, Xiaosong Jia, Chao Wu et al.
Proposes Steady-Forcing, combining V-Sink and EMA-Sink, to improve long-horizon natural video diffusion stability and fluid motion.
Matiur Rahman Minar, Seunghun Oh, GangHyeon Jeong et al.
MemoGen enhances image generation via experience memory, surpassing Nano Banana Pro and GPT-Image-1.
Wenshuo Chen, Kuimou Yu, Bowen Tian et al.
Any2Poster enables cross-modal, domain-general poster generation with 87.25% accuracy, integrating unified parsing, adaptive layout, and VLM-guided visual repair.
Amogh Vinaykumar, Aiden Li, Suozhi Huang et al.
Proposes SEIG, a staged framework leveraging pretrained vision-language models (VLMs) to reconstruct editable 3D scenes from a single image, achieving high fidelity in geometry, materials, and lighting.
Guangzhao He, Rundong Luo, Wei-Chiu Ma et al.
AdaCodec employs predictive visual coding, transmitting full reference frames only when prediction is costly, reducing visual tokens by 84.7% and boosting long-video understanding efficiency.
Haowen Hou, Zhen Huang, Zheming Liang et al.
This paper introduces VLM as a teacher for video reasoning via test-time online optimization, achieving a 16.7-point performance boost, surpassing traditional methods.
Junhao Cheng, Liang Hou, Tianxiong Zhong et al.
Li et al. propose DivIn, using Langevin dynamics to optimize initial noise, significantly boosting diversity in diffusion models, outperforming existing methods.
Xiang Li, Dianbo Liu, Kenji Kawaguchi
Proposes MSLoc, a boundary-sensitive, multi-modal framework for detecting manipulated segments in untrimmed long videos, achieving 78.5% mAP.
Yue Feng, Jingjing Li, Qijia Lu et al.
WebSpline employs structure-guided splines for real-time, high-fidelity 3D Gaussian reconstruction from monocular videos, outperforming prior methods in quality and speed.
Jongmin Park, Jeonghwan Yun, Minh-Quan Viet Bui et al.