GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
GS-Voxel: Fitting-free structured latents for large-scale 3DGS generation, supporting millions of voxels.
Ming Qian, Zijian Wang, Minchao Sun et al.
GS-Voxel: Fitting-free structured latents for large-scale 3DGS generation, supporting millions of voxels.
Ming Qian, Zijian Wang, Minchao Sun et al.
Proposes 'Entry Point Gates' to control cross-task usability; finds that binding location in understanding pathways enables concept transfer with minimal cost.
Zongyang Qiu, Yihan Wu, Kaixuan Fan et al.
PixRestore is a VAE-free pixel-space diffusion transformer with 50M parameters, enabling single-step high-fidelity image restoration.
Lingchen Sun, Rongyuan Wu, Xiangtao Kong et al.
GenRouter optimizes image generation workflows, reducing computational costs by 95%.
Harold Haodong Chen, Zhiyu Hou, Wen-Jie Shu et al.
Prototypical Network-based ID PAD achieves ~9% EER with only 4 samples, enabling cross-country domain generalization.
Mario Nieto-Hidalgo, Juan M. Espin, Juan E. Tapia
Ultra framework achieves 73.0% mIoU on ACDC using CTDN and CMIL for unsupervised cross-task optimization under adverse weather.
Shiqin Wang, Zhiqian Li, Haoyuan Du et al.
Using class-conditional GAN and DDPM to augment satellite images, boosting classification accuracy to 88%.
Marta Sumyk, Oleksandr Kosovan, Iryna Voitsitska
This study uses CRP and CRAFT to reveal concept-level biases and confusion trends in MS-COCO multi-label classification, aiding interpretability.
Haadia Amjad, Ronald Tetzlaff
ConceptFormer enhances visual document retrieval by learning adaptive latent concepts, achieving a 16.7% NDCG@10 improvement.
Chunyi Peng, Zhipeng Xu, Yukun Yan et al.
Proposes synchrony-aware sparse attention for audio-visual generation, achieving near 2× speedup while maintaining quality.
Shengchuan Gao, Teng Hu, Bohao Feng et al.
StructFlow introduces spatially structured noise in flow models, enhancing local control and structure preservation in image generation.
Arman Zarei, Mahdi M. Kalayeh
Proposed a unified backbone-expert framework achieving 67.28% accuracy on RML2016.10b.
Zhixiang Deng, Houbiao Li, Zongyong Cui
MOSS-VL is an open-source vision-language model supporting real-time interaction via gated cross-attention, with 11.3B parameters, excelling in proactive response tasks.
Pengyu Wang, Chenkun Tan, Shaojun Zhou et al.
Multiphase-Diff uses diffusion models to generate high-contrast multiphase physical system samples, significantly improving physical and distributional fidelity.
Yining Huang, Zhenyu Liang
Proposes a depth-constrained adaptation framework, analyzing encoder-decoder regions' impact on forgetting, with key focus on shallow encoder and deep decoder layers.
Amal Saqib, Tausifa Jan Saleem, Numan Saeed et al.
AutoDesign employs a meta-harness optimization framework to recursively improve long-horizon multimodal design quality, achieving a top score of 78.32.
Yaxin Luo, Haobin Jiang, Jialv Zou et al.
PlayWorld employs multi-modal agent players to evaluate 171 scenarios, revealing current world models' weaknesses in spatial consistency and persistent state evolution.
Kaixin Ding, Xi Chen, Minghong Cai et al.
TabSOM leverages Self-Organizing Maps to encode tabular data into images, achieving top classification performance and interpretability across multiple datasets.
David Chushig-Muzo, María Ángeles Rodríguez de Cara, Eva Milara et al.
RippleNet employs local differential signals with multi-scale, multi-directional modeling and frequency-guided attention to improve cross-generator fake image detection, achieving 89.0% accuracy.
Jiazhen Yang, Ruijin Jin, Junjun Zheng et al.
StateFlow employs a structured 3D world state model with three stages—construction, evolution, and access—to enable high-quality, controllable previsualization for film and game design.
Yuyang Yin, Zixiang Li, Longxuan Deng et al.