TanGO: Training-Free 3D Editing via Tangent-Space Guidance and Optimization
TanGO enables training-free 3D editing via tangent-space guidance, significantly reducing structural artifacts.
Siwoo Lim, Sunjae Yoon, Gwanhyeong Koo et al.
TanGO enables training-free 3D editing via tangent-space guidance, significantly reducing structural artifacts.
Siwoo Lim, Sunjae Yoon, Gwanhyeong Koo et al.
HyMobileAgent achieves efficient GUI interaction via data-environment co-scaling, 82.6% success on AndroidWorld.
Hy Vision Team, Huawen Shen, Zhengyang Tang et al.
AdaTurn improves low-budget active visual reasoning by explicit budget conditioning and boundary-aware reinforcement learning, boosting accuracy by over 10%.
Susan Liang, Chao Huang, Filippos Bellos et al.
Uni-AdaVD achieves concept erasure in visual generation via orthogonal value decomposition, showing strong performance across various models.
Qifan Zhou, Yuan Wang, Yanbin Hao et al.
G2SR leverages multi-view geometry with lightweight neural detection to achieve fast (69-89 Hz), low-memory (203MB) Gaussian surface reconstruction, outperforming end-to-end methods.
Dasong Gao, Vivienne Sze, Sertac Karaman
GMoT module explicitly extracts sparse motion cues, boosting micro-gesture video reasoning accuracy to 67.32% Top-1, surpassing baselines.
Taorui Wang, Wei Xia, Hui Ma et al.
UniPhys employs a unified physical grounding framework with a 40K dataset, achieving state-of-the-art articulation and physical property estimation.
Xian Li, Rong Wei, Lujie Yang et al.
FOLIO is a training-free focused semantic memory system combining short-term visual buffers with long-term entity-centric structured memory, boosting streaming video understanding.
Haoyang Fan, Dhruv Parikh, Anvitha Ramachandran et al.
Proposes Hy-Embodied-VLM-1.0, integrating action-centric taxonomy, optimized data pipeline, and Mixture-of-Experts architecture, achieving state-of-the-art multi-task embodied agent performance.
Ziyi Wang, Xumin Yu, Yongming Rao et al.
UniVR uses VR-GRPO for native visual reasoning, improving VR-X over Emu3.5 by 18.4% overall and 25.2% in long-term planning.
Zhongwei Ren, Yunchao Wei, Yao Zhao et al.
Proposes CoRe framework with structured rewards and auto-constructed dataset to improve fine-grained cross-image reasoning.
Lin Peng, Cong Wan, Zeyu Guo et al.
LookME introduces hierarchical two-level retrieval and sparse injection to enhance multimodal embeddings in vision-language models, outperforming text-only PLE methods with significant efficiency gains.
Zeyu Xu, Xingzhong Hou, Pengkai Guo et al.
CAtFM enhances flow matching with contrastive learning for style-content disentanglement, improving performance on datasets like ImageNet.
Yusong Li, Pingchuan Ma, Ming Gui et al.
SeamGen uses flow matching and Mesh Transformer to generate artist-aligned UV seams, outperforming traditional methods with 20% lower deviation.
Hao Xu, Yuqing Zhang, Yiqian Wu et al.
SARSI couples a self-model with evidence-gated recursive improvement; this paper is a design proposal, not an empirical result.
Chengshuai Yang
RegHead generates non-humanoid head blendshapes via feed-forward registration, faster than optimization methods.
Jiahao Luo, Hao Zhang, Jianqi Chen et al.
SymbOmni employs symbolic concept learning for continuous model evolution, boosting generation quality and efficiency.
Jinxiu Liu, Jianru Li, Tanqing Kuang et al.
ABot-3DWorld 0 introduces a unified spatial primitive for multimodal 3D scene generation, surpassing state-of-the-art in fidelity and versatility.
Mingchao Sun, Luyang Tang, Yu Liu et al.
DynEval enhances T2I model evaluation by dynamically assessing text-image alignment and quality using GenDB and DynEvalInstruct datasets.
Shyam Marjit, Dheeraj Baiju, Anuj Shikarkhane et al.
SynthDocBench uses synthetic long documents with controlled factors to reveal three key failure modes of vision-language models in long-context understanding.
Abhigya Verma, Khyati Mahajan, Amit Kumar Saha et al.