VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
Proposes Cross-Modal Verification to build VeriSciQA, a dataset with 20,272 high-quality scientific visual QA pairs.
Yuyi Li, Daoyuan Chen, Zhen Wang et al.
Proposes Cross-Modal Verification to build VeriSciQA, a dataset with 20,272 high-quality scientific visual QA pairs.
Yuyi Li, Daoyuan Chen, Zhen Wang et al.
GigaWorld-0 combines large-scale video synthesis and 3D reconstruction as a data engine, significantly enhancing embodied AI generalization.
GigaWorld Team, Angen Ye, Boyuan Wang et al.
Proposes min-Sliced Transport Plans (min-STP) for scalable, transferable optimal transport, achieving near-OT accuracy with reduced computation.
Xinran Liu, Elaheh Akbari, Rocio Diaz Martin et al.
Zhang et al. propose IF-Edit, a tuning-free framework leveraging pretrained video diffusion models for zero-shot image editing, excelling in non-rigid and reasoning tasks.
Zechuan Zhang, Zhenyuan Chen, Zongxin Yang et al.
COVT introduces continuous visual tokens into VLMs, boosting spatial reasoning by up to 16% on benchmarks.
Yiming Qin, Bomin Wei, Jiaxin Ge et al.
Introduces RLER, combining search and evolving rubrics, training DR Tulu-8B to outperform open-source models with 15.6% gain on benchmarks.
Rulin Shao, Akari Asai, Shannon Zejiang Shen et al.
This study systematically evaluates multilingual embedding models, contrastive learning, and re-ranking, showing dense retrieval surpasses translation-based methods with over 15% improvement in Recall@100.
Roksana Goworek, Olivia Macmillan-Scott, Eda B. Özyiğit
LAST integrates visual tools to enhance spatial-temporal reasoning, boosting VLMs' performance on 3D and long video understanding by 15.8% and 8.3% respectively.
Shuai Wang, Daoan Zhang, Tianyi Bai et al.
SENTINEL achieves end-to-end language-action control with a 99.45% success rate.
Yuxuan Wang, Haobin Jiang, Shiqing Yao et al.
Lehmer code enhances permutation search efficiency; theoretical bounds and empirical tests confirm its advantage over classical encodings.
Yuxuan Ma, Valentino Santucci, Carsten Witt
Proposes Distance-Diversified Top-k Subgraph Matching (DTkSM) with a partition framework, achieving up to 4 orders of magnitude speedup and high diversity.
Liuyi Chen, Yuchen Hu, Zhengyi Yang et al.
ViCoDR introduces multi-view consistency regularization, significantly improving 3D coherence in diffusion-based video generation.
Duolikun Danier, Ge Gao, Steven McDonagh et al.
Proposes VLP, combining VLM perception with program synthesis, achieving significant improvements in complex logical visual reasoning.
Antonia Wüst, Wolfgang Stammer, Hikaru Shindo et al.
Nemotron-Flash optimizes depth-width ratios and operator choices to significantly enhance latency and throughput of small language models.
Yonggan Fu, Xin Dong, Shizhe Diao et al.
DLR framework uses information-theoretic pattern discovery to generate diverse, high-success trajectories for VLA pretraining, improving downstream generalization.
Rushuai Yang, Zhiyuan Feng, Tianxiang Zhang et al.
HuggingR$^4$ introduces a progressive reasoning framework, achieving 92.03% workability, 82.46% reasonability, and reducing token use by 6.9×.
Shaoyin Ma, Chenggong Hu, Huiqiong Wang et al.
Proposes self-adaptive visual bases using Orthogonal Filtering to reduce visual tokens for efficient representation learning.
Shawn Young, Xingyu Zeng, Lijian Xu
CNN-based camera pose estimation for aircraft inspection achieves <0.24m and 2° error.
Xueyan Oh, Leonard Loh, Shaohui Foong et al.
SAMBA model leverages spatial embedding and differential Mamba for long-context EEG modeling, significantly improving accuracy.
Jiazhen Hong, Geoffrey Mackellar, Soheila Ghane
Splatblox uses Gaussian Splatting for outdoor robot navigation, achieving a 50% success rate increase.
Samarth Chopra, Jing Liang, Gershom Seneviratne et al.