GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.RO 2606.29917

Flying to Image-Specified Objects: 3D Quadrotor Navigation via Cross-Graph Memory and Viewpoint Planning

Hierarchical navigation with cross-graph memory and viewpoint planning improves quadrotor image-based object localization, achieving 88% success in simulation.

Junjie Gao, Yuqi Chen, Yongzhou Pan et al.

2026-06-29 76
cs.RO 2606.29898

Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies

CI-MSE improves offline validation by focusing on task-critical segments, enhancing correlation with real-world performance.

Haoxu Huang, Tongsam Zheng, Yifan Chen et al.

2026-06-29 60
cs.IR 2606.29894

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

SABER-Math constructs an automated, fine-grained benchmark for mathematical IR using large models, ontology, and preference ranking, outperforming traditional methods.

Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova et al.

2026-06-29 35
cs.LG 2606.29888

Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders

Proposes modality-specific sparse autoencoders and post-hoc alignment to address cross-modal feature heterogeneity, boosting retrieval accuracy to 85.3% on MS-COCO.

Chungpa Lee, Jihoon Kwon, Kyle Min et al.

2026-06-29 49
cs.LG 2606.29820

Dual-Flow Reinforcement Learning with State-Aware Exploration

Dual-Flow RL employs conditional flow matching to jointly model return distribution and multimodal policy, enhancing exploration and value estimation in continuous control.

Qijun Li, Zheng Fu, Qi Song et al.

2026-06-29 25
cs.CV 2606.29814

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

Nemotron-Labs-Diffusion-Image introduces a masked discrete diffusion model with token editing and GCE, achieving 0.90 on GenEval for high-res text-to-image synthesis.

Shufan Li, Greg Heinrich, Hanrong Ye et al.

2026-06-29 42
cs.RO 2606.29774

Analytic Concept-Centric Memory for Agentic Embodied Manipulation

Concept-centric memory framework improves long-horizon embodied manipulation by explicit object structure, achieving 70% success on RMBench.

Mingyang Sun, Xiujian Liang, Jiude Wei et al.

2026-06-29 42
cs.AI 2606.29705

GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots

GUICrafter enhances GUI agent cross-device generalization using weak supervision and massive unannotated screenshots.

Sunqi Fan, Lingshan Chen, Runqi Yin et al.

2026-06-29 12
cs.CV 2606.29699

Early Warning Signals for OpenVLA Failure under Visual Distribution Shift

By freezing OpenVLA policy, linear probes detect failure signals under visual distribution shift, achieving AUROC of 0.972.

Dipesh Tharu Mahato, Rachel Ren

2026-06-29 38
cs.AI 2606.29630

SFBench: The SciFy Scientific Feasibility Benchmark

SFBench is a benchmark dataset with 197 material science claims for evaluating scientific claim feasibility.

Cash Costello, James Mayfield, Elsbeth Turcan et al.

2026-06-29 10
cs.CV 2606.29600

One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models

Introduced MultiDepth-3k benchmark and Laplacian Visual Prompting (LVP) to probe depth-layer preferences; DAv2-L achieved 75.5% ML-SRA.

Xiaohao Xu, Feng Xue, Xiang Li et al.

2026-06-29 46
cs.SE 2606.29538

RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources

RESOURCE2SKILL distills multimodal human resources into executable skills, improving agent performance by +11.9 points on average.

Yijia Fan, Zonglin Di, Zimo Wen et al.

2026-06-29 51
cs.AI 2606.29495

Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents

CogWM tracks user BDI/E states for cognitive evaluation of conversational agents, trained on 150K samples, outperforming baselines.

Minghui Ma, Bin Guo, Hao Wang et al.

2026-06-29 38
cs.CV 2606.29462

MIRROR: Aligning Semantic Relations from Language to Image via Gromov--Wasserstein

MIRROR employs Gromov–Wasserstein regularization to transfer semantic relation geometry from language to vision, boosting multimodal relational reasoning.

Hong-Han Wang, Yuntao Wang, Hu Ding

2026-06-28 38
cs.CV 2606.29461

From Phase to Phenomenon: Self-Supervised Learning of Subsurface Scattering with Minimal Phase-shift Inputs

Proposed a self-supervised framework using eight phase-shift images to learn subsurface scattering representations.

Arjun Majumdar, Raphael Braun, Andreas Engelhardt et al.

2026-06-28 19
cs.CV 2606.29445

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

TASKER algorithm improves video understanding by task-driven and scene-aware keyframe extraction, achieving a 2.0% gain on EgoSchema.

Sunqi Fan, Qingle Liu, Runqi Yin et al.

2026-06-28 16
cs.CY 2606.29390

Toward Comprehensive Risk Assessments and Assurance of AI-Based Systems

Proposes an AI risk assessment framework integrating Operational Design Domains (ODD) to define safety boundaries, improving hazard identification.

Heidy Khlaaf

2026-06-28 59
physics.comp-ph 2606.29220

Latent Genetic Algorithm for Crystal Structure Prediction

Latent Genetic Algorithm (LGA) improves HfO2 ground-state recovery rate to 60-95% via latent space crossover.

Kaixin Zheng, Wanjian Yin, Hongyu Yu et al.

2026-06-28 60
cs.RO 2606.29173

TacGen: Touch Is a Necessary Dimension of Physical-World Representation -- Addressing Tactile Data Scarcity with Scalable Vision-to-Touch Alignment and Generation

TacGen aligns vision and touch via contrastive learning and latent diffusion, effectively addressing tactile data scarcity and improving physical property representation.

Wanghao Ye, Aarosh Das, Sihan Chen et al.

2026-06-28 53
cs.LG 2606.29164

Invariant Reasoning Directions in Latent Trajectories of Language Models

Introduces TILR, a low-rank subspace method that identifies stable reasoning directions, improving model consistency by ~10% and reducing trajectory variance by 50%.

Arun Vignesh Malarkkan, Manan Roy Choudhury, Utkarsh Byahut et al.

2026-06-28 65
Prev 1 ... 83 84 85 86 87 88 89 ... 552 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home