GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.LG 2606.29164

Invariant Reasoning Directions in Latent Trajectories of Language Models

Introduces TILR, a low-rank subspace method that identifies stable reasoning directions, improving model consistency by ~10% and reducing trajectory variance by 50%.

Arun Vignesh Malarkkan, Manan Roy Choudhury, Utkarsh Byahut et al.

2026-06-28 65
cs.CV 2606.29097

TrafficAlign: Aligning Large Language Models for Traffic Scenario Generation

TrafficAlign synthesizes traffic scenarios from videos, uses DSL validation, and fine-tunes LLMs for regional traffic distribution alignment, improving autonomous driving safety.

Zhi Tu, Liangkun Niu, Tianyi Zhang

2026-06-28 40
cs.CV 2606.29023

Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification

Efficient spatio-temporal grounding via second-level tracking and RL verification, enhancing localization quality.

Tianshu Zhang, Yan Wang, Ji Qi et al.

2026-06-28 4
cs.CL 2606.28938

EVLA: An Electro-Aware Multimodal Assistant for Physically-Grounded Driving Reasoning and Control

EVLA fuses multimodal perception with vehicle physics via UCSE and ESRC, achieving energy-efficient driving decisions with +0.0871 score improvement.

Yuxin Liu, Zihan Chen, Haoyu Wang et al.

2026-06-27 63
cs.CV 2606.28840

DLGStream: Dynamic Language-embedded Guassian Splatting for Open-vocabulary Enabled Free-viewpoint Video Streaming

DLGStream introduces dual-opacity language Gaussian representation with deformation fields, achieving 43KB/frame and 60FPS in open-vocabulary free-viewpoint video streaming.

Zhihui Ke, Yuyang Liu, Xiaobo Zhou et al.

2026-06-27 35
cs.CV 2606.28828

Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video

Ground4D integrates VGGT-based geometry initialization with deformable Gaussian Splatting for monocular 4D scene reconstruction and novel view synthesis, achieving state-of-the-art accuracy.

Qing Zhao, Weijian Deng, Pengxu Wei et al.

2026-06-27 36
cs.CV 2606.28820

CoGS: Compositional Dynamic Human-Object Scenes Gaussian Splatting from Monocular Video

CoGS introduces a compositional Gaussian splatting framework with a six-stage optimization for monocular human-object scene reconstruction, achieving state-of-the-art fidelity.

Jerrin Bright, John Zelek

2026-06-27 49
cs.RO 2606.28813

Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning

Human2Any employs object-centric interaction priors learned from human videos, enabling cross-embodiment robot manipulation without real robot demonstrations.

Shuo Cheng, Chuye Zhang, Alfred Cueva et al.

2026-06-27 53
cs.IR 2606.28780

Multimodal Graph RAG for Long-range Visually Rich Document Understanding

KG4VD leverages multimodal knowledge graphs and graph retrieval, achieving 44.52% accuracy on long-range document VQA benchmarks.

Yi-Cheng Wang, Chu-Song Chen

2026-06-27 51
cs.IR 2607.24789

NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model

NEXT framework uses NEXT-8B to improve video recommendations, achieving +0.53% watch time and +0.51% diversity.

Yuming Liu, Hongye Yang, Harrison Zhao et al.

2026-06-27 40
cs.AI 2606.28747

Self-Supervised Theorem Discovery in a Formal Axiomatic System

Proposes a self-supervised theorem discovery algorithm, autonomously generating thousands of meaningful theorems within a formal axiomatic system, enhancing proof capabilities.

Kazuki Ota, Takayuki Osa, Tatsuya Harada

2026-06-27 30
cs.AI 2606.28696

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

COMPASS uses shared token τc to connect composition recognition and generation, trained on 389,031 images and about 1.2M VQA pairs.

Ziqi Zhou, Weize Quan, Mining Tan et al.

2026-06-27 28
cs.CV 2606.28643

Obliviate: Erasing Concepts from Autoregressive Image Generation Models

Obliviate effectively erases concepts in autoregressive image generation, reducing nudity detection on RAB benchmark from 91.58% to 3.15%.

Hossein Shakibania, Jonas Henry Grebe, Tobias Braun et al.

2026-06-27 42
cond-mat.mtrl-sci 2606.28578

Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design

Using surrogate-gated generation and foundation-model embeddings, this study reduces evaluation calls by 90% in Bayesian materials design.

Sk Md Ahnaf Akif Alvi, Jan Janssen, Danny Perez et al.

2026-06-27 68
cs.AI 2606.28556

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

IMCBench evaluates multimodal LLMs in image-grounded medical dialogues; Claude Opus 4.6 scores highest at 3.61.

Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar et al.

2026-06-27 7
cs.CV 2606.28321

StructSplat: Generalizable 3D Gaussian Splatting from Uncalibrated Sparse Views

StructSplat is a pose-free 3D Gaussian reconstruction framework achieving +5.67 dB PSNR over SOTA on DL3DV.

Jia-Chen Zhao, Beiqi Chen, Xinyang Chen et al.

2026-06-27 47
cs.LG 2606.28274

Parameter Efficient Hybrid Transformer (PEHT) for Network Traffic Prediction via Dynamic Urban Congestion Integration

PEHT integrates LoRA-enhanced Transformer with urban mobility and congestion data, achieving state-of-the-art network traffic prediction with over 90% parameter reduction.

Abdolazim Rezaei, Mehdi Sookhak, Mahboobeh Haghparast

2026-06-27 211
cs.LG 2606.28194

COCOLogic-V2: Identifying Logical Inconsistencies via Truly Hard-Negatives

COCOLogic-V2 evaluates visual inductive reasoning with truly hard negatives, revealing models' struggles on boundary samples.

David Steinmann, Antonia Wüst, Kristian Kersting et al.

2026-06-26 51
cs.RO 2606.28192

PA-BiCoop: A Primary-Auxiliary Cooperative Framework for General Bimanual Manipulation

PA-BiCoop framework improves RLBench2 tasks by 48% on average and over 50% in real-world tasks.

Bai Qicheng, Wang Ziru, Ma Teli et al.

2026-06-26 18
cs.LG 2606.28182

LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

LLawCo framework uses failure reflection to extract behavioral laws, improving multi-agent cooperation success rate by 4.5%.

Qinhong Zhou, Chuang Gan, Anoop Cherian

2026-06-26 50
Prev 1 ... 84 85 86 87 88 89 90 ... 552 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home