GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.RO 2604.01158

SMASH: Mastering Scalable Whole-Body Skills for Humanoid Ping-Pong with Egocentric Vision

SMASH integrates onboard egocentric vision and scalable whole-body skill learning, enabling humanoid robots to perform continuous outdoor ping-pong without external sensors.

Junli Ren, Yinghui Li, Kai Zhang et al.

2026-04-02 62
cs.SE 2604.01029

Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines

Multi-LLM revision pipelines decompose second-pass gains into re-solving, scaffold, and content effects, enhancing MCQ and code generation tasks.

Jingjie Ning, Xueqi Li, Chengyu Yu

2026-04-01 11
cs.LG 2604.16411

CGCMA: Conditionally-Gated Cross-Modal Attention for Event-Conditioned Asynchronous Fusion

CGCMA uses conditionally-gated cross-modal attention to achieve event-conditioned asynchronous fusion, improving Sharpe ratio to +0.449 on CryptoMI dataset.

Yunxiang Guo

2026-04-01 31
cs.RO 2604.01259

Bench2Drive-VL: Benchmarks for Closed-Loop Autonomous Driving with Vision-Language Models

Proposes Bench2Drive-VL, integrating DriveCommenter for real-time closed-loop autonomous driving evaluation with multimodal reasoning.

Xiaosong Jia, Yuqian Shao, Zhenjie Yang et al.

2026-04-01 80
cs.IR 2604.00590

UniMixer: A Unified Architecture for Scaling Laws in Recommendation Systems

UniMixer introduces a unified recommendation scaling architecture with parameterized TokenMixer, combining attention, TokenMixer, and FM advantages for improved efficiency.

Mingming Ha, Guanchen Wang, Linxun Chen et al.

2026-04-01 53
cs.LG 2604.00556

HabitatAgent: An End-to-End Multi-Agent System for Housing Consultation

HabitatAgent achieves 95% accuracy in housing consultation using a multi-agent system.

Hongyang Yang, Yanxin Zhang, Yang She et al.

2026-04-01 15
cs.CV 2604.00279

The Geometry of Compromise: Unlocking Generative Capabilities via Controllable Modality Alignment

Introduced TPC-CMA method, reducing modality gap by 82.3% with only 4.84% accuracy drop.

Hongyuan Liu, Qinli Yang, Wen Li et al.

2026-04-01 11
cs.RO 2604.00202

DreamControl-v2: Simpler and Scalable Autonomous Humanoid Skills via Trainable Guided Diffusion Priors

DreamControl-v2 trains guided diffusion models directly in robot space, integrating diverse datasets to enhance scalability and automation for humanoid skills.

Sudarshan Harithas, Sangkyung Kwak, Pushkal Katara et al.

2026-04-01 66
math.FA 2603.30039

The Grothendieck Constant is Strictly Larger than Davie-Reeds' Bound

Using perturbative analysis, the paper proves that the Grothendieck constant K_G exceeds Davie-Reeds' lower bound by at least 10^−12, marking a first quantitative improvement since the 1980s.

Chris Jones, Giulio Malavolta

2026-04-01 5 citations 92
cs.AI 2603.29908

C-TRAIL: A Commonsense World Framework for Trajectory Planning in Autonomous Driving

C-TRAIL integrates LLMs with trust mechanisms via a closed-loop Recall-Plan-Update framework for autonomous driving trajectory planning.

Zhihong Cui, Haoran Tang, Tianyi Li et al.

2026-03-31 49
cs.CV 2603.29460

Square Superpixel Generation and Representation Learning via Granular Ball Computing

Proposes a granular ball-based square superpixel method enabling end-to-end deep learning integration, improving efficiency and structure.

Shuyin Xia, Meng Yang, Dawei Dai et al.

2026-03-31 45
cs.CV 2603.29281

PRISM: A Multi-View Multi-Capability Retail Video Dataset for Embodied Vision-Language Models

PRISM leverages multi-view videos and a 3D knowledge ontology to enhance embodied VLM's spatial, physical, and action understanding in retail environments.

Amirreza Rouhi, Parikshit Sakurikar, Satya Sai Reddy et al.

2026-03-31 42
cs.CV 2603.29194

Multi-Layered Memory Architectures for LLM Agents: An Experimental Evaluation of Long-Term Context Retention

Multi-Layer Memory Framework enhances LLMs' long-term context retention, achieving 46.85% success rate and 56.90% six-period retention.

Sunil Tiwari, Payal Fofadiya

2026-03-31 22
cs.RO 2603.29192

Efficient Camera Pose Augmentation for View Generalization in Robotic Policy Learning

GenSplat enhances view generalization in robotic policy learning using 3D Gaussian Splatting.

Sen Wang, Huaiyi Dong, Jingyi Tian et al.

2026-03-31 40
cs.CV 2603.29165

LatentPilot: Scene-Aware Vision-and-Language Navigation by Dreaming Ahead with Latent Visual Reasoning

LatentPilot internalizes future visual dynamics via privileged supervision, achieving SOTA in VLN benchmarks with 66.3% success rate.

Haihong Hao, Lei Chen, Mingfei Han et al.

2026-03-31 51
cs.CV 2603.29163

SparseDriveV2: Scoring is All You Need for End-to-End Autonomous Driving

SparseDriveV2 leverages dense static vocabularies with trajectory decomposition, achieving 92.0 PDMS, challenging the necessity of dynamic proposals.

Wenchao Sun, Xuewu Lin, Keyu Chen et al.

2026-03-31 44
cs.AI 2603.29112

GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification

GISTBench uses IG and IS to test whether LLM user profiles are supported by behavior; survey alignment reaches ρ=0.67.

Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan et al.

2026-03-31 26
cs.CV 2603.29009

MEDiC: Multi-objective Exploration of Distillation from CLIP

MEDiC unifies CLIP distillation and pixel reconstruction, reaching 73.9% kNN and 85.1% fine-tuning accuracy on ImageNet-1K.

Konstantinos Georgiou, Maofeng Tang, Hairong Qi

2026-03-31 30
cs.RO 2603.28565

StreamingVLA: Streaming Vision-Language-Action Model with Action Flow Matching and Adaptive Early Observation

StreamingVLA achieves 2.4× speedup and 6.5× halting reduction via Action Flow Matching and Adaptive Early Observation.

Yiran Shi, Dongqi Guo, Tianchen Zhao et al.

2026-03-30 42
cs.CR 2603.28551

"What Did It Actually Do?": Understanding Risk Awareness and Traceability for Computer-Use Agents

Proposes AgentTrace, a visualization framework that enhances user understanding and traceability of high-authority AI agents' actions, improving risk awareness and post-hoc auditability.

Zifan Peng, Mingchen Li

2026-03-30 63
Prev 1 ... 172 173 174 175 176 177 178 ... 573 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home