GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.AI 2608.23070

From Generation to Simulation: How Far Are World Models from Being True Simulators?

Using a capability-based framework, the study assesses how close generative world models are to being true simulators, highlighting gaps in physical laws and state feedback.

Tong Wang, Huan Deng, Mucheng Yang et al.

2026-08-24 40
cs.AI 2608.22847

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

GSAR combines self-evolving data synthesis and goal-state anchoring to improve GUI agent training with over 90% accuracy in trajectory verification.

Long Zhang, Yuhan Chen, Chaoran Zhang et al.

2026-08-24 48
cs.AI 2608.21357

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

VIALS benchmark evaluates scientific artifact interpretation, with top models achieving only 26.5% accuracy, highlighting major gaps.

Elaine Lau, Thanuka Udumulla, Lee Izhaki-Tavor et al.

2026-08-22 73
cs.AI 2608.21292

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

AUSO unifies skill internalization and utilization via action-aware progressive reinforcement learning, improving long-horizon task performance.

Huizu Lin, Chengkai Huang, Tianqi Gao et al.

2026-08-22 67
cs.AI 2608.21278

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

CLEAR uses input-conditioned latent gating to reduce HarmBench ASR from 32.3% to 0.5%, maintaining high utility on GSM8K and other tasks.

Chengxiao Wang, Enyi Jiang, Xiaojing Liao et al.

2026-08-22 62
cs.AI 2608.21233

Fine-Grain GPU Parallelization of the Generalized Partition Crossover for Large-Scale Traveling Salesman Problems

GPU fine-grain parallelization of GPX partitioning accelerates large-scale TSP solving by 48-625×, reducing memory use significantly.

Swetha Varadarajan, Darrell Whitley

2026-08-21 56
cs.AI 2608.21218

Enhancing LLMs in Predictive Political QA with Semi-Structured Data

Proposed PSL framework combines semi-structured political records' stance and structural signals, significantly improving predictive political QA.

Yinan Liu, Zihan Zhou, Zichun Jin et al.

2026-08-21 67
cs.AI 2608.21027

Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents

COTA enhances runtime intervention in LLM agents by comparing alternatives, improving performance in environments like WebShop.

Yanze Jiang, Mingxuan Li, Yuhao Wang et al.

2026-08-21 5
cs.AI 2608.20918

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

UpgradeBench offers a decision-centric benchmark for upgrading LLM specialists, with a mean quality regret of 0.37pp.

Ye Chen, Weining Zhang

2026-08-21 4
cs.AI 2608.20845

RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

Proposes ingest-time semantic compilation (ISC) with maintained semantic index, outperforming query-time interpretation in accuracy and cost efficiency.

Kyle Wild, Yusuke Takahashi, Asako Uraki

2026-08-21 67
cs.AI 2608.20743

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

This paper systematically evaluates diffusion-based block-parallel decoding in multimodal models, highlighting potential speedups up to 3.6× with current architectures.

Yantao Li, Huanlin Gao, Fang Zhao et al.

2026-08-21 36
cs.AI 2608.20320

An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction

Proposed a multi-agent framework integrating conversational data collection, structured processing, and large language model prediction for weather-sensitive travel behavior, achieving up to 71.5% accuracy.

Narges Ahmadi, Yubo Jiao, Jônatas Augusto Manzolli et al.

2026-08-21 80
cs.AI 2608.20318

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

AI4AI-Bench evaluates LLM agents' recursive self-improvement in training algorithms; mean score 0.166, top 0.250, across 10 algorithm families.

Yizhe Chi, Wenyi Li, Deyao Hong et al.

2026-08-21 75
cs.AI 2608.20290

Phantom Gains: Auditing Self-Improvement Against a Measured Null

Introduces null-model based exact tests and multi-criteria trajectory analysis to reliably measure genuine self-improvement in language models, avoiding measurement artifacts.

Cheng Xu, Nan Yan, Liming Chen et al.

2026-08-21 91
cs.AI 2608.20274

Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents

This study compares task-level and subtask-level skill induction, finding subtask and text-format skills transfer more reliably, and introduces a skill utility score.

Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian et al.

2026-08-21 77
cs.AI 2608.20271

Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning

Using XGBoost with first 5-minute trading data, predict Solana memecoin rug pulls within 1 hour.

Jianghai Li, Pavel Kuznetsov, Yury Yanovich et al.

2026-08-21 86
cs.AI 2608.19804

ADAPT: Physics-Aware Diffusion-based World Models for Adaptive Predictive Transferable HVAC Control

Proposes ADAPT, a physics-informed diffusion-based indoor model, achieving 7.3% energy savings and 30.2% occupant comfort improvement.

Xu Yang, Kailai Sun, Dianyu Zhong et al.

2026-08-20 40
cs.AI 2608.19535

From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

Proposes telemetry-guided adaptive compression for edge RAG, reducing GPU energy by up to 53.2% and SoC energy by 48.2%, with negligible quality loss.

Zlatan Feric, Amir Taherin, Yanzhi Wang et al.

2026-08-20 79
cs.AI 2608.19072

What is Missing from AI Post-Training AI: An Empirical Analysis

Empirical analysis of 1,338 LLM post-training trajectories reveals strategies are locked in early, lacking in-execution self-reassessment mechanisms, limiting AI self-improvement.

Joy Jia Yin Lim, Xin Huang, Hao Peng et al.

2026-08-20 149
cs.AI 2608.18836

Verifiable abstention makes AI leak diagnosis accountable in urban water distribution networks

Verifiable abstention makes AI leak diagnosis accountable, with 214 correct actions out of 550 events.

Tianwei Mu, Yue Wang, Mingzhe Yuan et al.

2026-08-19 6
Prev 1 2 3 4 5 6 7 ... 42 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home