GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.AI 2609.20804

An Empirical Study of Harness Design for Coding Agents

Study improves coding agents' long-term performance via lightweight harness design; context management proves most effective.

Run-Ze Fan, Zihao Zhang, Simin Ma et al.

2026-09-18 14
cs.AI 2609.19680

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

FINSKILLOPS improves SEC filing QA correctness through self-evolving multi-agent system, raising it from 3.70 to 4.55.

Yanzhang Ma, Zhenghan Tai, Hanwei Wu et al.

2026-09-17 13
cs.AI 2609.18723

Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning

Introduces Mahalanobis-Ensemble Decoding to optimize LLM decoding via ensemble pruning, enhancing semantic diversity and generation quality.

Dunyao Xue, Chengshuo Du, Zhengbo Wang et al.

2026-09-16 15
cs.AI 2609.18004

Missing Bridges: Composition-Aware Active Imitation Learning

AALT selects demonstrations to maximize start-goal connectivity, improving task success rate.

Maxwell J. Jacobson, Ahmed H Qureshi, Yexiang Xue

2026-09-16 12
cs.AI 2609.12851

MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations

MedRoundsQA evaluates medical diagnosis through multi-turn dialogues, with accuracy dropping 13-39 points.

Youssef Mohamed, Ahmed Heakl, Qinrong Cui et al.

2026-09-11 0
cs.AI 2609.10824

Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

The study introduces META-AGENT for task-agnostic preprocessing in unknown environments, achieving highest Avg@3 reward on five benchmarks.

Vinay Samuel, Varun Ursekar, Vijay S. Kalmath et al.

2026-09-10 9
cs.AI 2609.10451

JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task Composition

JarvisGUI evaluates cross-device GUI agents with dynamic task composition, revealing capability gaps in real-world workflows.

Zixiang Chen, Yuheng Lu, Zihao Cheng et al.

2026-09-10 94
cs.AI 2609.10441

ConvMem: Convolutional Memory for Long-Context Reasoning

ConvMem reformulates long-context reasoning as hierarchical convolution, outperforming training-free baselines on RULER-HotpotQA.

Hongming Zhang, Zhaozhen Gu, Fengshuo Bai et al.

2026-09-10 95
cs.AI 2609.10413

Fortunate Recall: Ontology-Driven Memory Lifecycle Management for Persistent Coherence in LLMs

Fortunate Recall uses a 10+1 behavioral ontology and lifecycle policies to enhance LLM memory management, achieving 76.9% on LifecycleBench.

Ansuman Mullick, Eray Tüzün

2026-09-10 81
cs.AI 2609.07987

When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability

LLM digital twins reduce human measurement via statistical substitutability, but behavioral fidelity is insufficient.

Steven Wang, Kyle Hunt, Shaojie Tang et al.

2026-09-08 6
cs.AI 2609.05396

A Deep Generative Model for Synthesizing Labeled Wireless Signals

Introduced IIns-GAN for generating labeled wireless signals, enhancing model training.

Yuxiao Li, Keke Hu, Santiago Mazuelas et al.

2026-09-05 96
cs.AI 2609.05395

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

EDGE framework synthesizes tool-calling data using dynamic graphs, enhancing performance.

Dain Kim, Eungi Cho, Kyumin Kim et al.

2026-09-05 97
cs.AI 2609.05385

Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence

Evaluates LLM explanations' necessity and sufficiency using behavioral evidence, finding limited correlation with model behavior.

Urja Pawar, Rajitha Ramanayake, Nabeel Kemal et al.

2026-09-05 93
cs.AI 2609.05381

Molecular Déjà Vu: Digit-Level Retrieval of Published Values in Frontier Language Models

Study finds widespread verbatim retrieval in LLMs on molecular regression benchmarks, affecting prediction accuracy.

Matthias Busch, Marius Tacke, Sviatlana V. Lamaka et al.

2026-09-05 91
cs.AI 2609.05374

CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI Agents

CUA-Universe uses App-Forge, Task-Weave, and Path-Steer to create hybrid GUI+CLI environments, enhancing efficiency and success rates.

Haoting Shi, Wenhao Wang, Weicheng Fang et al.

2026-09-05 99
cs.AI 2609.05339

Does Your Agent's Memory Survive a Model Upgrade? A Controlled Study of Memory Portability

The study examines memory portability during model upgrades, finding fixed-schema knowledge graphs remain stable.

Ankit Goyal, Jaideep Ray

2026-09-05 101
cs.AI 2609.05333

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

Toolkit for measuring contextual individuation in Transformer models using bridge forms.

José Luciano Verçosa Marques, Frederico Jorge Heitmann, Daniel Omar Perez et al.

2026-09-05 86
cs.AI 2609.05314

Large Language Models for HVAC Operations in Building Energy Systems: A Critical Review of Methods, Applications, and Deployment Readiness

Large language models for HVAC operations in building energy systems; none ready for industry deployment.

Alexander Neubauer, Tianzhen Hong, Han Li et al.

2026-09-05 42
cs.AI 2609.05040

Towards Efficient Evaluation of Evolutionary Transfer Optimization: Case Studies on Task-Parameterized Applications

Achieved efficient ETO evaluation via accumulation and blending matrix reformulations, with up to 256.72x speedup.

Yanchen Li, Xiaoming Xue, Kay Chen Tan

2026-09-04 55
cs.AI 2609.04981

A Tree-based RAG Framework for Evidence-Intensive QA via Adaptive Planning and Topology-Aware Evidence Gathering

APT-RAG framework excels in evidence-intensive QA, achieving a 40.69% F1 score improvement.

Songeun Lee, Kyungjin Min, Injae Na et al.

2026-09-04 34
1 2 3 4 ... 42 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home