GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.AI 2511.15830

Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions

Mini Amusement Parks (MAPs) is a testbed for modeling business decisions, with experts outperforming AI by 11.4x.

Stéphane Aroca-Ouellette, Ian Berlot-Attwell, Panagiotis Lymperopoulos et al.

2025-11-20 10
cs.CV 2511.15661

VisPlay: Self-Evolving Vision-Language Models from Images

VisPlay uses self-evolving RL to improve vision-language models' reasoning via unlabeled image data.

Yicheng He, Chengsong Huang, Zongxia Li et al.

2025-11-20 35
cs.IR 2511.15389

Unveiling Inference Scaling for Difference-Aware User Modeling in LLM Personalization

Proposed DRP framework enhances LLM personalization via inference scaling, achieving a 23% BLEU improvement.

Suyu Chen, Yimeng Bai, Yulong Huang et al.

2025-11-19 18
cs.CL 2511.15304

Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

Poetic style as a universal single-turn jailbreak, achieving 62% attack success across 25 LLMs, with high cross-domain transferability.

Piercosma Bisconti, Matteo Prandi, Federico Pierucci et al.

2025-11-19 62
cs.LG 2511.15248

EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control

EntroPIC employs proportional-integral control to stabilize entropy, enhancing exploration in large language model training.

Kai Yang, Xin Xu, Yangkun Chen et al.

2025-11-19 50
cs.DB 2511.15090

SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning

SciEGQA introduces semantic-region grounding and Grounding–Crop–then–Answer, with 1,623 human QA pairs and 30K+ training pairs.

Wenhan Yu, Zhaoxi Zhang, Wang Chen et al.

2025-11-19 32
cs.AI 2511.14730

Heterogeneous Multi-Agent Proximal Policy Optimization for Power Distribution System Restoration

HAPPO achieves over 95% load restoration in large-scale distribution systems using heterogeneous multi-agent PPO with centralized critic.

Parya Dolatyabi, Ali Farajzadeh Bavil, Mahdi Khodayar

2025-11-19 50
cs.CL 2511.14460

Agent-R1: A Unified and Modular Framework for Agentic Reinforcement Learning

Agent-R1 introduces step-level trajectory abstraction and flexible context management, enabling multi-turn reinforcement learning with diverse optimization strategies, achieving state-of-the-art results.

Mingyue Cheng, Shuo Yu, Daoyu Wang et al.

2025-11-18 31 citations 34
cs.CV 2511.14349

ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries

ARC-Chapter leverages large-scale multimodal data and hierarchical annotations, achieving 14% F1 and 11.3% SODA improvements on long video chaptering.

Junfu Pu, Teng Wang, Yixiao Ge et al.

2025-11-18 49
cs.LG 2511.14117

Distributions In, Distributions Out: The Case for Soft-Label Training

Introduced soft-label training, reducing KL divergence by 32% on datasets like ChaosNLI, enhancing model uncertainty expression.

Agamdeep Singh, Ashish Tiwari, Hosein Hasanbeig et al.

2025-11-18 17
cs.CV 2511.13648

PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image

PhysX-Anything uses VLM and a novel 3D tokenization to generate high-quality, physically grounded assets from a single image, enabling direct simulation deployment.

Ziang Cao, Fangzhou Hong, Zhaoxi Chen et al.

2025-11-18 30
cs.AI 2511.13288

Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO

M-GRPO extends Group Relative Policy Optimization for hierarchical multi-agent systems, aligning heterogeneous trajectories to improve reasoning performance.

Haoyang Hong, Jiajun Yin, Yuan Wang et al.

2025-11-17 42
cs.CV 2511.13269

Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation

This study introduces SpatialSky-Bench and Sky-VLM, achieving SOTA in 13 UAV spatial reasoning tasks with 53.3 average score, surpassing baselines by over 130%.

Lingfeng Zhang, Yuchen Zhang, Hongsheng Li et al.

2025-11-17 69
cs.CV 2511.13259

GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models

GeoX-Bench benchmarks large multimodal models for cross-view geo-localization and pose estimation, showing strong localization but challenges in pose accuracy.

Yushuo Zheng, Jiangyong Ying, Huiyu Duan et al.

2025-11-17 31
cs.CV 2511.13108

DGS-Net: Distillation-Guided Gradient Surgery for CLIP Fine-Tuning in AI-Generated Image Detection

DGS-Net uses gradient decomposition and distillation to improve CLIP fine-tuning for AI image detection, achieving a 6.6% accuracy boost.

Jiazhen Yan, Ziqiang Li, Fan Wang et al.

2025-11-17 49
cs.CL 2511.13043

Spark-Prover-X1: Formal Theorem Proving Through Diverse Data Training

Spark-Prover-X1 enhances formal theorem proving via diverse data training, solving 27 problems on PutnamBench.

Xinyuan Zhou, Yi Lei, Xiaoyu Zhou et al.

2025-11-17 23
cs.CV 2511.13032

Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts

Uni-Inter synthesizes 3D human motion across diverse interaction contexts using a Unified Interactive Volume (UIV).

Sheng Liu, Yuanzhi Liang, Jiepeng Wang et al.

2025-11-17 36
cs.AI 2511.13027

Scaling Generative Verifiers For Natural Language Mathematical Proof Verification And Selection

Scaled GenSelect and LLM-as-a-Judge to millions of tokens, enhancing math proof verification.

Sadegh Mahdavi, Branislav Kisacanin, Shubham Toshniwal et al.

2025-11-17 30
cs.AI 2511.13007

GEM: Generative Entropy-Guided Preference Modeling for Few-shot Alignment of LLMs

GEM uses entropy-guided generative preference modeling for few-shot LLM alignment, improving preference prediction by 5-10% and task performance significantly.

Yiyang Zhao, Huiyu Bai, Xuejiao Zhao

2025-11-17 44
cs.SE 2511.12884

Agent READMEs: An Empirical Study of Context Files for Agentic Coding

Empirical analysis of 2,303 agent context files reveals their structure, maintenance, and content biases, highlighting insufficient emphasis on security and performance.

Worawalan Chatlatanagulchai, Hao Li, Yutaro Kashiwa et al.

2025-11-17 34
Prev 1 ... 233 234 235 236 237 238 239 ... 573 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home