GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.CL 2605.03858

MCJudgeBench: A Benchmark for Constraint-Level Judge Evaluation in Multi-Constraint Instruction Following

MCJudgeBench evaluates LLM judge accuracy and stability at constraint level, distinguishing intrinsic and procedural inconsistencies.

Jaeyun Lee, Junyoung Koh, Zeynel Tok et al.

2026-05-05 41
cs.LG 2605.03677

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Uni-OPD unifies on-policy distillation with dual strategies, boosting multi-task and multi-modal model performance.

Wenjin Hou, Shangpin Peng, Weinong Wang et al.

2026-05-05 33 citations 83
cs.SE 2605.03546

ProgramBench: Can Language Models Rebuild Programs From Scratch?

ProgramBench evaluates LMs' ability to rebuild programs from scratch; the best model passes 95% of tests on only 3% of tasks, spanning CLI tools to complex software.

John Yang, Kilian Lieret, Jeffrey Ma et al.

2026-05-05 61
cs.CV 2605.03463

First Shape, Then Meaning: Efficient Geometry and Semantics Learning for Indoor Reconstruction

FSTM employs a two-stage training process with a single SDF to achieve indoor scene geometry and semantic reconstruction, improving speed by 2.3×.

Remi Chierchia, Léo Lebrat, David Ahmedt-Aristizabal et al.

2026-05-05 43
cs.AI 2605.03308

Revisiting the Travel Planning Capabilities of Large Language Models

By decomposing travel planning tasks, the study reveals LLMs' deficiencies in implicit constraint inference.

Bo-Wen Zhang, Jin Ye, Peng-Yu Hua et al.

2026-05-05 8
cs.LG 2605.16318

Investigating Action Encodings in Recurrent Neural Networks in Reinforcement Learning

This study compares additive and multiplicative RNN architectures incorporating action embeddings for RL, showing multiplicative models improve efficiency and accuracy.

Matthew Schlegel, Volodymyr Tkachuk, Adam White et al.

2026-05-05 30
cs.CV 2605.03175

DINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery

Proposes CAFe-DINO, leveraging DINOv3 for open-vocabulary remote sensing segmentation without fine-tuning, outperforming supervised models.

Ryan Faulkenberry, Saurabh Prasad

2026-05-05 49
cs.AI 2605.02819

SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering

Proposes SCPRM, integrating schema-aware distance for risk-sensitive multi-hop KG reasoning, improving accuracy by 1.18%.

Jiujiu Chen, Yazheng Liu, Sihong Xie et al.

2026-05-05 44
cs.CV 2605.02730

Perceptual Flow Network for Visually Grounded Reasoning

PFlowNet achieves visual reasoning via variational reinforcement learning, scoring 90.6% on V* Bench.

Yangfu Li, Yuning Gong, Hongjian Zhan et al.

2026-05-04 11
cs.AI 2605.02592

Foundation-Model-Based Agents in Industrial Automation: Purposes, Capabilities, and Open Challenges

Foundation-model-based industrial agents leverage large language models for autonomous decision-making, enhancing human interaction (+37%) and uncertainty handling (+35%).

Vincent Henkel, Felix Gehlhoff, David Kube et al.

2026-05-04 48
cs.CV 2605.02580

Hyp2Former: Hierarchy-Aware Hyperbolic Embeddings for Open-Set Panoptic Segmentation

Hyp2Former employs hyperbolic hierarchical embeddings to improve open-set panoptic segmentation, achieving state-of-the-art results.

Yao Lu, Rohit Mohan, Florian Drews et al.

2026-05-04 46
cs.RO 2605.02525

A Semantic Autonomy Framework for VLM-Integrated Indoor Mobile Robots: Hybrid Deterministic Reasoning and Cross-Robot Adaptive Memory

Proposes a six-layer semantic autonomy framework combining hybrid deterministic-VLM reasoning and cross-robot memory transfer, enabling efficient indoor robot navigation.

Bogdan Felician Abaza, Andrei-Alexandru Staicu, Cristian Vasile Doicin

2026-05-04 65
stat.ML 2605.02462

Black-box optimization of noisy functions with unknown smoothness

POO algorithm optimizes noisy functions with unknown smoothness, error within √ln n of best algorithms.

Jean-Bastien Grill, Michal Valko, Rémi Munos

2026-05-04 4
cs.CL 2605.02364

InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and Repetition

Introduces InfoLaw, a data-aware scaling law using information theory to predict large language model performance with quality-weighted data and repetition, achieving 0.15% loss prediction error.

Fengze Liu, Weidong Zhou, Binbin Liu et al.

2026-05-04 44
cs.CV 2605.02212

NTIRE 2026 Challenge on Efficient Low Light Image Enhancement: Methods and Results

NTIRE 2026 Challenge introduces RetinexFormerRefine, achieving SSIM of 0.5654 with models under 1MB for low-light image enhancement.

Jiebin Yan, Chenyu Tu, Weixia Zhang et al.

2026-05-04 29
cs.AI 2605.02168

Planner Matters! An Efficient and Unbalanced Multi-agent Collaboration Framework for Long-horizon Planning

Proposed a planner-centric multi-agent framework, significantly improving long-horizon task success rates.

Wenyi Wu, Sibo Zhu, Kun Zhou et al.

2026-05-04 9
stat.ML 2605.02014

MIRA: A Score for Conditional Distribution Accuracy and Model Comparison

MIRA scores condition distribution accuracy via sample regions, enabling scalable Bayesian model validation and comparison.

Sammy Sharief, Justine Zeghal, Gabriel Missael Barco et al.

2026-05-04 67
cs.SD 2605.01809

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

TMD-Bench introduces a multi-level evaluation framework for music-dance co-generation, showcasing RhyJAM's competitive rhythmic synchronization.

Xiaoda Yang, Majun Zhang, Changhao Pan et al.

2026-05-03 35
cs.CL 2605.01749

Only Say What You Know: Calibration-Aware Generation for Long-Form Factuality

Decoupling Exploration and Commitment with Calibration-Aware Generation boosts factuality by up to 13% in long-form outputs.

Wen Luo, Guangyue Peng, Liang Wang et al.

2026-05-03 55
cs.RO 2605.01518

VOFA: Visual Object Goal Pushing with Force-Adaptive Control for Humanoids

VOFA integrates vision-based policies with force-adaptive control, achieving over 90% success in pushing unknown objects up to 17kg in real and simulated environments.

Zichao Hu, Zifan Xu, Dongsik Chang et al.

2026-05-03 53
Prev 1 ... 150 151 152 153 154 155 156 ... 573 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home