GptGet
Features PaperForge Apps Papers Blog Contact AI Chat 中文
Sort: Latest Popular Citations
All Artificial Intelligence Computation and Language Computer Vision Information Retrieval Machine Learning Machine Learning (Stats) Neural and Evolutionary Computing Robotics
cs.SE 2504.14757

SWE-Synth: Synthesizing Verifiable Bug-Fix Data to Enable Large Language Models in Resolving Real-World Bugs

SWE-Synth uses LLMs to simulate debugging workflows, synthesizing verifiable bug-fix data, improving open-source model training by 2.3%.

Minh V. T. Pham, Huy N. Phan, Hoang N. Phan et al.

2025-04-21 42
cs.CL 2504.14692

OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding

OmniV-Med unifies 2D/3D images and videos in a medical vision-language model, using rotary position encoding and token pruning to achieve SOTA performance across 7 benchmarks.

Songtao Jiang, Yuan Wang, Sibo Song et al.

2025-04-21 39
cs.RO 2504.14604

RoboOcc: Enhancing the Geometric and Semantic Scene Understanding for Robots

RoboOcc integrates Opacity-guided Self-Encoder and Geometry-aware Cross-Encoder to improve 3D scene understanding, outperforming SOTA with 8.47% IoU gain.

Zhang Zhang, Qiang Zhang, Wei Cui et al.

2025-04-20 54
cs.NI 2504.14411

Planet as a Brain: Towards Internet of AgentSites based on AIOS Server

Proposes AIOS server for decentralized AgentSite network, enabling scalable peer-to-peer communication and discovery.

Xiang Zhang, Yongfeng Zhang

2025-04-20 57
cs.AI 2504.13837

Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

This study critically examines RLVR's ability to induce novel reasoning in LLMs, revealing paths are bounded by base models, with distillation offering better expansion.

Yang Yue, Zhiqi Chen, Rui Lu et al.

2025-04-19 1027 citations 37
cs.CV 2504.13820

CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning

CheXWorld employs a self-supervised world model with multi-task learning, capturing local structures, global layouts, and domain shifts, achieving state-of-the-art results on 8 benchmarks.

Yang Yue, Yulin Wang, Chenxin Tao et al.

2025-04-19 51
cs.SE 2504.13472

CodeVisionary: An Agent-based Framework for Evaluating Large Language Models in Code Generation

CodeVisionary employs a two-stage multi-agent framework combining requirement-driven context distillation and collaborative scoring for complex code evaluation.

Xinchen Wang, Pengfei Gao, Chao Peng et al.

2025-04-18 29
cs.CV 2504.13180

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

PerceptionLM employs multi-stage training with 2.8M human-annotated video QA and spatio-temporal captions, advancing transparent vision-language modeling.

Jang Hyun Cho, Andrea Madotto, Effrosyni Mavroudi et al.

2025-04-18 92 citations 38
cs.LG 2504.13173

It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization

Miras framework unifies sequence models via associative memory, diverse attentional biases, and retention gates, outperforming some SOTA models.

Ali Behrouz, Meisam Razaviyayn, Peilin Zhong et al.

2025-04-18 31
cs.AI 2504.13171

Sleep-time Compute: Beyond Inference Scaling at Test-time

Introduces sleep-time compute to reduce test-time compute requirements by pre-computing context, improving accuracy.

Kevin Lin, Charlie Snell, Yu Wang et al.

2025-04-18 6
cs.LG 2504.16109

Representation Learning for Tabular Data: A Comprehensive Survey

Deep neural networks excel in representation learning for tabular data, enhancing accuracy in classification and regression tasks.

Jun-Peng Jiang, Si-Yang Liu, Hao-Run Cai et al.

2025-04-18 21
cs.CL 2504.13161

Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training

Nemotron-CLIMB employs semantic embedding and iterative search to optimize data mixtures, boosting language model pre-training performance.

Shizhe Diao, Yu Yang, Yonggan Fu et al.

2025-04-18 35
cs.CV 2504.13109

UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

Proposes Uni-Inv and Uni-Edit for accurate inversion and region-aware editing in flow models, achieving low error rates and natural edits.

Guanlong Jiao, Biqing Huang, Kuan-Chieh Wang et al.

2025-04-18 37
cs.CY 2504.12914

In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?

This study assesses international AI safety cooperation risks, highlighting verification mechanisms and shared protocols as promising areas for secure collaboration.

Ben Bucknall, Saad Siddiqui, Lara Thurnherr et al.

2025-04-17 35
cs.CL 2504.12845

Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks

MLRBench evaluates LLMs' reasoning over multilingual long contexts, revealing resource gaps.

Amey Hengle, Prasoon Bajpai, Soham Dan et al.

2025-04-17 28
cs.CL 2504.12663

Persona-judge: Personalized Alignment of Large Language Models via Token-level Self-judgment

Persona-judge achieves personalized alignment via self-judgment, enhancing alignment efficiency by 98%.

Xiaotian Zhang, Ruizhe Chen, Yang Feng et al.

2025-04-17 16
cs.CV 2504.12574

ForgetMe: Evaluating Selective Forgetting in Generative Models

Proposes ForgetMe dataset and Entangled metric for selective forgetting in diffusion models, using prompt-layered editing and training-free local feature removal.

Zhenyu Yu, Mohd Yamani Inda Idris, Pei Wang

2025-04-17 43
cs.CL 2504.12522

Evaluating the Diversity and Quality of LLM Generated Content

Proposes an effective semantic diversity metric for LLM outputs, revealing preference tuning enhances high-quality content diversity.

Alexander Shypula, Shuo Li, Botong Zhang et al.

2025-04-17 57 citations 45
cs.CL 2504.12516

BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

BrowseComp is a challenging benchmark for AI web browsing, comprising 1266 questions requiring persistent, creative search for hard-to-find information.

Jason Wei, Zhiqing Sun, Spencer Papay et al.

2025-04-17 42
cs.LG 2504.12501

Reinforcement Learning from Human Feedback

RLHF uses preference models and PPO to align language models with human preferences, improving safety and style consistency.

Nathan Lambert

2025-04-17 56
Prev 1 ... 291 292 293 294 295 296 297 ... 573 Next

© 2026 GptGet.net - Paper Insights Platform

Paper List Submit Paper Help GptGet Home