Geometric Uncertainty for Detecting and Correcting Hallucinations in LLMs
Proposes geometric volume and suspicion metrics for improved hallucination detection and correction in LLMs.
Edward Phillips, Sean Wu, Soheila Molaei et al.
Proposes geometric volume and suspicion metrics for improved hallucination detection and correction in LLMs.
Edward Phillips, Sean Wu, Soheila Molaei et al.
WebWeaver employs dynamic outlines and dual agents, significantly improving open-ended deep research accuracy and efficiency.
Zijian Li, Xin Guan, Bo Zhang et al.
Proposes Agentic CPT framework, leveraging large-scale multi-task data to enhance foundational models' agentic abilities, achieving 39.9% on BrowseComp-en.
Liangcai Su, Zhen Zhang, Guangyu Li et al.
WebResearcher employs iterative research and WebFrontier data engine to achieve state-of-the-art long-horizon reasoning, surpassing 6 benchmarks with significant improvements.
Zile Qiao, Guoxin Chen, Xuanzhong Chen et al.
Conan-Embedding-v2 is a 1.4B parameter from-scratch trained model, using soft masks and dynamic hard negative mining to achieve SOTA text embeddings.
Shiyu Li, Yang Tang, Ruijie Liu et al.
Using linear discriminant analysis (LDA), the study identifies orthogonal directions encoding tense and aspect in LLM residual space, enabling causal control during multi-token generation.
Alina Klerings, Jannik Brinkmann, Daniel Ruffinelli et al.
AgenticIE system extracts information from complex regulatory documents, achieving a 16% EM score improvement.
Gaye Colakoglu, Gürkan Solmaz, Jonathan Fürst
Proposes a framework for LLM-based social simulation using importance sampling and optimal transport to generate population-aligned persona sets.
Zhengyu Hu, Jianxun Lian, Zheyuan Xiao et al.
FHIR-AgentBench evaluates LLM performance in EHR QA under HL7 FHIR, highlighting data retrieval and reasoning challenges.
Gyubok Lee, Elea Bach, Eric Yang et al.
Proposes CAI ratio for unsupervised LLM evaluation, effectively guiding model selection in dynamic environments.
Cheng Chen, Haiyan Yin, Ivor Tsang
DiverValue-Bench evaluates LLMs' alignment with diverse human values, revealing geographic and demographic disparities.
Yao Liang, Dongcheng Zhao, Feifei Zhao et al.
PersonaFuse framework enhances human-LLM interactions by activating personality traits, significantly improving emotional intelligence.
Yixuan Tang, Yi Yang, Ahmed Abbasi
This study analyzes failure modes in multi-agent debate, revealing models tend to conform and reinforce errors, causing performance decline.
Andrea Wynn, Harsh Satija, Gillian Hadfield
LLMs combined pairwise comparison and fine-tuning improve scalar construct measurement in social science, surpassing direct scoring.
Hauke Licht, Rupak Sarkar, Patrick Y. Wu et al.
DARLING framework integrates a semantic classifier into reinforcement learning to jointly optimize language model diversity and quality, significantly improving creative and mathematical tasks.
Tianjian Li, Yiming Zhang, Ping Yu et al.
ConspirED dataset reveals cognitive traits of conspiracy theories, assesses LLM safety.
Luke Bates, Max Glockner, Preslav Nakov et al.
DeepScholar-bench is a real-time benchmark for generative research synthesis, evaluating knowledge integration, retrieval relevance, and citation verifiability.
Liana Patel, Negar Arabzadeh, Harshit Gupta et al.
Introduces T2R-bench, a benchmark for industrial table-to-report tasks, with models scoring only 62.71%, highlighting room for improvement.
Jie Zhang, Changzai Pan, Kaiwen Wei et al.
This study evaluates LVLMs' knowledge boundary perception using probabilistic, answer consistency, and verbal confidence signals, revealing significant room for improvement.
Zhikai Ding, Shiyu Ni, Keping Bi
Pandora uses Python Pandas for unified structured knowledge reasoning, enhancing cross-task performance.
Yongrui Chen, Junhao He, Linbo Fu et al.