FinHarness: An Inline Lifecycle Safety Harness for Finance LLM Agents
FinHarness reduces ASR to 15% on FinVault with 4.7× fewer advanced judge calls via inline lifecycle safety harness
Haoxuan Jia, Yang Liu, Bin Chong et al.
FinHarness reduces ASR to 15% on FinVault with 4.7× fewer advanced judge calls via inline lifecycle safety harness
Haoxuan Jia, Yang Liu, Bin Chong et al.
Introduces ENPMR-Bench, a benchmark for proactive memory retrieval based on Maslow’s hierarchy, revealing significant deficiencies with current models, and outlining future directions.
Xing Fu, Yulin Hu, Mengtong Ji et al.
BhashaSetu improves low-resource translation quality via corpus deduplication, reducing performance by 1.17 BLEU and 2.21 chrF++.
Param Thakkar, Anushka Yadav, Michael Tiemann et al.
SurfaceLoRA method reduces sensitive information inference risk in LLM-exported vectors while maintaining summarization performance.
Weixin Liu, Bowen Qu, Juming Xiong et al.
Study reveals algorithmic fragility and persona bias in LLM-generated autistic communication using a multi-agent qualitative analysis framework.
Naba Rizvi, Mohammed Rizvi, Harper Strickland et al.
StakeBench evaluates language understanding via market commitment, linking 560,876 comments to four diagnostic tasks.
Yunhua Pei, Jingyu Hu, Yiwei Shi et al.
Proposes MixT, a Hamiltonian-inspired local operator approach, achieving over 50% parameter reduction in billion-parameter LLMs.
Ying Lu, Peng-Fei Zhou, Qi-Xuan Fang et al.
Mimir is a 1.6B parameter multilingual concept model trained on 388 billion sentences, shifting from token to concept-level understanding.
Elio Musacchio, Lucia Siciliani, Pierpaolo Basile
OpenSkillEval introduces a dynamic evaluation framework using over 600 generated tasks and 30 open-source skills, revealing that skill availability does not guarantee effective utilization or performance gains.
Jiahao Ying, Boxian Ai, Wei Tang et al.
ConvexTok uses convex relaxations to optimize tokenisation, achieving near-optimal compression within 1% at 128k vocabulary, improving BpB significantly.
Jan Tempus, Philip Whittington, Craig W. Schmidt et al.
Evaluated six commercial AI chatbots on 2,100 BBC news questions across six languages, achieving up to 95.6% accuracy on emerging facts.
Mirac Suzgun, Emily Shen, Federico Bianchi et al.
6B-parameter LLMs pretrained sequentially on Common Crawl show 15% F1 improvement on KairosQA for temporal knowledge over shuffled baselines.
Pilchen Hippolyte, Fabre Romain, Signe Talla Franck et al.
Self-Policy Distillation (SPD) extracts low-rank capability subspaces via gradients, guiding self-generation without external signals, improving performance by 13%.
Guangya Hao, Yitong Shang, Yunbo Long et al.
SynAE employs multi-metric evaluation to assess synthetic data quality for tool-calling agents, focusing on validity, fidelity, and diversity, enabling robust benchmarking.
Shuaiqi Wang, Aadyaa Maddi, Zinan Lin et al.
RankJudge generates multi-turn dialogue benchmarks, using the Bradley-Terry model to evaluate 21 LLM judges.
Zhenwei Tang, Zhaoyan Liu, Rasa Hosseinzadeh et al.
Auto-Dreamer employs offline memory consolidation via region rewriting, achieving 7% higher success and 12× smaller memory on ScienceWorld.
Chongrui Ye, Yuxiang Liu, Yu Wang et al.
ThoughtTrace collects user-reported thoughts during multi-turn conversations, improving user behavior prediction (+41.7%) and model alignment (+25.6%) in large-scale datasets.
Chuanyang Jin, Binze Li, Haopeng Xie et al.
SkillsVote manages open-source agent skills through lifecycle stages, improving task performance by 2.6% on benchmarks.
Hongyi Liu, Haoyan Yang, Tao Jiang et al.
EPIC constructs preference-aligned compact index, reducing memory by 2404× and improving accuracy by 18.79%, enabling efficient on-device personalization.
Changmin Lee, Jaemin Kim, Taesik Gong
DRIFT and SAPLMA effectively detect hallucinations after removing benchmark artifacts, achieving AUROC of 0.91.
Khizar Hussain, Murat Kantarcioglu