Embedding Models Measure in Peculiar Ways
Study finds embedding models weakly reflect physical measurements, influenced by superficial string similarity.
Juri Opitz, Andrianos Michail
Study finds embedding models weakly reflect physical measurements, influenced by superficial string similarity.
Juri Opitz, Andrianos Michail
JEPA-Anything uses Orthogonal Predictive Factorization (OPF) to enhance cross-domain predictive models, improving prediction accuracy by 34.8%.
Taoyong Cui, Zhongyao Wang, Xinyue Xu et al.
RetireOPD method improves RL model success rate by 18.8% in ALFWorld using adaptive retirement strategy.
Yan Yu, Zhengxi Lu, Yizhou Liu et al.
Study shows gender bias in GPT models is transformed, not reduced, introducing 'harm laundering'.
Sarah Wyer, Sue Black, Noura Al Moubayed
dQwen3.5 uses hybrid attention to achieve training loss in half the tokens compared to full-attention models.
Anton Xue, Litu Rout, Aditya Akella et al.
Introduces On-Demand Attention (ODA) method, significantly reducing global reads and enhancing long-context inference efficiency.
Haibo Feng, Ruiqi Liang, Hanyang Peng et al.
Study evaluates sequence-likelihood signal in retrieval-dominated QA, finding it insufficient to enhance accuracy.
Gunwoo Lee, Changmin Sung, Sang-Hwan Gwak et al.
EB-Decode accelerates dLLM inference with adaptive block sizes and parallel sampling, achieving 3.53-18.76× throughput improvement.
Lixuan Wei, Wei Zhou, Jianwen Wu et al.
The study examines truncation sampling's impact on large language models' accuracy at high temperatures.
Francesco La Rosa
Using semantic uncertainty to predict transition relevance places (TRPs) significantly outperforms baseline methods.
Muhammad Umair, Jan P. de Ruiter
IBIB protocol evaluates enterprise AI systems via serving routes, not model identifiers, improving reliability measurement.
Blake Stenstrom, Charangan Vasantharajan, Brian Sathianathan
SocialRL combines multi-turn PPO with dynamic process rewards, raising goal achievement by an average of 9.2 percentage points.
Jianing Wang, Xintao Wang, Aili Chen et al.
CantoneseLLM v2 combines CPT, DPO, and RLVR, reaching 73.16 on HKCanto-Eval with Cantonese Traditional-Chinese reasoning.
Tsz Chung Cheng, Chung Shing Cheng, Chaak Ming Lau et al.
Proposes DH-BPE combining minimum-token segmentation exposure with hierarchical BPE for fixed vocab optimization.
Kenny Shao
WearableQA benchmark evaluates AI's health reasoning on real wearable data, with performance ranging from 19.6% to 72.9%.
Ji Soo Lee, Xilun Chen, Pierce Chuang et al.
LexFlip uses lexical perturbations to detect legal meaning changes; NLI performs best.
Gaurab Baral
Introduces DualCWE, a self-supervised framework learning lexical representations from IPA lists, enabling rapid phylogenetic inference for 3,399 languages.
Tim Wientzek
Improves reasoning depth to 72.20% using Gold-Anchored QLoRA and RLVR.
Thi Kim Trang Vo, Nam Tien Le, Thi Kim Nguyet Vo et al.
FlexPension-LLM predicts pension enrollment among China's flexible workers, achieving 0.9316 F1.
Yumiao Li, Peixin Liu, Donglin Di et al.
TeQHallu uses low-level symbolic competence for unsupervised hallucination detection, excelling on RAGTruth dataset.
Renato Vukovic, Hsien-chin Lin, Carel van Niekerk et al.