Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics
Proposes distributional divergence metrics (e.g., MAUVE) to address gen-PPL's failure in evaluating text quality.
Antonio Franca, Alexander Tong
Proposes distributional divergence metrics (e.g., MAUVE) to address gen-PPL's failure in evaluating text quality.
Antonio Franca, Alexander Tong
TrustMargin is a training-free arbitration method combining parametric priors and evidence support, improving answer selection in large language models by 4-6 points F1.
Jingyan Xu, Hong Shi, Yi Shan et al.
CATPO uses tree informativeness scoring and critique-guided repair to improve mathematical reasoning accuracy to 37.5%.
Ayush Singh, Umang Goyal, Ankur Dahiya
PoE-Bridge enables parallel decoding in diffusion language models, achieving 5x speedup and recovering 95% performance.
Juntong Shi, Brian L. Trippe, Jure Leskovec et al.
Agentopia introduces a long-term multi-agent society simulation over 10 years, leveraging life reward-based reinforcement learning to enhance social behaviors and anthropomorphic capabilities of LLMs.
Xintao Wang, Sirui Zheng, Hongqiu Wu et al.
Proposes EmbedFilter, a linear transformation that filters the latent subspace encoding high-frequency, uninformative tokens, improving zero-shot text embedding performance by up to 14%.
Songhao Wu, Zhongxin Chen, Yuxuan Liu et al.
This paper introduces LLM-guided MAP-Elites evolution for optimizing medical decision pipelines, improving accuracy and safety metrics significantly across tasks.
Ivan Sviridov, Artem Oskin, Ivan Panin et al.
FinEvolveBench tests delayed-feedback self-evolution on 31 Chinese A-share industries and 177,324 news articles; generic memory rarely improves IC.
Zihao Deng, Yining Zhu, Leiming Wang et al.
CRAFT introduces bidirectional counterfactual reasoning, achieving 82.4% accuracy on WikiTQ and 94.6% on TabFact, outperforming baselines.
Chenshuo Pan, Yu Zhao, Jie Zhang et al.
Proposes emergent language in multi-agent RL to study consciousness-related structures without prior biases, revealing self-referential communication and echo-mismatch circuits.
Zengqing Wu, Chuan Xiao
Introduced a benchmark for data snapshot detection, evaluated open-source models, revealing significant gaps in real-world institutional document understanding.
AJ Carl P. Dy, Aivin V. Solatorio
IR3DE uses ridge regression for fast, low-cost model routing, achieving 98.4% performance in multi-domain LLM selection.
Eros Fanì, Oğuzhan Ersoy
MARDoc employs structured memory to improve multimodal long document QA, achieving 57.1% accuracy and outperforming baselines.
Kaifeng Chen, Hongtao Liu, Qiyao Peng et al.
AdaPLD achieves efficient decoding with adaptive retrieval and reuse, boosting speed by 3.10×.
Runheng Liu, Jincheng Xie, Wen Hu et al.
Introduced α-STOP for early failure alerting in dialogues and LLM-agent trajectories, improving frontier quality by 1-42%.
Avinash Baidya, Xinran Liang, Ruocheng Guo et al.
Introduced PERSUASIONTRACE, leveraging Bayesian networks to model multi-turn belief updates, scoring near humans (81 vs 80).
Jared Moore, Noah Goodman, Nick Haber et al.
Depth-Attention enhances language models by cross-layer value mixing, improving accuracy by up to 2.3 points.
Boyi Zeng, Yiqin Hao, Zitong Wang et al.
LifeSide uses multi-session Memory-Emotion-Environment loops to evaluate lifelong digital companions' long-term memory, understanding, privacy, and emotional support, based on 2000 personas and 111K tasks.
Yuqian Wu, Zhijie Deng, Wei Chen et al.
RAMPART optimizes LLM agent memory with priority-aware runtime transformation, enhancing task success.
Nikodem Tomczak
VCIFBench evaluates complex video instruction following, covering 306 tasks with multi-constraint challenges, revealing significant gaps in current models.
Huangchen Xu, Yuan Wu, Yi Chang