cs.CL 2501.13956

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Zep employs Graphiti's time-aware knowledge graph, dynamically integrating conversational and business data, achieving 94.8% accuracy and 90% latency reduction, outperforming MemGPT.

Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais et al.

2025-01-21 39
cs.CL 2501.04686

Unlocking Multimodal Mathematical Reasoning via Process Reward Model

Introduces URSA framework with large-scale datasets MMathCoT-1M and DualMath-1.1M, combined with PS-GRPO algorithm, significantly enhancing multimodal mathematical reasoning, outperforming GPT-4o by 8.4%.

Ruilin Luo, Zhuofan Zheng, Yifan Wang et al.

2025-01-09 36 citations 33
cs.CL 2501.06246

A Partition Cover Approach to Tokenization

Partition cover-based tokenization algorithm GREEDTOK outperforms BPE and Unigram, achieving ~3% better compression on real-world corpora and lower bits per byte in large-scale pretraining.

Jia Peng Lim, Shawn Tan, Davin Choo et al.

2025-01-09 6 citations 40
cs.CL 2501.00273

Echoes in AI: Quantifying lack of plot diversity in LLM outputs

Introduces Sui Generis score to quantify plot diversity in LLM stories; experiments with GPT-4 and LLaMA-3 show high plot element repetition, indicating lack of creativity.

Weijia Xu, Nebojsa Jojic, Sudha Rao et al.

2024-12-31 60 citations 48
cs.CL 2412.13788

Open Universal Arabic ASR Leaderboard

Introduced Open Universal Arabic ASR Leaderboard to evaluate multi-dialect models and reveal generalization capabilities.

Yingzhi Wang, Anas Alhmoud, Muhammad Alqurishi

2024-12-18 5
cs.CL 2412.13717

Towards Automatic Evaluation for Image Transcreation

Proposed object, embedding, and VLM-based metrics achieve 0.55-0.87 correlation with human ratings for automatic image transcreation evaluation.

Simran Khanuja, Vivek Iyer, Claire He et al.

2024-12-18 6 citations 37