Gender and Race Bias in Consumer Product Recommendations by Large Language Models
Combines Prompt engineering, Marked Words, SVM, and JSD to detect biases in LLM-generated recommendations.
Ke Xu, Shera Potka, Alex Thomo
Combines Prompt engineering, Marked Words, SVM, and JSD to detect biases in LLM-generated recommendations.
Ke Xu, Shera Potka, Alex Thomo
Proposes PlugMem, a task-agnostic plugin memory module using knowledge graphs, outperforming task-specific methods with higher information density.
Ke Yang, Zixi Chen, Xuan He et al.
LinGO uses linguistic graph optimization with LLMs to interpret online uncivil discourse, improving accuracy and F1 scores.
Yuan Zhang, Thales Bertaglia
RexBERT excels in e-commerce with Ecom-niverse corpus and three-phase training.
Rahul Bajaj, Anuj Garg
ImpRIF formalizes implicit reasoning as verifiable graphs, boosting instruction following by over 10% on benchmarks.
Yuancheng Yang, Lin Yang, Xu Wang et al.
Proxy compression enhances language model efficiency, significantly outperforming byte-level baselines.
Lin Zheng, Xinyu Li, Qian Liu et al.
Introduces SalamaBench, a benchmark using MLCommons taxonomy, evaluating 12 safety categories across 8170 prompts for Arabic models.
Omar Abdelnasser, Fatemah Alharbi, Khaled Khasawneh et al.
InfMem enhances long-context reasoning accuracy by 10.17 points using the PreThink-Retrieve-Write protocol.
Xinyu Wang, Mingze Li, Peng Lu et al.
Reinforcement learning-based divide-and-conquer (DAC) training boosts LLM reasoning scalability, surpassing chain-of-thought (CoT) by 8.6%.
Xiao Liang, Zhong-Zhi Li, Zhenghao Lin et al.
Proposed RAG-based DziriBOT with DziriBERT achieves 88% accuracy on Algerian dialect understanding.
El Batoul Bechiri, Dihia Lanasri
The paper reveals refusal behaviors in LLMs are controlled by multiple geometrically distinct directions, yet linear steering along any yields similar refusal effects.
Faaiz Joad, Majd Hawasly, Sabri Boughorbel et al.
Proposes xMemory, a hierarchical decoupling and aggregation method that improves agent memory retrieval, boosting answer quality and reducing redundancy by 15-20%.
Zhanghao Hu, Qinglin Zhu, Runcong Zhao et al.
Proposes Rubric-ARM, an alternating RL framework jointly optimizing rubric generator and judge, achieving 4.7% improvement in reward modeling accuracy.
Ran Xu, Tianci Liu, Zihan Dong et al.
DeALOG employs a decentralized multi-agent framework with shared natural language logs, achieving robustness and transparency in multimodal QA.
Abhijit Chakraborty, Ashish Raj Shekhar, Shiven Agarwal et al.
Study finds self-evolving LLMs rely on raw experience but struggle to utilize condensed experience effectively.
Weixiang Zhao, Yingshuo Wang, Yichen Zhang et al.
Proposes SUSTAINSCORE to quantify instruction-induced task performance drops; experiments show performance declines of 15-35%.
Yunjia Qi, Hao Peng, Xintong Shi et al.
Leviathan replaces input embedding with a learned vectorization layer, improving language modeling perplexity by 9% at 1.2B parameters, especially aiding rare words.
Reza T. Batley, Sourav Saha
OVD: trajectory matching with verbal scores reduces memory, improves Web QA and math reasoning by up to 25.7%.
Jing Xiong, Hui Shen, Shansan Gong et al.
Introduces Temporal Guidance (TeGu), leveraging temporal contrast to enhance LLM generation quality, achieving a 3.03% improvement on GSM8K.
Hong-Kai Zheng, Piji Li
This paper introduces Persona Prompting (PP) to analyze its impact on social reasoning in LLMs, focusing on bias, rationale quality, and task performance using hate speech datasets.
Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov et al.