High-Dimension Human Value Representation in Large Language Models
Introduced UniVaR, a high-dimensional representation of human values in LLMs across 25 languages and cultures.
Samuel Cahyawijaya, Delong Chen, Yejin Bang et al.
Introduced UniVaR, a high-dimensional representation of human values in LLMs across 25 languages and cultures.
Samuel Cahyawijaya, Delong Chen, Yejin Bang et al.
ResearchAgent leverages LLMs, academic graphs, and entity knowledge bases to automatically generate and iteratively refine novel scientific research ideas, outperforming baseline models.
Jinheon Baek, Sujay Kumar Jauhar, Silviu Cucerzan et al.
RULER benchmark evaluates long-context models across multiple tasks, revealing performance drops at extreme lengths.
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman et al.
Khayyam Challenge evaluates Persian LLMs using 20,192 multiple-choice questions.
Omid Ghahroodi, Marzia Nouri, Mohammad Vali Sanian et al.
MiniCPM showcases small language models' potential with scalable training strategies, rivaling 7B-13B models.
Shengding Hu, Yuge Tu, Xu Han et al.
AutoWebGLM leverages HTML simplification and reinforcement learning to create a web navigation agent surpassing GPT-4 in performance.
Hanyu Lai, Xiao Liu, Iat Long Iong et al.
Proposes a multi-level irrelevant information framework using Wikipedia, evaluates LLM robustness; finds models are easily misled by semantically related distractors.
Siye Wu, Jian Xie, Jiangjie Chen et al.
Using Liang et al.'s (2024) distributional GPT framework, analyzed 950,965 papers (2020-2024), revealing rapid growth of LLM-modified content, especially in CS (up to 17.5%).
Weixin Liang, Yaohui Zhang, Zhengxuan Wu et al.
Proposes RQ-RAG, which refines queries via rewriting, decomposition, and disambiguation, boosting 7B Llama2 performance on single/multi-hop QA, surpassing SOTA.
Chi-Min Chan, Chunpu Xu, Ruibin Yuan et al.
Gecko employs a two-step distillation from LLMs, achieving 66.31 on MTEB with only 256 dimensions, outperforming larger models.
Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren et al.
LlamaFactory unifies 100+ models for efficient fine-tuning, enabling codeless customization.
Yaowei Zheng, Richong Zhang, Junhao Zhang et al.
Introduces the concept of 'effective cutoff' by probing perplexity across data versions, revealing discrepancies between reported and actual knowledge cutoffs in LLMs.
Jeffrey Cheng, Marc Marone, Orion Weller et al.
SHROOM detects hallucinations in NLG outputs using a 4000-sample dataset, leveraging fine-tuning and zero-shot prompting strategies.
Timothee Mickus, Elaine Zosa, Raúl Vázquez et al.
Proposes Claim-Conditioned Probability (CCP) for token-level uncertainty quantification, improving factual accuracy detection across 7 models and 4 languages.
Ekaterina Fadeeva, Aleksandr Rubashevskii, Artem Shelmanov et al.
Design2Code benchmarks multimodal models on real webpages, revealing gaps in visual element recall and layout accuracy, with GPT-4o leading but still limited.
Chenglei Si, Yanzhe Zhang, Ryan Li et al.
MathScale uses topic extraction and concept graphs to generate 2 million math QA pairs, boosting LLMs' reasoning by 43%.
Zhengyang Tang, Xingxing Zhang, Benyou Wang et al.
Proposes exploration-based trajectory optimization (ETO) for LLM agents, leveraging failure trajectories with contrastive learning, outperforming baselines by significant margins.
Yifan Song, Da Yin, Xiang Yue et al.
Proposed Key-Point-Driven Data Synthesis (KPDDS), creating 800K+ math reasoning QA pairs, significantly boosting model performance.
Yiming Huang, Xiao Liu, Yeyun Gong et al.
This study systematically compares multiple scoring methods for LLMs in multiple-choice tasks, revealing high sensitivity and variability across methods and models.
Polina Tsvilodub, Hening Wang, Sharon Grosch et al.
CLLMs improve Jacobi decoding to achieve 2.4x to 3.4x speedup while maintaining quality.
Siqi Kou, Lanxiang Hu, Zhezhi He et al.