The Limits of Automatic Evaluation of Creativity in Large Language Models

TL;DR

This study critically evaluates automatic creativity metrics and LLM-based judgments, revealing weak correlation with human assessments and systemic biases, highlighting evaluation limitations.

cs.CL 🔴 Advanced 2026-08-25 50 views
Alessandro Tutone Giorgio Franceschelli Mirco Musolesi
AI creativity assessment NLP automatic metrics large language models

Key Findings

Methodology

This work combines human ratings and automated metrics on 200 short stories from WritingPrompts, with 100 human and 100 AI-generated texts. Eleven creativity dimensions were rated via surveys, while five automatic metrics (Perplexity, EAD, SBERT-Div, Creativity Index, template scores) and LLM judgments were computed. Correlation analyses (Pearson, Kendall) assessed alignment. Results show poor correlation between automatic metrics and human ratings, with LLM judgments biased toward AI texts, favoring stylistic features over semantic depth.

Key Results

  • Automated metrics exhibit near-zero correlation with human ratings (ρ<0.2), with Creativity Index showing almost no relation (ρ≈0.07). LLM judgments favor AI texts, scoring them systematically higher, especially on Effectiveness, Elaboration, and other subjective dimensions.
  • Finer analysis reveals LLM bias: scores for AI texts cluster at maximum, while human scores vary considerably. Correlations between human and LLM evaluations are weak (τ≈0.23 for Creativity), indicating limited reliability of LLM as a judge.
  • Dimension correlation analysis shows LLM evaluations align mainly with surface features like Elaboration, but poorly with core creative qualities such as Surprise or Usefulness, emphasizing superficial assessment over semantic understanding.

Significance

The findings expose fundamental flaws in current automatic evaluation methods for creativity, emphasizing their inability to capture subjective, multi-faceted qualities. This limits their application in AI-generated content assessment, urging the development of more nuanced, human-aligned metrics. The bias in LLM judgments also raises concerns about over-reliance on models trained on stylistic patterns, which may distort true creative evaluation. These insights are vital for advancing AI's role in creative industries and content quality assurance.

Technical Contribution

This research systematically compares multiple automatic metrics and LLM judgments against human evaluations, revealing their weak correlation and biases. It highlights the necessity of multi-dimensional, semantic-aware evaluation frameworks. The study introduces a comprehensive statistical analysis of biases, providing a foundation for future improvements in automated creativity assessment, and underscores the importance of integrating human feedback for more reliable evaluation systems.

Novelty

This is the first large-scale, systematic comparison of automatic and human creativity evaluations across multiple dimensions, demonstrating the inadequacy of existing metrics. It also critically examines LLMs as evaluators, exposing their bias toward stylistic features, a novel insight that challenges the assumption of LLMs as reliable judges of creativity.

Limitations

  • Sample size limited to short stories, which may not generalize to other creative forms like poetry or novels. Future work should include diverse genres.
  • Metrics focus on surface-level features, lacking deep semantic or emotional assessment. Incorporating richer semantic understanding remains a challenge.
  • LLM bias may stem from training data; future research should explore bias correction and multi-model consensus to improve reliability.

Future Work

Future directions include developing multi-modal, multi-dimensional evaluation frameworks that incorporate semantic, emotional, and contextual cues. Combining human-in-the-loop systems with automated metrics could enhance reliability. Cross-cultural and cross-genre studies are needed to generalize findings, and more robust bias mitigation strategies should be explored to ensure fair and accurate creativity assessment.

AI Executive Summary

Recent advances in artificial intelligence, especially large language models like GPT-4, have revolutionized text generation, producing outputs that often rival human creativity. However, assessing the true creative quality of these outputs remains a significant challenge. Traditional automatic metrics such as Perplexity, EAD, and SBERT-Div, though computationally efficient, have shown limited correlation with human judgments, often failing to capture the nuanced qualities that define creativity—novelty, surprise, and value.

This study undertook a comprehensive evaluation using a dataset of 200 short stories—half human-authored, half AI-generated—assessing 11 dimensions of creativity through human surveys, automatic metrics, and LLM-based judgments. The results revealed a stark disconnect: automatic metrics exhibited near-zero correlation with human ratings, and LLM judges systematically favored AI texts, scoring them higher across multiple dimensions. Notably, the Creativity Index, designed to quantify innovation, correlated weakly (ρ≈0.07), underscoring the difficulty of formalizing creativity.

The bias of LLMs toward AI-generated content suggests that current models may rely heavily on stylistic cues learned during training, rather than genuine creative qualities. These findings highlight the urgent need for more sophisticated, semantically aware evaluation frameworks that can better reflect human perceptions. Improving automated assessment tools is crucial for advancing AI applications in creative industries, content curation, and educational tools, ensuring that machine evaluations align more closely with human standards.

While the study exposes significant limitations, it also paves the way for future research focused on multi-dimensional, human-aligned evaluation methods, integrating semantic understanding and bias mitigation. Such efforts are essential for realizing AI’s full potential in supporting and enhancing human creativity in diverse domains.

Deep Analysis

Background

随着深度学习的发展,人工智能在自然语言生成方面取得巨大突破,代表性模型如GPT-3、BERT推动了内容自动化生产。早期研究多关注文本的流畅性与语法正确性,逐步扩展到创造性任务,如诗歌、故事、剧本生成。尽管如此,创造性作为人类独有的复杂认知能力,难以用单一指标衡量。传统评估方法多依赖人类主观评价,存在成本高、主观性强的问题。自动指标如Perplexity、EAD、SBERT-Div等被提出,试图量化文本的创新性、丰富性,但其与人类感知的差距逐渐显现。近年来,LLMs被用作自动评判者,试图模拟人类判断,推动自动化评价体系的发展,但其偏差和局限性也逐渐暴露。

Core Problem

核心问题在于自动评价指标难以准确反映文本的创造性,尤其是主观性强、多维度的特性。现有指标多关注表层特征,如语法、词汇多样性,忽视深层语义、意外性和价值感。另一方面,LLM作为评判者表现出偏向性,偏好生成风格,导致评价结果偏差。这些问题限制了自动评价在实际应用中的可靠性,阻碍了生成模型的优化和创新。

Innovation

本研究首次系统性比较了多种自动指标与人类主观评价的相关性,揭示其局限性。引入LLM作为评判者,发现其偏向性明显,偏好AI生成文本,偏离人类认知。提出结合多维度评价体系,强调深层语义理解的重要性,推动自动评价方法的改进。研究还采用统计分析,量化偏差,为未来设计更符合人类认知的自动评价机制提供理论基础。

Methodology

  • �� 收集WritingPrompts中的100个人类短故事与100个AI生成故事。• 采用11个创造性维度(如新颖性、意外性、价值感)进行问卷调查,获得人类评分。• 计算五种自动指标(Perplexity、EAD、SBERT-Div、Creativity Index、模板评分)对文本进行量化。• 使用LLM作为评判者,给出相同维度的评分。• 采用Pearson和Kendall相关系数分析自动指标与人类评分的相关性。• 比较不同方法的偏差,分析偏向性和相关性弱的原因。

Experiments

实验采用WritingPrompts数据集,包含100个人类故事和100个AI生成故事。自动指标计算包括Perplexity(模型内部不确定性)、EAD(词汇多样性调整)、SBERT-Div(语义多样性)、Creativity Index(创新度)和模板评分(结构复杂性)。人类评价由多名评审基于11个维度打分。LLM作为评判者,采用预训练模型(如GPT-4)对文本进行评分。通过统计分析,比较自动指标与人类评分的相关性,验证自动指标的有效性与偏差。

Results

自动指标与人类评分相关性极低(ρ多在0.2以下),Creativity Index几乎无相关(ρ≈0.07),显示其难以反映创造性核心特质。LLM偏向AI文本,评分趋于满分,表现出系统性偏差,尤其在Surprise和Usefulness维度。不同维度间相关性分析表明,自动指标偏重表层特征,忽视深层语义,导致评价偏差。这些结果表明,自动指标和LLM评判不能作为可靠的创造性评价工具。

Applications

该研究为内容生成与评估提供科学依据,强调需开发更符合人类认知的多维度评价体系。未来可应用于自动内容筛选、创意辅助、内容质量控制等场景,提升内容产业的创新效率。还可结合人机合作,优化自动评价模型,推动AI在艺术、教育等领域的深度应用。

Limitations & Outlook

样本规模局限于短故事,未覆盖诗歌、小说等多样文本类型。自动指标偏重表层特征,缺乏深层语义理解。LLM偏向性可能受训练数据偏差影响,未来需多模型融合与偏差校正。研究未考虑跨文化差异,未来应拓展多语种、多文化背景下的评价体系。

Plain Language Accessible to non-experts

想象你在一家工厂里,生产各种不同的产品。工厂里有很多机器(代表自动评价指标),它们可以快速检查产品的某些特征,比如颜色、大小、形状,但它们不能真正理解产品的创新或美感。工厂还雇了一个工人(代表人类评判者),他可以用经验和感觉判断产品是否新颖、惊喜或有价值。现在,机器虽然快,但经常会误判,偏爱某些特定的外观,而工人则能更全面地感受到产品的独特性。这个比喻说明,自动评价工具虽然方便,但很难真正理解创造力的深层次含义,只有人类的主观判断才能更准确地反映作品的价值。

ELI14 Explained like you're 14

想象你在学校里参加一个比赛,你画了一幅画。老师和同学们都可以评价你的画,但每个人的喜好不同。有的人喜欢颜色鲜艳的,有的人喜欢画里的故事。现在,假设有个机器人可以帮你评分,它只看颜色和线条的漂亮程度,但根本不懂画的故事和创意。这就像自动评价一样,虽然快,但不能真正理解作品的深意。人们发现,这个机器人总是偏爱某些风格,忽视了作品的真正创新和惊喜。其实,真正的创造力很复杂,包含新奇、意外和价值感,只有人类的感觉才能判断得准。自动工具虽然方便,但还不能完全取代人类的眼睛和心灵。

Abstract

Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations. Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.

cs.CL cs.AI cs.CY