Enhancing Diversity of LLM-Generated Educational Tasks

TL;DR

Proposes CreativeDC, a two-stage prompting framework, boosting Python task diversity by 1.6× with structured divergent-convergent reasoning.

cs.AI 🟡 Intermediate 2025-12-30 53 views
Manh Hung Nguyen Sebastian Tschiatschek Adish Singla
education LLMs task generation diversity creativity

Key Findings

Methodology

This study draws on creativity theories, implementing a two-stage prompting process: first encouraging broad idea exploration (divergent thinking), then selecting and refining promising ideas into valid tasks (convergent thinking). Applied to Python programming, the approach uses theme and concept inputs, combined with the Qwen-3-235B model, to generate diverse high-utility tasks. Evaluation involves automated Vendi Score and expert assessments, confirming significant diversity improvements. The framework promotes exploration before commitment, leading to about 1.6× more distinct high-quality tasks compared to baselines.

Key Results

  • At K=50, CreativeDC achieves a Vendi Score of 13 and 16 task clusters in expert evaluation, outperforming baselines by approximately 1.6× in diversity. Tasks also scored higher in novelty and interest, with expert ratings of 0.81 versus 0.50.
  • Across themes and concepts, the method consistently enhances diversity and utility, demonstrating robustness and generalization.
  • Automated and expert metrics strongly correlate, validating the effectiveness of the structured prompting approach.

Significance

This work addresses the critical challenge of content homogeneity in large language models, offering a systematic framework that balances creativity and utility. It advances the field of automated educational content generation, enabling scalable, diverse, and engaging learning materials. The approach can be extended to various disciplines, fostering personalized and innovative education at scale, with broad implications for AI-assisted teaching and curriculum design.

Technical Contribution

The core innovation lies in integrating a staged reasoning process—divergent idea generation followed by convergent refinement—within prompting strategies. This structured approach guides models to explore broader semantic spaces, resulting in richer outputs. The introduction of the 'Effective Diversity' metric, combining automated and human evaluations, provides a comprehensive measure of high-quality content variation. Unlike traditional single-step prompts, this method systematically enhances creative exploration, setting a new standard for diversity in task generation.

Novelty

This is the first systematic application of a staged divergent-convergent reasoning framework in educational task generation with large models. It moves beyond simple prompt engineering, embedding a structured thought process that significantly improves output diversity and relevance. The combination of automated metrics and expert validation further distinguishes this work from prior approaches that focus solely on correctness or difficulty.

Limitations

  • The current validation is limited to introductory Python tasks; applicability to more complex or interdisciplinary subjects remains untested.
  • Expert evaluations are based on small samples, necessitating larger-scale user studies for broader validation.
  • The two-stage process is not iterative; future work could explore multi-round refinement to further enhance quality and diversity.

Future Work

Future research will extend this framework to other disciplines like mathematics and science, incorporating learner profiles for personalized task generation. Multi-round iterative prompting and reinforcement learning could further improve task quality. Additionally, integrating student feedback and performance data will help tailor content dynamically, advancing personalized AI-driven education.

AI Executive Summary

Large language models (LLMs) have revolutionized automated content creation, especially in education. However, a persistent challenge is content homogeneity, which limits engagement and coverage. Traditional prompt methods often lead models to produce similar, repetitive outputs, constraining the diversity essential for personalized learning. To address this, this study introduces CreativeDC, a structured two-stage prompting framework inspired by creativity theories. The first stage encourages models to explore a broad semantic space, generating diverse ideas without constraints. The second stage refines and filters these ideas into valid, high-utility tasks aligned with input requirements. Applied to Python programming, the approach leverages theme and concept inputs, combined with the Qwen-3-235B model, to produce significantly more diverse tasks—about 1.6 times more—than baseline methods. Evaluation via automated Vendi Score and expert assessments confirms the method's effectiveness, showing consistent improvements across multiple metrics. This framework not only enhances content variety but also maintains task relevance and quality, addressing a core bottleneck in scalable, personalized education. Its implications extend beyond programming to broader educational domains, paving the way for AI-assisted curriculum design that is both rich and adaptable. Future work aims to generalize this approach to other subjects, incorporate learner-specific data, and develop iterative prompting strategies for even higher quality and diversity.

Deep Analysis

Background

Recent advances in large language models (LLMs) like GPT-4 and CodeX have enabled automatic generation of educational content, including questions, exercises, and explanations. Prior works such as GPT-3-based question generators and CodeX for programming tasks have shown promising results. However, these models tend to produce repetitive, homogeneous outputs due to their training on large datasets with dominant patterns, leading to the 'artificial hivemind' phenomenon. Researchers have attempted fine-tuning, reinforcement learning, and multi-agent debate frameworks to promote diversity, but these often compromise content quality or require extensive data. The challenge remains to generate diverse yet high-utility educational tasks at scale, which is crucial for personalized learning, curriculum coverage, and engagement.

Core Problem

Despite the capabilities of modern LLMs, their tendency toward content homogeneity hampers the creation of varied educational materials. Existing prompt strategies often lead to similar outputs, limiting the scope of personalized and comprehensive education. The core problem is how to systematically enhance the diversity of generated tasks without sacrificing their relevance, correctness, or difficulty. This is particularly difficult because encouraging diversity can conflict with maintaining utility, and naive approaches may result in irrelevant or low-quality tasks. Developing a structured prompting framework that balances exploration and refinement is essential to overcome these bottlenecks.

Innovation

The main innovation is the integration of a staged reasoning process—divergent and convergent thinking—within prompt design. First, the model explores a wide range of ideas related to the theme without constraints, promoting creative diversity. Then, it selects and refines the most promising ideas into valid tasks that meet input criteria. This approach mimics human creative processes and encourages broader exploration before focusing on feasibility. Additionally, the paper introduces the 'Effective Diversity' metric, combining automated embeddings-based scores and expert assessments, to evaluate the quality of diverse high-utility tasks. This systematic framework surpasses traditional single-step prompts, providing a scalable solution for generating varied educational content.

Methodology

  • �� Define a two-stage prompting process: first, instruct the model to brainstorm broadly around the theme, generating multiple diverse ideas without constraints. • Use theme and concept inputs to guide the initial exploration, encouraging unusual or surprising ideas. • In the second stage, prompt the model to select one idea from the brainstormed list and refine it into a valid programming task, ensuring it satisfies all input requirements and focuses on the specified concept. • Employ the Qwen-3-235B model with temperature set to 1.0 for generation. • Calculate the Vendi Score based on embeddings from Qwen-Embedding-0.6B to quantify diversity automatically. • Conduct expert evaluations on a subset of tasks, rating utility, novelty, and interest. • Combine automated metrics and expert assessments to compute 'Effective Diversity,' emphasizing high-utility, diverse tasks.

Experiments

The experimental setup involved generating 50 tasks per context across 20 themes and concepts, comparing CreativeDC with baseline methods: direct prompt (BASE) and a thought-step prompt (COT). The evaluation used the Vendi Score for automated diversity measurement and expert panels for qualitative assessment of task relevance, novelty, and engagement. Model parameters included a temperature of 1.0, with multiple random seeds to ensure robustness. Statistical tests confirmed the significance of diversity improvements. The experiments demonstrated that CreativeDC consistently outperformed baselines in both automated and human evaluations, with a focus on high-utility task diversity and quality.

Results

At K=50, CreativeDC achieved an average Vendi Score of 13, with expert evaluations indicating 16 distinct high-utility task clusters, both significantly higher than baseline methods (~8 clusters). Tasks generated by CreativeDC scored higher in novelty (0.81 vs. 0.50) and interest, with expert ratings confirming better relevance and engagement. The results validate that the staged prompting approach effectively broadens creative exploration while maintaining utility, leading to richer, more diverse educational content. The consistency across themes and concepts suggests strong generalization potential.

Applications

This method can be directly applied in automated question banks, personalized tutoring systems, and curriculum development platforms. It enables educators and AI systems to produce a wide array of engaging tasks tailored to different difficulty levels and learner profiles, fostering active learning and curiosity. Long-term, it supports scalable, adaptive education environments where content adapts dynamically to student needs, making learning more interactive and effective.

Limitations & Outlook

The current validation is limited to introductory Python tasks, and its effectiveness in more complex or interdisciplinary domains remains untested. Expert evaluations are based on small samples, which may not fully capture real-world variability. The two-stage process is not iterative, potentially limiting refinement; future work could incorporate multi-round feedback. Additionally, computational costs and model biases may affect scalability and fairness, requiring further investigation.

Plain Language Accessible to non-experts

想象你在厨房里做菜,平常你可能只会做几道菜,比如炒蛋或煮面。现在,厨师用一种特别的方法,先让自己想出很多奇怪的菜,比如用水果做披萨、用糖做汤,然后再从中挑出最有趣、最可行的菜,认真准备。这样做出来的菜不仅丰富多样,还特别新颖。这个方法就像让模型先疯狂想各种点子,然后筛选出最棒的,确保每次都能带来新鲜感和惊喜。通过这样的步骤,厨师可以不断创新,做出令人惊喜的美味佳肴。

ELI14 Explained like you're 14

你知道吗?在学校里,老师布置作业,有时候题目都差不多,没有新意。科学家们也遇到这个问题,他们用一种特别的方法帮电脑设计出很多不同的题目。这个方法就像你在玩游戏时,先想很多奇怪的点子,比如让猫变成超人,或者机器人去火星,然后再挑出最酷、最合理的点子,写成题目。这样一来,学生们就能遇到各种新奇的题目,不会觉得无聊。科学家用这个技巧,让电脑先发散思维,想出很多点子,再收敛成最好的题目,效果特别棒!

Glossary

发散-收敛思维 (Divergent-Convergent Thinking)

一种创造力策略,先广泛探索各种想法(发散),再筛选和完善(收敛)。在论文中用于引导模型生成多样任务。

引导模型在任务生成中实现多样性与实用性的平衡。

Vendi Score (文迪评分)

一种自动衡量内容多样性的指标,基于内容嵌入的余弦相似度和熵计算,反映任务的差异程度。

用于评估生成任务的多样性,确保内容不重复。

高效多样性 (Effective Diversity)

在保证任务实用性的前提下,衡量高质量任务的多样性指标,结合自动指标和专家评估。

核心评价标准之一,确保生成内容丰富且有用。

自动化指标 (Automated Metrics)

利用模型和算法自动评估内容的多样性和质量,减少人工干预。

在实验中用于快速筛查和比较不同方法。

专家评估 (Expert Evaluation)

由领域专家对生成内容的实用性、趣味性和新颖性进行主观评分,确保内容符合教育需求。

验证自动指标的有效性。

Open Questions Unanswered questions from this research

  • 1 如何将该方法推广到数学或科学等其他学科,尤其在内容复杂度和难度层级上的适应性问题仍未解决。
  • 2 模型在极端主题或高度专业化内容中的表现尚未充分验证,可能存在内容偏差或不合理的风险。

Applications

Immediate Applications

教育内容自动生成

教师或教育平台可利用此方法快速生成多样化练习题,提升教学效率和学生兴趣。

个性化学习路径

结合学生模型,为不同学生定制符合其水平和兴趣的练习,增强学习效果。

Long-term Vision

智能教育系统

未来可构建全自动化、个性化的教育内容生产平台,实现大规模定制化学习资源。

Abstract

Large language models (LLMs) have shown the potential for generating educational content at scale, assisting educators in creating practice tasks or synthesizing data for training educational models. However, LLMs suffer from the ``Artificial Hivemind'' effect, where they produce homogeneous content. This homogeneity limits the diversity of LLM-generated tasks, a crucial factor in these educational settings. In this paper, we investigate how to increase the diversity of generated tasks while keeping their utility high. Inspired by the divergent--convergent thinking stages in creativity literature, we propose a prompting framework with two reasoning stages: (1) exploring the creative space, and (2) satisfying the input requirements. We evaluate CreativeDC, a method instantiated from this framework in the domain of Python programming, using both automated metrics and expert evaluation. Results show that CreativeDC produces significantly more distinct high-utility tasks (about $1.6\times$) than baselines. Our work offers an effective approach for generating and evaluating more diverse tasks at scale.

cs.AI