Towards Scalable Measurement of Durable Skills

TL;DR

Using Executive LLM framework enhances scalable measurement of durable skills, significantly increasing evidence density.

cs.HC 🔴 Advanced 2026-09-15 3 views
Amir Globerson Amy Keeling Anisha Choudhury Anna Iurchenko Aviad Segal Avinatan Hassidim Ayça Çakmakli Ben Gomes Benn Witt Cathy Cheunga Cristine Legare Diana Akrong Eliad Carmi Elisabeth Bauer Gal Elidan Hadas Gelbart Hairong Mu Katherine Chou Lev Borovoi Nir Kerem Niv Efron Noa Kerrem Gilo Preeti Singh Rajvi Kapadia Rena Levitt Roni Rabin Ronit Levavi Morad Rotem Yulzary Shashank Agarwal Sophie Allweis Tracey Lee-Joe Tzvika Stein Yael Bar Moshe Yael Haramaty Yaniv Carmel Yishay Mor Yoav Bar Sinai Yoav Bergner Yossi Matias Yuri Lev
large language model durable skills scalability psychometrics educational technology

Key Findings

Methodology

The study proposes a framework based on large language models (LLMs) for measuring durable skills like collaboration, creativity, and critical thinking. The framework involves subjects interacting with AI teammates in a human-like manner to maintain ecological validity, while the Executive LLM guides conversations to gather high-density evidence of skills. An AI evaluator analyzes transcripts to assess skill levels.

Key Results

  • Using Executive LLM, evidence density in conversations significantly increased, with conflict resolution skills showing an 85% evidence rate.
  • AI automated scoring closely matches expert ratings, with Cohen's Kappa values between 0.45-0.64.
  • For creativity tasks, the AI evaluator matches human expert ratings, proving effective in complex tasks.

Significance

This research demonstrates the potential of LLMs for measuring complex social and cognitive constructs, addressing the traditional conflict between ecological validity and psychometric rigor. By introducing the Executive LLM, the study provides a new technological path in educational assessment, potentially influencing future standards.

Technical Contribution

Technical contributions include the development of the Executive LLM framework, which guides conversations to gather more evidence while maintaining natural dialogue. Additionally, the AI evaluator's automated scoring matches human expert ratings, proving feasible for large-scale assessments.

Novelty

First to use Executive LLM to guide conversations for skill evidence, offering a more natural and efficient assessment method compared to existing hard-coded rule simulations.

Limitations

  • In complex dialogues, AI may fail to fully capture subtle human social cues.
  • Further validation is needed for applicability across different cultural contexts.

Future Work

Future research could explore applicability across different cultural contexts and expand to more durable skills assessment areas like emotional intelligence and leadership.

AI Executive Summary

In the modern workplace, durable skills like collaboration, creativity, and critical thinking are crucial, yet measuring them remains challenging. Traditional assessment methods struggle to balance ecological validity with psychometric rigor. This study proposes a new framework based on large language models (LLMs) that simulates human dialogue, maintaining the authenticity of natural interactions while providing control and reproducibility.

The core of this framework is the Executive LLM, which not only acts as an AI teammate but also guides conversations to gather high-density evidence of skills. Combined with an AI evaluator, the framework can automatically analyze transcripts and assess skill levels. Experimental results show that using the Executive LLM significantly increases evidence density in conversations, and AI automated scoring closely matches expert ratings.

This research offers a new technological path in educational assessment, demonstrating the potential of LLMs for measuring complex social and cognitive constructs. Future research could explore applicability across different cultural contexts and expand to more durable skills assessment areas like emotional intelligence and leadership.

Deep Analysis

Background

Durable skills like collaboration, creativity, and critical thinking are increasingly important in the modern workplace. However, measuring these skills has been challenging because traditional methods struggle to balance ecological validity with psychometric rigor. Previous research has focused on automated scoring and detailed process data analysis, but these methods have limited application in real-world settings.

Core Problem

The core problem is how to conduct skill assessments that maintain natural interactions while being controllable and reproducible. Traditional methods often fail to balance ecological validity and psychometric rigor, leading to these skills being overlooked in educational curricula.

Innovation

The core innovation of this study is the introduction of the Executive LLM framework, which simulates human dialogue to maintain the authenticity of natural interactions while providing control and reproducibility. Compared to existing hard-coded rule simulations, it offers a more natural and efficient assessment method.

Methodology

  • �� Use large language models (LLMs) to simulate human dialogue, maintaining ecological validity.
  • �� Introduce Executive LLM to guide conversations to gather high-density evidence of skills.
  • �� Use AI evaluator to automatically analyze transcripts and assess skill levels.

Experiments

The experimental design includes participants interacting with AI teammates in tasks like scientific discussions and debates. Cohen's Kappa is used to evaluate the consistency between AI automated scoring and expert ratings, showing high agreement.

Results

Experimental results show that using the Executive LLM significantly increases evidence density in conversations, with conflict resolution skills showing an 85% evidence rate. AI automated scoring closely matches expert ratings, with Cohen's Kappa values between 0.45-0.64.

Applications

The framework can be used in educational assessments to help teachers better understand students' durable skill levels and provide data support for curriculum design.

Limitations & Outlook

While the framework performs well in experiments, in complex dialogues, AI may fail to fully capture subtle human social cues. Further validation is needed for applicability across different cultural contexts.

Plain Language Accessible to non-experts

Imagine working in a team where members need to solve problems together. Normally, interactions are natural, but if we want to assess everyone's collaboration skills, we need a method that maintains natural interactions while allowing effective assessment. This study is like equipping the team with a smart assistant that not only participates in discussions but also guides them to clearly showcase everyone's abilities. This way, we can better understand each team member's skills.

ELI14 Explained like you're 14

Imagine playing a game with friends where you need to work together to win. This study is like adding a super-smart AI teammate to your game, which not only plays with you but also helps you show off your best teamwork skills. This way, teachers can better understand your collaboration abilities and help you develop these skills in future learning.

Glossary

Large Language Model (LLM)

An AI model capable of understanding and generating natural language text, widely used in dialogue systems and text generation.

Used to simulate human dialogue, maintaining ecological validity.

Executive LLM

A specialized large language model that guides conversations to gather high-density evidence of skills.

Used to guide conversations to gather more skill evidence.

AI Evaluator

An AI tool for analyzing transcripts and assessing skill levels.

Used to automatically analyze transcripts and assess skill levels.

Cohen's Kappa

A statistical measure of inter-rater agreement, with higher values indicating better agreement.

Used to evaluate the consistency between AI automated scoring and expert ratings.

Durable Skills

Skills like collaboration, creativity, and critical thinking that are crucial in the modern workplace.

The core measurement target of the study.

Open Questions Unanswered questions from this research

  • 1 How to apply this framework across different cultural contexts to ensure its universality and effectiveness.
  • 2 How AI can better capture subtle human social cues in more complex social interactions.

Applications

Immediate Applications

Educational Assessment

Teachers can use this framework to better understand students' durable skill levels and provide data support for curriculum design.

Long-term Vision

Intelligent Education Systems

Future systems could automatically assess and enhance students' durable skills, transforming educational practices.

Abstract

Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an "Executive LLM" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.

cs.HC