Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
CAV and TopK SAE auditing show that concept recoverability is not equivalent to influence on speaking scores.
Key Findings
Methodology
The paper probes internal layers of a BERT transcript grader and a Whisper speech-text grader with Concept Activation Vectors. A class-balanced, L2-regularized linear hinge-loss classifier yields a CAV; balanced accuracy measures concept recoverability, while the gradient cosine-distance metric Bgr measures alignment between the concept direction and the score gradient. The authors also train TopK Sparse Autoencoders, learn CAVs in sparse latent space, and map them back through the linear decoder.
Key Results
- Recoverability is strongly layer- and architecture-dependent. On BULATS, Whisper’s 2048-dimensional dense.in recovers every tested concept, whereas more than half of L1 concepts are degenerate in the two BERT layers and Whisper’s 64-dimensional act.out. BERT proficiency CAV accuracy is 88.1%/87.4% for ≥A2.
- Influence is not implied by encoding. BERT’s ≥A2 sensitivity changes from Bgr=0.81 at Layer 1 to 0.46 at Layer 2; Dutch L1 changes from 0.91 to 0.52. Whisper demographic and L1 concepts remain near the no-sensitivity baseline of 1, while its strongest proficiency sensitivity is about 0.83.
- TopK SAEs improve separability: no concept remains degenerate in the latent-space analysis. However, mapped SAE-CAV sensitivity generally moves toward 1, especially in BERT’s 20-dimensional Layer 2; Dutch Bgr there shifts from 0.52 to about 0.92. Better recoverability can therefore weaken influence estimates.
Significance
The study establishes an important distinction for high-stakes language assessment: a model may encode an attribute without using it to score a response. This matters for L1, age, and gender, whose presence in training data can create shortcuts. The results also show that fairness conclusions depend on modality, layer, and probe representation. For education, immigration, and employment certification, the work provides a more granular audit logic than aggregate accuracy or group score gaps alone.
Technical Contribution
The authors extend regression-based TCAV analysis to modern Transformer speaking graders and jointly report recoverability and gradient influence. Their main engineering contribution is a frozen-model TopK SAE pipeline: activations are encoded into a wider sparse space, CAVs are learned there, and the direction is mapped back as WD dz for the original Bgr calculation. This demonstrates that the probe representation itself changes measured sensitivity, rather than merely improving interpretability.
Novelty
Relative to prior work detecting deliberately injected L1 bias in feature-based graders, this study applies CAV auditing to both BERT text and Whisper multimodal graders and compares dense activations with SAE latents. Its central novelty is conceptual and empirical: recoverability and influence are separable, architecture-dependent properties. The paper therefore challenges the common practice of treating a successful demographic probe as direct evidence of scoring bias.
Limitations
- BULATS contains uneven L1 and demographic groups, and several concepts correlate with proficiency; consequently, Bgr can reflect legitimate data correlations rather than unfair causal use of an attribute.
- SAEs are trained on frozen activations and map directions only through the decoder’s reconstruction span. CAVs and local gradients remain associational evidence, not a causal proof of discriminatory decisions.
- Speak & Improve lacks full demographic metadata; gender is inferred with an ECAPA-TDNN classifier. This limits the strength of cross-corpus conclusions and excludes reliable analysis of several other attributes.
Future Work
Future work should replicate the analysis across examinations, languages, accents, and intersectional groups, combining CAVs with counterfactual speech editing, calibration, and causal interventions. Systematic sweeps over SAE width, TopK sparsity, and layer selection could clarify why low-dimensional layers attenuate influence. A practical outcome would be a standardized audit protocol combining representation recoverability, gradient influence, subgroup performance, and human review.
AI Executive Summary
Automatic speaking assessment is increasingly used in immigration, employment certification, and education. Yet a grader can be accurate overall while exploiting irrelevant cues such as first language, age, or gender. Aggregate RMSE and correlation cannot reveal which internal signals drive a score. Labroo, Qian, and Knill address this problem by auditing a BERT transcript grader and a Whisper speech-text grader with Concept Activation Vectors (CAVs).
A CAV is a direction separating examples that exhibit a human-defined concept from those that do not. Linear-classifier accuracy asks whether a concept is represented; Bgr, the cosine distance between that direction and the score gradient, asks whether moving along it changes the prediction. The study also trains TopK Sparse Autoencoders (SAEs), learns CAVs in sparse latent space, and maps them back through a linear decoder. On BULATS, every concept is recoverable at Whisper’s 2048-dimensional dense.in, but many L1 concepts collapse in other layers, showing that recoverability belongs to a representation, not simply to a concept.
Influence is equally architecture-specific. BERT’s ≥A2 Bgr falls from 0.81 at Layer 1 to 0.46 at Layer 2, and Dutch L1 falls from 0.91 to 0.52; Whisper demographic and L1 scores remain near the neutral value 1. SAEs remove degenerate probes but usually attenuate mapped sensitivity, most strongly in BERT’s 20-dimensional layer. The practical message is clear: an audit must ask both whether a model can decode an attribute and whether that attribute changes the score.
Deep Analysis
Background
Transformer encoders such as BERT, wav2vec 2.0, and Whisper have improved automatic L2 speaking assessment. TCAV introduced concept directions for neural interpretation, while Wei et al. adapted CAVs to regression and exposed injected L1 bias in feature-based graders. Modern Transformer representations, however, are distributed and entangled through superposition, making single-direction probing less reliable.
Core Problem
A fair grader should use fluency, grammatical control, and content, not L1, age, or gender. The technical difficulty is that an attribute may be represented without affecting the output, or may correlate with proficiency in the data and appear influential without being an unfair shortcut. Neither overall accuracy nor a demographic probe alone resolves this distinction.
Innovation
The paper contributes three elements. First, it extends regression CAV auditing to BERT and Whisper graders. Second, it compares recoverability and score influence across modalities and layers. Third, it tests TopK SAE representations. SAEs can make concepts more linearly separable, but the authors also measure whether decoder mapping preserves influence, revealing that improved separability may attenuate the original gradient relationship.
Methodology
- �� Split each grader into representation h=Fh(x) and output ŷ=Fy(h).
- �� Train a balanced, L2-regularized linear classifier with hinge loss on binary concept labels; its normal vector is the CAV d.
- �� Use balanced accuracy for recoverability: below 60% is weakly usable, while a class accuracy below 10% is degenerate.
- �� Scale the CAV by mean activation norm and compute Bgr, the mean cosine distance to ∇hFy. Bgr≈1 denotes insensitivity; values below or above 1 indicate score-increasing or score-decreasing alignment.
- �� Train a TopK SAE, z=TopK(ReLU(WEh+bE)), with reconstruction ĥ=WDz+bD. Learn a latent CAV dz and map it as WDdz before evaluating Bgr.
- �� Select SAE hyperparameters using reconstruction error, sparsity, dead-feature fraction, and decoder coherence.
Experiments
The primary corpus is BULATS Part 1: 1,657 training, 188 development, and 465 evaluation speakers, scored from 0 to 6. Speak & Improve 2025 provides 3,068/438/300 speakers for train/development/evaluation. BERT is probed at 600-dimensional Layer 1 and 20-dimensional Layer 2; Whisper at 2,048-dimensional dense.in and 64-dimensional act.out. Concepts include CEFR thresholds, gender, age, and L1. Metrics are RMSE, PCC, within-0.5/1.0 accuracy, balanced CAV accuracy, and Bgr, with standard CAV versus SAE-CAV comparisons.
Results
On BULATS, BERT and Whisper achieve PCC 0.756 and 0.786, with RMSE 0.817 and 0.781. Whisper dense.in is the only layer recovering every concept; other layers contain many degenerate L1 probes. BERT’s ≥A2 Bgr shifts from 0.81 to 0.46, while Dutch shifts from 0.91 to 0.52. Whisper demographic and L1 values stay near 1. SAE probes eliminate degeneracy but generally move Bgr toward 1, most strongly in the 20-dimensional BERT layer.
Applications
Assessment providers can audit L1, age, gender, and proficiency at every internal layer before deployment. Developers can identify where a signal first appears and combine Bgr with subgroup calibration and score gaps to decide whether to rebalance data, modify ASR, retrain an encoder, or adjust the regression head. This is particularly relevant to immigration, employment, and high-stakes educational certification.
Limitations & Outlook
The evidence is concentrated in BULATS, where small L1 groups and concept–proficiency correlations complicate fairness interpretation. Speak & Improve lacks rich demographic metadata and uses ECAPA-TDNN gender inference. SAE-CAVs depend on reconstruction quality and decoder span; local gradients and CAVs do not establish causal discrimination. Future evaluations should add counterfactual audio, intersectional groups, significance tests, more languages, and longitudinal monitoring after deployment.
Plain Language Accessible to non-experts
Imagine a teacher who grades an English speaking test. The teacher should listen for fluent delivery, accurate sentences, and clear ideas. But the teacher may also notice an accent, age, or background. The researchers first draw a “clue arrow”: if examples from Dutch speakers point in one direction inside the teacher’s mental notes, that arrow is a CAV.
Finding the arrow does not mean the teacher used it to grade. The researchers therefore push the clue slightly and watch the score. If the score moves with the clue, the clue may influence the decision; if the clue is at a right angle to the score direction, the teacher may recognize it without relying on it.
They also use an organizer that sorts messy notes into sparse drawers, like arranging a crowded filing cabinet. This makes clues easier to find, but the sorting process can lose some relationships. That is what happens: the SAE makes concepts easier to separate, yet often makes their measured effect on scores smaller. A fair audit must therefore ask two questions: what can the system recognize, and what does it actually use?
ELI14 Explained like you're 14
Imagine a robot teacher grading your English speaking test. It should care about smooth speaking, grammar, and ideas—not whether your family speaks Arabic, how old you are, or what you look like. But the robot’s brain is not a neat answer sheet; it is more like a giant pile of tangled gaming cables.
Researchers use a trick called a CAV, basically a direction-finder. They give the robot examples from Dutch speakers and non-Dutch speakers and check whether its hidden numbers can draw a boundary between them. But wait—recognizing a clue is not the same as using it! So they nudge the clue and see whether the final score changes.
They also use an SAE, like organizing your game inventory into labeled slots. After organizing, more clues become easy to find. Surprisingly, the clues often seem less connected to the score. For BERT, the A2 proficiency measure changes from 0.81 in an earlier layer to 0.46 later; Whisper’s demographic and L1 measures stay close to the neutral value 1.
The big lesson is like checking a referee’s decision. Seeing your team’s jersey does not prove the referee favored your team. You must check both whether the referee noticed the jersey and whether the final call changed because of it. That double check can make speaking tests safer and fairer!
Glossary
Concept Activation Vector (CAV)
A direction in activation space representing a human-interpretable concept. Technically, it is the normal vector of a linear classifier separating concept-positive and concept-negative examples.
Used for L1, gender, age, and proficiency probes.
TCAV
Testing with Concept Activation Vectors measures whether concept directions align with model predictions. The original framework targeted classification; this paper uses its regression adaptation.
Provides the conceptual basis for the audit.
Bgr sensitivity
The mean cosine distance between a concept direction and the gradient of the predicted score. A value near 1 indicates little local sensitivity, whereas deviation from 1 indicates directional influence.
Separates representation from score influence.
TopK Sparse Autoencoder
An autoencoder that expands activations into a wider latent space and retains only the K largest latent entries. It aims to expose more disentangled features.
Used to learn sparse-space CAVs and map them back to activations.
Concept recoverability
The ability to decode a concept from an internal representation using a linear classifier. It indicates encoding, not necessarily use in prediction.
Measured by balanced classification accuracy.
Whisper-Fuse
A multimodal grader that combines log-mel audio embeddings with transcript-token representations through decoder cross-attention. It therefore has access to acoustic and lexical information.
Serves as the speech-text comparison architecture.
Open Questions Unanswered questions from this research
- 1 CAVs and local gradients reveal associations, not causal discrimination. Counterfactual speech transformations, controlled interventions, and causal estimators are needed to determine whether a demographic attribute changes a real decision.
- 2 The interaction between SAE width, TopK sparsity, decoder coherence, and sensitivity remains unclear. In particular, why low-dimensional layers lose mapped influence requires a stronger geometric or theoretical account.
- 3 Real candidates have intersecting attributes such as L1, accent, age, and socioeconomic background. Binary single-concept probes do not yet capture these mixtures reliably.
Applications
Immediate Applications
Pre-deployment grader audit
Testing providers can train CAVs for L1, gender, age, and proficiency at multiple BERT or Whisper layers, then report balanced accuracy and Bgr together. An irrelevant concept that is both recoverable and far from Bgr=1 should trigger rebalancing, calibration, or model intervention.
Model-development diagnosis
Engineering teams can compare transcript layers, multimodal fusion layers, and output heads to locate where demographic signals emerge. Combined with human scores and subgroup gaps, this can distinguish ASR, encoder, fusion, and regression-head problems.
Long-term Vision
Auditable high-stakes language assessment
A future standard could combine CAV recoverability, gradient influence, subgroup calibration, counterfactual tests, and human review in immigration and employment certification. Main obstacles include multilingual metadata, privacy, causal validation, and monitoring cost.
Abstract
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tests, making it critical to show that their scores depend on speaking proficiency rather than irrelevant speaker attributes such as first language (L1) or age. Transformer-based foundation models have improved the accuracy of these L2 speaking graders, but their black-box representations make fairness and interpretability analysis more difficult. Building on prior work that used Concept Activation Vectors (CAVs) to detect bias towards unwanted attributes (`concepts') in feature-based graders, we extend CAV-based analysis to two neural speaking assessment systems: a text-based BERT grader and a speech-and-text multimodal grader based on Whisper. CAVs represent human-interpretable concepts as directions in a model's activation space, allowing us to distinguish between whether a concept is encoded in a model's internal representations and whether it influences the predicted score, the latter quantified using a gradient-based sensitivity metric. Since CAVs rely on linear separability, which is less likely in complex neural embedding spaces, we also investigate whether sparse autoencoders (SAEs) provide cleaner concept directions by learning CAVs in a sparse latent space and mapping them back to activation space. Our analysis shows that concept recoverability depends strongly on the representation and architecture being probed, rather than on the concept alone. Sensitivity to concepts is also architecture-dependent. SAEs make concepts more linearly recoverable, but attenuate the original activation-space sensitivity, especially in low-dimensional layers. These findings highlight the need to distinguish concept recoverability from concept influence when auditing bias in speaking assessment systems.