
Self-supervised models unlock low-resource L2 speech assessment
Researchers demonstrate that self-supervised speech models can assess L2 learner pronunciation without labeled training data, a shift that unlocks assessment in resource-constrained regions. Using DTW alignment over WavLM embeddings, the approach evaluates phonetic accuracy, rhythm, and intonation across English and Japanese learners, matching or exceeding human rater consistency on phonetic tasks. This work signals how foundation models trained on unlabeled speech can be repurposed for pedagogical evaluation, reducing the annotation burden that has historically gatekept language assessment tools to well-funded institutions.58




.jpg)



















