ORCID

0000-0002-4925-0591

Subject Area

Education

Abstract

Prosody, the ability to read with appropriate variation in pitch, duration, pausing, and stress, is an important (and challenging to measure) dimension of oral reading fluency (ORF) and comprehension. ORF prosody is currently measured through human raters and rubric-based scoring systems (e.g., the National Assessment of Educational Progress (NAEP) or the Multidimensional Fluency Scale (MDFS scales). The rating process is labor and resource-intensive and introduces rater bias. Recent work has shown acoustic features to predict prosody ratings; however, their models had low generalization performance on cross-domain data (unseen passages and student samples).

The present study investigated the automated assessment of prosody in ORF by developing a machine learning-based scoring method that integrates acoustic, sentence-level, and chunk-level silence features to estimate students' prosody scores. Using 5,841 oral reading recordings from 1,811 Grade 2–4 students in the CORE project, this study extracted 87 acoustic features from the eGeMAPS v02 set and combined them with between-sentence and chunk-level silence features generated through natural language processing (NLP)-based chunking procedures. Silence features were additionally residualized with respect to reading rate and grade level to isolate prosodic structure independent of the global fluency measure, Words Correct Per Minute (WCPM). Multiple machine learning algorithms, including Elastic Net, Random Forest, XGBoost, and LightGBM, are trained and evaluated using Quadratic Weighted Kappa (QWK) as the main model-selection metric to capture ordinal agreement with human ratings.

On average, models trained with chunk-level silence features on top of acoustic-only baselines outperformed acoustic-only models by a margin that varied by model and cross-domain test set. For all models, ensemble methods achieved the best performance, particularly XGBoost, with a QWK of about 0.76 on one held-out test set. Random Forest and LightGBM models also showed reliable improvement with the addition of chunk-based features. Feature importance analysis showed that pause features aligned with syntactic units (particularly noun phrases) were often among the most important features. In sum, the findings showed evidence that linguistically meaningful chunk-level pauses provide additional predictive signal beyond sentence-level silence and acoustic features alone in modeling expressive reading.

Degree Date

Summer 2026

Document Type

Dissertation

Degree Name

Ph.D.

Department

Education Policy & Leadership

Advisor

Akihito Kamata

Number of Pages

119

Format

.pdf

Creative Commons License

Creative Commons Attribution-Noncommercial 4.0 License
This work is licensed under a Creative Commons Attribution-Noncommercial 4.0 License

Share

COinS