ORCID
0000-0002-4925-0591
Subject Area
Education
Abstract
Prosody, the ability to read with appropriate variation in pitch, duration, pausing, and stress, is an important (and challenging to measure) dimension of oral reading fluency (ORF) and comprehension. ORF prosody is currently measured through human raters and rubric-based scoring systems (e.g., the National Assessment of Educational Progress (NAEP) or the Multidimensional Fluency Scale (MDFS scales). The rating process is labor and resource-intensive and introduces rater bias. Recent work has shown acoustic features to predict prosody ratings; however, their models had low generalization performance on cross-domain data (unseen passages and student samples).
The present study investigated the automated assessment of prosody in ORF by developing a machine learning-based scoring method that integrates acoustic, sentence-level, and chunk-level silence features to estimate students' prosody scores. Using 5,841 oral reading recordings from 1,811 Grade 2–4 students in the CORE project, this study extracted 87 acoustic features from the eGeMAPS v02 set and combined them with between-sentence and chunk-level silence features generated through natural language processing (NLP)-based chunking procedures. Silence features were additionally residualized with respect to reading rate and grade level to isolate prosodic structure independent of the global fluency measure, Words Correct Per Minute (WCPM). Multiple machine learning algorithms, including Elastic Net, Random Forest, XGBoost, and LightGBM, are trained and evaluated using Quadratic Weighted Kappa (QWK) as the main model-selection metric to capture ordinal agreement with human ratings.
On average, models trained with chunk-level silence features on top of acoustic-only baselines outperformed acoustic-only models by a margin that varied by model and cross-domain test set. For all models, ensemble methods achieved the best performance, particularly XGBoost, with a QWK of about 0.76 on one held-out test set. Random Forest and LightGBM models also showed reliable improvement with the addition of chunk-based features. Feature importance analysis showed that pause features aligned with syntactic units (particularly noun phrases) were often among the most important features. In sum, the findings showed evidence that linguistically meaningful chunk-level pauses provide additional predictive signal beyond sentence-level silence and acoustic features alone in modeling expressive reading.
Degree Date
Summer 2026
Document Type
Dissertation
Degree Name
Ph.D.
Department
Education Policy & Leadership
Advisor
Akihito Kamata
Number of Pages
119
Format
Creative Commons License

This work is licensed under a Creative Commons Attribution-Noncommercial 4.0 License
Recommended Citation
Wang, Kuo, "Automated Scoring of Prosody in Oral Reading Fluency: Integrating Acoustic and Sentence/Chunk-Level Silence Features with Machine Learning" (2026). Education Policy and Leadership Theses and Dissertations. 24.
https://scholar.smu.edu/simmons_depl_etds/24
Included in
