Findings
Every result from the study, including the ones that did not go our way. Based on 167 recordings from 42 pupils at a single school — so these numbers describe that setting, not every school.
Finding 1 — Do the teachers agree with each other?
01On at least one measure the teachers' numbers genuinely differ. Human scores are noisy there, which caps how well any ASR system can appear to agree.
Finding 2 — Which model size should each device use?
02`medium` is the most accurate model tested, and the table below shows what you give up on each smaller device.
Finding 3 — Can we correct the computer's scores toward the teacher's?
03Yes — for reading speed (WCPM), the computer's average error against teachers drops from **-13.603** to **0.149 WCPM** on recordings it has never seen.
Finding 5 — One bilingual model, or one per language?
05The pre-chosen option was right — a one bilingual model + language flag works best overall, and the differences between the three options are small.
Finding 6 — Do readers fall into natural groups?
06No — when the computer groups the readings without being told the answer, the groups it finds are only slightly related, which suggests Phil-IRI's three levels are a useful convention rather than three real, separate kinds of reader.