READSmart

Finding 01

Finding 1 — Do the teachers agree with each other?

In one sentence: On at least one measure the teachers' numbers genuinely differ. Human scores are noisy there, which caps how well any ASR system can appear to agree.

Why this matters

Every score this system produces is judged against a teacher's score. So before asking "does the computer agree with the teacher?" we have to ask "do teachers agree with each other?" If two teachers listening to the same child write down very different numbers, then there is no single correct answer to compare against, and no amount of engineering can fix that.

To check this, some recordings were deliberately scored by more than one teacher.

What was measured

  • 204 scoring sheets in total, covering 167 recordings
  • 25 of those recordings were scored by two or more different teachers
  • 5 teachers took part
  • Breakdown: 18 recording(s) scored by 2 teachers, 7 recording(s) scored by 3 teachers

The headline numbers

The score below is called an ICC. Read it like a percentage of "how much of the difference between scores is real difference between students, rather than disagreement between teachers." 1.00 is perfect agreement. Above 0.75 is good.

The range matters more than the score itself. With only 25 double-scored recordings, a single number is not enough to trust — the range says how sure we can be. A range that spans both "bad" and "good" means we cannot tell yet.

What was scoredAgreement (ICC)How sure (95% range)VerdictTwo teachers usually within
Words read correctly0.9800.95 – 0.99excellent±12.029 words
Reading time1.0001.00 – 1.00excellent±0.0 seconds
Total miscues0.5880.14 – 0.77too uncertain to call±11.409 miscues
Word reading accuracy0.427-0.09 – 0.71too uncertain to call±11.734 %
Words Correct Per Minute (derived)0.9600.90 – 0.98excellent±9.553 WCPM

The last column is the plain-English one: 95% of the time, two teachers scoring the same recording land within that much of each other.

⚠️ Reading time shows perfect agreement with zero variation. That almost never happens with independent human judgement — it means the value was recorded once and shared between the scorers, not judged separately by each. It cannot be counted as evidence that teachers agree.

⚠️ Words Correct Per Minute is calculated from other fields on the sheet, not written down independently. Teachers who agree on the inputs agree on it automatically, so its high score is inherited, not independent confirmation.

Words read correctly ███████████████████████████· 0.980 Reading time ████████████████████████████ 1.000 Total miscues ████████████████············ 0.588 Word reading accuracy ████████████················ 0.427 Words Correct Per Minute ███████████████████████████· 0.960

Important: a low score here does not always mean teachers disagreed

ICC is a ratio. It compares how much students differ from each other against how much teachers differ from each other. So if all the students score similarly on a measure, ICC drops — even when the two teachers wrote down almost the same number. The table below separates those two situations.

What was scoredHow far apart the teachers wereHow far apart the students wereReading
Words read correctly±3.154 words±31.094 wordsstudents differ clearly, so ICC is meaningful
Reading time±0.0 seconds±42.814 secondsstudents differ clearly, so ICC is meaningful
Total miscues±2.859 miscues±5.912 miscuesstudents differ clearly, so ICC is meaningful
Word reading accuracy±2.624 %±4.786 %students are similar, so ICC understates agreement
Words Correct Per Minute±2.439 WCPM±17.593 WCPMstudents differ clearly, so ICC is meaningful

A useful rule: if the middle column is small in real-world terms, the teachers agreed in any way that matters to a child's assessment, whatever the ICC says.

Reading level (Independent / Instructional / Frustration)

  • Teachers picked the exact same level 40.0% of the time (25 recordings).
  • Kappa (agreement corrected for lucky guesses): 0.133.

Reading level is a category, not a number, so it uses percent agreement and kappa rather than ICC.

Why the level disagrees more than the numbers do

Phil-IRI turns a percentage into a label using hard cutoffs — 97% and above is Independent, below 90% is Frustration. A cutoff means a hair's-breadth difference in scoring can flip the label completely: 96.8% and 97.1% are practically the same reading performance but land in different categories.

  • Teachers disagreed on the level for 15 recordings.
  • 15 of those had a score sitting within 3 points of a cutoff.

So most level disagreements are boundary cases, not real disagreements about the child's reading. This is a property of the Phil-IRI instrument itself.

What this means for the app: showing a bare label like "Frustration" hides how close the call was. The interface should show the percentage alongside the label, and flag borderline results, so a teacher knows when their own judgement matters most.

Is any single teacher scoring harder or easier than the rest?

Positive = that teacher gives higher scores than their colleagues on the same recording. Negative = lower. Small numbers here mean nobody is an outlier.

Words read correctly

TeacherDifference vs. colleaguesBased on
T01+0.9 words5 shared recordings
T02+3.111 words9 shared recordings
T03-5.367 words15 shared recordings
T04+0.417 words12 shared recordings
T05+2.688 words16 shared recordings

Reading time

TeacherDifference vs. colleaguesBased on
T010.0 seconds6 shared recordings
T020.0 seconds9 shared recordings
T030.0 seconds16 shared recordings
T040.0 seconds12 shared recordings
T050.0 seconds16 shared recordings

Total miscues

TeacherDifference vs. colleaguesBased on
T01-1.25 miscues6 shared recordings
T02-2.778 miscues9 shared recordings
T03+4.188 miscues16 shared recordings
T04+0.917 miscues12 shared recordings
T05-2.844 miscues16 shared recordings

Word reading accuracy

TeacherDifference vs. colleaguesBased on
T01+1.083 %6 shared recordings
T02+3.0 %9 shared recordings
T03-5.1 %15 shared recordings
T04+0.8 %10 shared recordings
T05+2.333 %15 shared recordings

Words Correct Per Minute

TeacherDifference vs. colleaguesBased on
T01+0.732 WCPM6 shared recordings
T02+2.197 WCPM9 shared recordings
T03-3.966 WCPM16 shared recordings
T04+0.678 WCPM12 shared recordings
T05+1.948 WCPM16 shared recordings

What we do next because of this

  • Do not adjust or rescale teacher scores. Beyond the evidence above, each teacher shares only a handful of recordings with the others — far too few to estimate a reliable per-teacher correction. Inventing one would add noise, not remove it.
  • Use WCPM as the primary agreement measure. It is the most reliably scored number in this dataset, so it is the fairest yardstick for judging the system.
  • Report accuracy and miscue agreement honestly as a study limitation, with the narrow-spread explanation above so the numbers are not misread.
  • Do not make reading level the headline claim. The category is unstable near the Phil-IRI cutoffs even between two human teachers, so it is the wrong thing for an automated system to be judged on. Report the underlying percentage.

What we are measuring against (the "gold standard" answer)

Every claim about this system is of the form "the computer agreed with the teacher." That only means something if we are clear about what the teacher's score actually is. This study's position, stated plainly:

1. We compare against what teachers actually judged, not what was calculated for them. The independent human judgement in this dataset is words read correctly, and teachers agree on it strongly. Words-per-minute is arithmetic on top of that, and is reported as such.

2. A teacher's score is not a single true number — it is a range of reasonable professional judgement. Two qualified teachers hearing the same child write down somewhat different numbers. Pretending one of them is "correct" would be false precision.

3. So the system is judged by the right comparison: how far the computer sits from a teacher, against how far two teachers sit from each other.

For words read correctly, two teachers agree to within about ±12.029 words. That is the target. If the system lands inside that band, the honest and strongest available claim is:

The system disagrees with a teacher no more than two teachers disagree with each other.

That is a stronger result than any accuracy percentage, and this dataset supports it. Teacher disagreement stops being a weakness and becomes the yardstick the system is measured by.

Honest limitations

  • Only 25 recordings were double-scored, so these numbers carry real uncertainty. They indicate a direction, not a precise value.
  • Teachers were not assigned at random — each recording was scored by whoever was available, so this is a one-way (ICC1) reliability model, not a full crossed design.
  • Per-teacher differences are based on only a handful of shared recordings each and should not be used to judge any individual teacher.
  • Teacher names are replaced with T01, T02, ... in this file.

Generated by pipeline/teacher_agreement.py. Regenerate with uv run python teacher_agreement.py.

All findings