Finding 01
Finding 1 — Do the teachers agree with each other?
In one sentence: On at least one measure the teachers' numbers genuinely differ. Human scores are noisy there, which caps how well any ASR system can appear to agree.
Why this matters
Every score this system produces is judged against a teacher's score. So before asking "does the computer agree with the teacher?" we have to ask "do teachers agree with each other?" If two teachers listening to the same child write down very different numbers, then there is no single correct answer to compare against, and no amount of engineering can fix that.
To check this, some recordings were deliberately scored by more than one teacher.
What was measured
- 204 scoring sheets in total, covering 167 recordings
- 25 of those recordings were scored by two or more different teachers
- 5 teachers took part
- Breakdown: 18 recording(s) scored by 2 teachers, 7 recording(s) scored by 3 teachers
The headline numbers
The score below is called an ICC. Read it like a percentage of "how much of the difference between scores is real difference between students, rather than disagreement between teachers." 1.00 is perfect agreement. Above 0.75 is good.
The range matters more than the score itself. With only 25 double-scored recordings, a single number is not enough to trust — the range says how sure we can be. A range that spans both "bad" and "good" means we cannot tell yet.
| What was scored | Agreement (ICC) | How sure (95% range) | Verdict | Two teachers usually within |
|---|---|---|---|---|
| Words read correctly | 0.980 | 0.95 – 0.99 | excellent | ±12.029 words |
| Reading time | 1.000 | 1.00 – 1.00 | excellent | ±0.0 seconds |
| Total miscues | 0.588 | 0.14 – 0.77 | too uncertain to call | ±11.409 miscues |
| Word reading accuracy | 0.427 | -0.09 – 0.71 | too uncertain to call | ±11.734 % |
| Words Correct Per Minute (derived) | 0.960 | 0.90 – 0.98 | excellent | ±9.553 WCPM |
The last column is the plain-English one: 95% of the time, two teachers scoring the same recording land within that much of each other.
⚠️ Reading time shows perfect agreement with zero variation. That almost never happens with independent human judgement — it means the value was recorded once and shared between the scorers, not judged separately by each. It cannot be counted as evidence that teachers agree.
⚠️ Words Correct Per Minute is calculated from other fields on the sheet, not written down independently. Teachers who agree on the inputs agree on it automatically, so its high score is inherited, not independent confirmation.
Words read correctly ███████████████████████████· 0.980
Reading time ████████████████████████████ 1.000
Total miscues ████████████████············ 0.588
Word reading accuracy ████████████················ 0.427
Words Correct Per Minute ███████████████████████████· 0.960
Important: a low score here does not always mean teachers disagreed
ICC is a ratio. It compares how much students differ from each other against how much teachers differ from each other. So if all the students score similarly on a measure, ICC drops — even when the two teachers wrote down almost the same number. The table below separates those two situations.
| What was scored | How far apart the teachers were | How far apart the students were | Reading |
|---|---|---|---|
| Words read correctly | ±3.154 words | ±31.094 words | students differ clearly, so ICC is meaningful |
| Reading time | ±0.0 seconds | ±42.814 seconds | students differ clearly, so ICC is meaningful |
| Total miscues | ±2.859 miscues | ±5.912 miscues | students differ clearly, so ICC is meaningful |
| Word reading accuracy | ±2.624 % | ±4.786 % | students are similar, so ICC understates agreement |
| Words Correct Per Minute | ±2.439 WCPM | ±17.593 WCPM | students differ clearly, so ICC is meaningful |
A useful rule: if the middle column is small in real-world terms, the teachers agreed in any way that matters to a child's assessment, whatever the ICC says.
Reading level (Independent / Instructional / Frustration)
- Teachers picked the exact same level 40.0% of the time (25 recordings).
- Kappa (agreement corrected for lucky guesses): 0.133.
Reading level is a category, not a number, so it uses percent agreement and kappa rather than ICC.
Why the level disagrees more than the numbers do
Phil-IRI turns a percentage into a label using hard cutoffs — 97% and above is Independent, below 90% is Frustration. A cutoff means a hair's-breadth difference in scoring can flip the label completely: 96.8% and 97.1% are practically the same reading performance but land in different categories.
- Teachers disagreed on the level for 15 recordings.
- 15 of those had a score sitting within 3 points of a cutoff.
So most level disagreements are boundary cases, not real disagreements about the child's reading. This is a property of the Phil-IRI instrument itself.
What this means for the app: showing a bare label like "Frustration" hides how close the call was. The interface should show the percentage alongside the label, and flag borderline results, so a teacher knows when their own judgement matters most.
Is any single teacher scoring harder or easier than the rest?
Positive = that teacher gives higher scores than their colleagues on the same recording. Negative = lower. Small numbers here mean nobody is an outlier.
Words read correctly
| Teacher | Difference vs. colleagues | Based on |
|---|---|---|
| T01 | +0.9 words | 5 shared recordings |
| T02 | +3.111 words | 9 shared recordings |
| T03 | -5.367 words | 15 shared recordings |
| T04 | +0.417 words | 12 shared recordings |
| T05 | +2.688 words | 16 shared recordings |
Reading time
| Teacher | Difference vs. colleagues | Based on |
|---|---|---|
| T01 | 0.0 seconds | 6 shared recordings |
| T02 | 0.0 seconds | 9 shared recordings |
| T03 | 0.0 seconds | 16 shared recordings |
| T04 | 0.0 seconds | 12 shared recordings |
| T05 | 0.0 seconds | 16 shared recordings |
Total miscues
| Teacher | Difference vs. colleagues | Based on |
|---|---|---|
| T01 | -1.25 miscues | 6 shared recordings |
| T02 | -2.778 miscues | 9 shared recordings |
| T03 | +4.188 miscues | 16 shared recordings |
| T04 | +0.917 miscues | 12 shared recordings |
| T05 | -2.844 miscues | 16 shared recordings |
Word reading accuracy
| Teacher | Difference vs. colleagues | Based on |
|---|---|---|
| T01 | +1.083 % | 6 shared recordings |
| T02 | +3.0 % | 9 shared recordings |
| T03 | -5.1 % | 15 shared recordings |
| T04 | +0.8 % | 10 shared recordings |
| T05 | +2.333 % | 15 shared recordings |
Words Correct Per Minute
| Teacher | Difference vs. colleagues | Based on |
|---|---|---|
| T01 | +0.732 WCPM | 6 shared recordings |
| T02 | +2.197 WCPM | 9 shared recordings |
| T03 | -3.966 WCPM | 16 shared recordings |
| T04 | +0.678 WCPM | 12 shared recordings |
| T05 | +1.948 WCPM | 16 shared recordings |
What we do next because of this
- Do not adjust or rescale teacher scores. Beyond the evidence above, each teacher shares only a handful of recordings with the others — far too few to estimate a reliable per-teacher correction. Inventing one would add noise, not remove it.
- Use WCPM as the primary agreement measure. It is the most reliably scored number in this dataset, so it is the fairest yardstick for judging the system.
- Report accuracy and miscue agreement honestly as a study limitation, with the narrow-spread explanation above so the numbers are not misread.
- Do not make reading level the headline claim. The category is unstable near the Phil-IRI cutoffs even between two human teachers, so it is the wrong thing for an automated system to be judged on. Report the underlying percentage.
What we are measuring against (the "gold standard" answer)
Every claim about this system is of the form "the computer agreed with the teacher." That only means something if we are clear about what the teacher's score actually is. This study's position, stated plainly:
1. We compare against what teachers actually judged, not what was calculated for them. The independent human judgement in this dataset is words read correctly, and teachers agree on it strongly. Words-per-minute is arithmetic on top of that, and is reported as such.
2. A teacher's score is not a single true number — it is a range of reasonable professional judgement. Two qualified teachers hearing the same child write down somewhat different numbers. Pretending one of them is "correct" would be false precision.
3. So the system is judged by the right comparison: how far the computer sits from a teacher, against how far two teachers sit from each other.
For words read correctly, two teachers agree to within about ±12.029 words. That is the target. If the system lands inside that band, the honest and strongest available claim is:
The system disagrees with a teacher no more than two teachers disagree with each other.
That is a stronger result than any accuracy percentage, and this dataset supports it. Teacher disagreement stops being a weakness and becomes the yardstick the system is measured by.
Honest limitations
- Only 25 recordings were double-scored, so these numbers carry real uncertainty. They indicate a direction, not a precise value.
- Teachers were not assigned at random — each recording was scored by whoever was available, so this is a one-way (ICC1) reliability model, not a full crossed design.
- Per-teacher differences are based on only a handful of shared recordings each and should not be used to judge any individual teacher.
- Teacher names are replaced with T01, T02, ... in this file.
Generated by pipeline/teacher_agreement.py. Regenerate with uv run python teacher_agreement.py.