Model evaluation
Does it flag the right moments?
Inkling's flags scored against hand labels of where the student was confused. A flag counts when it lands within ±15 s of a label; each label matches at most one flag (nearest first).
Small labeled demo set — 2 sessions, 3 labels. A sanity check that the pipeline does what it claims, not a benchmark.
At the shipped threshold (0.55)
Hesitation spikes
The scoring model alone: 10 s windows whose smoothed score crossed the threshold.
- Precision
- 1.00
- Recall
- 0.33
- F1
- 0.50
Full pipeline (moments)
What the student sees: spikes plus erase → rewrite corrections paired from the ink.
- Precision
- 1.00
- Recall
- 1.00
- F1
- 1.00
| Detector | TP | FP | FN | Precision | Recall | F1 |
|---|---|---|---|---|---|---|
| Hesitation spikes | 1 | 0 | 2 | 1.00 | 0.33 | 0.50 |
| Full pipeline (moments) | 3 | 0 | 0 | 1.00 | 1.00 | 1.00 |
Label by label
| Label | Hesitation spike | Moment |
|---|---|---|
| 01:023x+1: stopped after the outer derivative, erased 2(3x+1)Maya — Session 1 | FNinside the baseline — never flagged by design | TP01:02 (±0 s) corrected |
| 01:40sin(x²): forgot the inner derivative, erased cos(x²)Maya — Session 1 | FNinside the baseline — never flagged by design | TP01:40 (±0 s) corrected |
| 04:06x²e^{3x}: product vs chain — slow attempt, erased, never rewrittenMaya — Session 1 | TP04:15 (+8.5 s) | TP04:10 (+3.5 s) unresolved gap |
Maya — Session 2: no labels (calm notes) — any flag there would be a false positive; none were raised.
Spikes can't fire while Inkling is still learning a student's normal (the first two minutes of their writing), so confusion that early is caught only by pairing an erase with its rewrite. That is why the full pipeline is the number students actually experience.
Sensitivity: moving the spike threshold
Every point re-runs the scoring model on the same ink and transcript with a different spikeScore. A spike also needs two signals on at once, so lowering the threshold alone adds few flags.
| Spike threshold | Precision | Recall | F1 | True positives | False positives | Missed |
|---|---|---|---|---|---|---|
| 0.35 | 1.00 | 0.33 | 0.50 | 1 | 0 | 2 |
| 0.40 | 1.00 | 0.33 | 0.50 | 1 | 0 | 2 |
| 0.45 | 1.00 | 0.33 | 0.50 | 1 | 0 | 2 |
| 0.50 | 1.00 | 0.33 | 0.50 | 1 | 0 | 2 |
| 0.55 (shipped) | 1.00 | 0.33 | 0.50 | 1 | 0 | 2 |
| 0.60 | 0.00 | 0.00 | 0.00 | 0 | 1 | 3 |
| 0.65 | — | 0.00 | — | 0 | 0 | 3 |
| 0.70 | — | 0.00 | — | 0 | 0 | 3 |
| 0.75 | — | 0.00 | — | 0 | 0 | 3 |
| 0.80 | — | 0.00 | — | 0 | 0 | 3 |
About the labels
Small hand-labeled demo set, not a benchmark. Each label is a lecture time where the student would have marked '?' in her notes. The labels are derived from the seeded Maya scenario's intended confusion moments in lib/demoScenario.ts (the file header and DEMO_TIMES), written down before looking at the detector's output: (1) 01:02 — she wrote only the outer derivative 2(3x+1) and erased it at DEMO_TIMES.breakthroughEraseMs (62000); (2) 01:40 — she wrote cos(x²) without the inner derivative and erased it at DEMO_TIMES.correctionEraseMs (100000); (3) ~04:07 — the product-vs-chain attempt x²e^{3x} starts at 246500 ms in buildMayaSession1 (slow writing, erased, never rewritten; inside DEMO_TIMES.gapFromMs..gapToMs). Session 2 is calm writing through the same material, so it carries no labels (every flag there is a false positive). A detection matches a label when it is within toleranceMs; matching is one-to-one, nearest pairs first. A hesitation spike's time is the centre of its 10 s scoring window.
Ink: the seeded sessions as stored in the database. Labels: public/demo/ground-truth.json · also served as JSON at /api/evaluation.
See the signals for Maya — Session 1