Model evaluation

Model evaluation

Does it flag the right moments?

Inkling's flags scored against hand labels of where the student was confused. A flag counts when it lands within ±15 s of a label; each label matches at most one flag (nearest first).

Small labeled demo set — 2 sessions, 3 labels. A sanity check that the pipeline does what it claims, not a benchmark.

At the shipped threshold (0.55)

Hesitation spikes

The scoring model alone: 10 s windows whose smoothed score crossed the threshold.

Precision
1.00
Recall
0.33
F1
0.50

Full pipeline (moments)

What the student sees: spikes plus erase → rewrite corrections paired from the ink.

Precision
1.00
Recall
1.00
F1
1.00
True positives, false positives and missed labels per detector
DetectorTPFPFNPrecisionRecallF1
Hesitation spikes1021.000.330.50
Full pipeline (moments)3001.001.001.00

Label by label

Each ground-truth label and the detection matched to it
LabelHesitation spikeMoment
01:023x+1: stopped after the outer derivative, erased 2(3x+1)Maya — Session 1FNinside the baseline — never flagged by designTP01:02 (±0 s) corrected
01:40sin(x²): forgot the inner derivative, erased cos(x²)Maya — Session 1FNinside the baseline — never flagged by designTP01:40 (±0 s) corrected
04:06x²e^{3x}: product vs chain — slow attempt, erased, never rewrittenMaya — Session 1TP04:15 (+8.5 s)TP04:10 (+3.5 s) unresolved gap

Maya — Session 2: no labels (calm notes) — any flag there would be a false positive; none were raised.

Spikes can't fire while Inkling is still learning a student's normal (the first two minutes of their writing), so confusion that early is caught only by pairing an erase with its rewrite. That is why the full pipeline is the number students actually experience.

Sensitivity: moving the spike threshold

Every point re-runs the scoring model on the same ink and transcript with a different spikeScore. A spike also needs two signals on at once, so lowering the threshold alone adds few flags.

PrecisionRecallGaps in the precision line: nothing was flagged at that threshold.
Precision and recall of hesitation spikes as the spike threshold goes from 0.35 to 0.80. Threshold 0.55 (shipped): precision 1.00, recall 0.33, F1 0.50; 1 true positives, 0 false positives, 2 missed.
Spike thresholdPrecisionRecallF1True positivesFalse positivesMissed
0.351.000.330.50102
0.401.000.330.50102
0.451.000.330.50102
0.501.000.330.50102
0.55 (shipped)1.000.330.50102
0.600.000.000.00013
0.65—0.00—003
0.70—0.00—003
0.75—0.00—003
0.80—0.00—003

About the labels

Small hand-labeled demo set, not a benchmark. Each label is a lecture time where the student would have marked '?' in her notes. The labels are derived from the seeded Maya scenario's intended confusion moments in lib/demoScenario.ts (the file header and DEMO_TIMES), written down before looking at the detector's output: (1) 01:02 — she wrote only the outer derivative 2(3x+1) and erased it at DEMO_TIMES.breakthroughEraseMs (62000); (2) 01:40 — she wrote cos(x²) without the inner derivative and erased it at DEMO_TIMES.correctionEraseMs (100000); (3) ~04:07 — the product-vs-chain attempt x²e^{3x} starts at 246500 ms in buildMayaSession1 (slow writing, erased, never rewritten; inside DEMO_TIMES.gapFromMs..gapToMs). Session 2 is calm writing through the same material, so it carries no labels (every flag there is a false positive). A detection matches a label when it is within toleranceMs; matching is one-to-one, nearest pairs first. A hesitation spike's time is the centre of its 10 s scoring window.

Ink: the seeded sessions as stored in the database. Labels: public/demo/ground-truth.json · also served as JSON at /api/evaluation.

See the signals for Maya — Session 1