AI-assisted teaching should measure skill transfer, not answer throughput
AI-assisted teaching should measure skill transfer, not answer throughput. A student completing more exercises with an assistant may be learning faster, delegating more, or merely accepting plausible output. Completion counts cannot distinguish those cases. The real test is whether the learner can retrieve, adapt, and compose the underlying skills when the surface form changes and the assistant’s support is reduced.
Cognitive capacity makes assistance tempting. Working memory holds only a small number of chunks; well-designed support can remove incidental load and preserve attention for the structure of a problem. But the same support can hide whether the learner formed reusable chunks at all. If the assistant supplies the decomposition, selects the method, and checks the result, the student may produce a correct answer without building the retrieval cues or control knowledge needed next time.
I therefore separate performance during assistance from transfer after assistance. The first can measure whether the tool helps the joint system complete work. The second asks whether the person can solve a novel problem that requires the same primitive in a different composition. Change the domain language, withhold the familiar template, alter the order of subskills, or introduce a misleading cue. If performance collapses, the training optimized local fluency rather than portable competence.
The compositional-learning literature offers the right model. Skills are not a flat inventory; they have prerequisites, preconditions, postconditions, and dependency edges. Mastering two components separately does not prove that a learner can select and join them under pressure. Benchmarks such as SCAN, CLEVR, and BabyAI are valuable precisely because their splits test systematic recombination rather than memorized instances. Teaching should borrow that discipline: grade unfamiliar combinations, not just familiar components.
This changes how an AI tutor should intervene. Early support can make the task tractable, but it should fade along named dimensions: fewer hints, less decomposition, weaker retrieval cues, and delayed verification. The tutor should sometimes ask the learner to predict the next step before revealing it, explain why a tempting alternative fails, and reconstruct the solution after a gap. Progressive complexity works when the learner owns more of the control loop over time, not when the assistant quietly absorbs every harder decision.
Evaluation needs delayed and counterfactual measures. Re-test after time has passed, vary the context, and distinguish recognition from generation. Record which hint level was required, whether the learner detected an error, and whether they could transfer the method without the original vocabulary. Evaluation literacy compounds because learners who can judge an answer become less dependent on fluent output. Teaching judgment rather than tools is what lets that ability survive the next interface change.
There is one precise concession: answer throughput is a legitimate operational metric when the immediate objective is supported performance—accessibility, time-sensitive work, or a task that will always remain human-plus-AI. It can show whether the tool reduces friction. It becomes misleading only when presented as evidence that unaided or differently aided skill has transferred.
I would rather see fewer assisted answers followed by successful novel composition than a dashboard full of completed exercises that disappear when the prompt changes. Calibration and escalation are part of the same outcome: knowing when a skill applies, when confidence is warranted, and when to seek help. AI-assisted teaching earns its name when capability moves into the learner, not merely through the learner’s screen.