How Do You Know an AI Tutor Taught Anyone Anything?
Every AI tutor dashboard can prove the tutor was used. Almost none can prove it taught. Usage, test scores and changed behaviour are three different claims with three different instruments, and only the third supports the word taught. This essay separates them, then lets you build the strongest claim your own evidence can actually defend.
A learning platform can tell you, to the minute, how long a person spent inside it. It can tell you how many sessions they opened, how many modules they finished, how many questions they asked and how highly they rated the experience. What almost none of them can tell you is the only thing the organisation buying the platform actually wants to know: whether anyone now does their work differently, and whether the tutor is the reason.
The dashboard answers a different question
This is not a new failure and it is not specific to AI. A decade ago, in a paper whose title was already a warning, Dragan Gašević, Shane Dawson and George Siemens argued that learning analytics had drifted toward what was easy to instrument rather than what mattered, and that without a theory of learning behind them the counts quietly become the outcome.[1] Clicks are cheap to log. Minutes are cheap to log. Whether a person can now do something they could not do before is expensive to establish, so the field measured the cheap thing and reported it in the language of the expensive one.
Usage telemetry is not a fraud. It is a real measurement of a real quantity, and you need it: a tutor nobody opens cannot teach anybody anything, so usage is a necessary condition. The error is the slide from the necessary to the sufficient, from “our tutor was used for 40,000 hours this quarter” to “our tutor is working”. Those are different sentences supported by different evidence, and the dashboard is designed so that you never notice the gap between them.
Foundations · Why usage looks like evidence
Usage is a proxy, and proxies are legitimate when the link between the proxy and the thing you care about has been established. Nobody has established that link here. Time in a tutor can rise because the tutor is engaging, because it is confusing, because it is slow, or because a manager mandated it. Each of those produces the same number on the chart. A proxy with several incompatible explanations is not weak evidence of learning; it is not evidence of learning at all, until you find out which explanation is true.
The ladder nobody climbs
The standard way to separate these claims is nearly seventy years old. In 1959 Donald Kirkpatrick proposed four levels at which a training programme can be evaluated: reaction, what people thought of it; learning, what they can now do on a test; behaviour, what they do differently in the work; and results, what changed for the organisation.[2][3] The model is taught in every instructional design programme in the world. It is also, in practice, a ladder whose first rung takes almost all the traffic.
ASTD, now ATD, surveyed the profession on exactly this and found that 92 percent of organisations evaluate at least at level 1, reaction, while use of the model drops off dramatically at every level after that, with only 18 percent reaching return on investment.[4] The interesting number is not that one. It is the second one, which almost nobody quotes: when the same practitioners were asked how valuable each level actually is, level 1 came last. Only 36 percent rated reaction data as high or very high value, while 75 percent said that of behaviour and of results.[4]
Read those two findings together and you get a clean statement of the problem. The profession measures the level it trusts least, and trusts the levels it does not measure. That is not ignorance. It is what happens when the cost of an instrument, rather than the value of the claim it supports, decides what gets instrumented. Every AI tutor dashboard shipping in 2026 inherits that shape.
Three claims, three instruments
Strip the dashboard back and there are only three families of evidence an AI tutor can generate, and each one licenses exactly one sentence. Usage telemetry licenses “it was used”. Assessment inside the learning flow licenses “they could do it, at that moment”. Observed change in the work afterwards licenses “it taught”. Taught is a level 3 word, and it is the only one that anyone outside the learning team is buying.
Attribution is the second half of the problem. Even a change in the work does not belong to the tutor unless something rules out the alternatives: the people who chose to use it were the motivated ones, the team also got a new manager, the quarter was easier. The instrument that rules those out is not a bigger sample. It is assignment the learner does not control, which is why the design question, who gets the tutor and who does not, has to be settled before the first user logs in. After launch it is too late; there is no comparison to construct.
The levels are not a staircase
There is a tempting shortcut here, and the literature closed it thirty years ago. If the levels are a ladder, surely a good score on the bottom rung predicts the higher ones, so a happy sheet can stand in for evidence you did not gather. George Alliger and Elizabeth Janak examined that assumption directly and found that the model’s causal and correlational claims had simply never been established.[5] The later meta-analysis by Alliger and colleagues put numbers on it: the correlations among training criteria are weak, and reaction measures in particular are poor predictors of whether anyone learned anything.[6] People can enjoy a course that taught them nothing, and resent one that changed how they work.
So the levels are not a staircase you can climb by inference. They are four separate questions, each needing its own instrument, and skipping one means you do not have its answer. That is the discipline, and the strongest evidence in the field respects it. When Kestin and colleagues ran a randomised controlled trial of an AI tutor against active learning in a Harvard physics course, they reported that students learned significantly more in less time with the tutor, with a median of 49 minutes on task.[7] It is a careful, well designed study and it is exactly the level 2 evidence its instrument can support: the learning was measured by tests taken immediately after each lesson. The authors do not claim more than that. The vendors quoting them routinely do.
What I instrument instead
This is the measurement design I have built into ALI, the adaptive learning platform behind my doctoral work at the University of Cyprus, and it is deliberately boring. Three signal classes are kept separate and never merged into a single score. Usage telemetry is logged in full and is never reported as learning. Assessment evidence is generated inside the learning flow, on items written backwards from the behaviour each tutorial is meant to change, which is constructive alignment doing its ordinary job.[8] Behaviour is the change in how people work afterwards, and it is the only signal permitted to carry the word taught.
Two design rules make the third one obtainable rather than aspirational. First, arms are randomised, so that a difference has somewhere to come from; the comparison is built before launch, because it cannot be recovered later. Second, the platform hosts the teaching and the subsequent work, so the behavioural evidence is a record of what people actually did rather than a survey asking them to recall it. None of this is technically hard. It is a decision that has to be made early, and the reason so few dashboards can answer the question in the title is that the decision was never made at all.
Exercise · Read a real dashboard
Below are four lines from AI tutor reporting decks. For each, decide what the evidence licenses: usage, learning, or behaviour. Then take the three question audit into your next vendor call.
- Usage, test score and changed behaviour are three claims with three instruments. Only the third supports the word taught [2][3].
- The profession measures the level it trusts least: 92 percent evaluate reaction, only 36 percent think reaction data is valuable, and behaviour and results are rated most valuable by 75 percent [4].
- The levels are not a staircase. Reaction predicts learning poorly, so a satisfaction score cannot stand in for evidence you did not collect [5][6].
- Attribution is a design decision, not an analysis technique. The comparison group has to exist before the first learner logs in.
The follow-up quiz arrives in three days.
Spaced retrieval is the best evidenced way to keep what you just read. Leave an email and you will get three recall questions on this essay in three days, plus the next essay when it publishes. That is the whole mechanism; unsubscribe any time.
References
Primary and peer reviewed sources only. Links go to publisher pages.
- Gašević, D., Dawson, S., and Siemens, G. (2015). Let’s not forget: Learning analytics are about learning. TechTrends, 59(1), 64-71. link.springer.com/article/10.1007/s11528-014-0822-x
- Kirkpatrick, D. L. (1959). Techniques for evaluating training programs. Journal of the American Society of Training Directors, 13.
- Kirkpatrick, J. D., and Kirkpatrick, W. K. (2016). Kirkpatrick’s Four Levels of Training Evaluation. ATD Press. td.org
- ASTD with i4cp (2009). The Value of Evaluation: Making Training Evaluations More Effective. Findings reported by ASTD. td.org
- Alliger, G. M., and Janak, E. A. (1989). Kirkpatrick’s levels of training criteria: thirty years later. Personnel Psychology, 42(2), 331-342. onlinelibrary.wiley.com
- Alliger, G. M., Tannenbaum, S. I., Bennett, W., Traver, H., and Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology, 50(2), 341-358. onlinelibrary.wiley.com
- Kestin, G., Miller, K., Klales, A., Milbourne, T., and Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. nature.com/articles/s41598-025-97652-6
- Biggs, J., and Tang, C. (2011). Teaching for Quality Learning at University. Open University Press.