Insights  /  Learning Science and AI  /  How Do You Know an AI Tutor Taught Anyone Anything?
Learning Science and AIEssayINTERACTIVE

How Do You Know an AI Tutor Taught Anyone Anything?

Every AI tutor dashboard can prove the tutor was used. Almost none can prove it taught. Usage, test scores and changed behaviour are three different claims with three different instruments, and only the third supports the word taught. This essay separates them, then lets you build the strongest claim your own evidence can actually defend.

Before you read · 30 seconds3 questions · adapts this page
1. In the four level evaluation model, a score on a test taken immediately after a lesson is evidence at…
2. A tutor reports 12,000 sessions and a 4.6 satisfaction score. The strongest claim that evidence supports is…
3. Two groups differ on an outcome after one used an AI tutor. What makes the difference attributable to the tutor?

What just happened, and why it matters: three retrieval questions estimated what you already know, and this page branched on the result. Note what it did not do. It did not ask how you like to read, or whether you prefer video. It assessed, then adapted to the assessment. That is the whole argument of this essay in miniature: a system that logs your preferences knows what you enjoy, and a system that assesses you knows what you can do. Only the second one is in a position to teach you, and only the second one can later show it did.

A learning platform can tell you, to the minute, how long a person spent inside it. It can tell you how many sessions they opened, how many modules they finished, how many questions they asked and how highly they rated the experience. What almost none of them can tell you is the only thing the organisation buying the platform actually wants to know: whether anyone now does their work differently, and whether the tutor is the reason.

The dashboard answers a different question

This is not a new failure and it is not specific to AI. A decade ago, in a paper whose title was already a warning, Dragan Gašević, Shane Dawson and George Siemens argued that learning analytics had drifted toward what was easy to instrument rather than what mattered, and that without a theory of learning behind them the counts quietly become the outcome.[1] Clicks are cheap to log. Minutes are cheap to log. Whether a person can now do something they could not do before is expensive to establish, so the field measured the cheap thing and reported it in the language of the expensive one.

Usage telemetry is not a fraud. It is a real measurement of a real quantity, and you need it: a tutor nobody opens cannot teach anybody anything, so usage is a necessary condition. The error is the slide from the necessary to the sufficient, from “our tutor was used for 40,000 hours this quarter” to “our tutor is working”. Those are different sentences supported by different evidence, and the dashboard is designed so that you never notice the gap between them.

Foundations · Why usage looks like evidence

Usage is a proxy, and proxies are legitimate when the link between the proxy and the thing you care about has been established. Nobody has established that link here. Time in a tutor can rise because the tutor is engaging, because it is confusing, because it is slow, or because a manager mandated it. Each of those produces the same number on the chart. A proxy with several incompatible explanations is not weak evidence of learning; it is not evidence of learning at all, until you find out which explanation is true.

The ladder nobody climbs

The standard way to separate these claims is nearly seventy years old. In 1959 Donald Kirkpatrick proposed four levels at which a training programme can be evaluated: reaction, what people thought of it; learning, what they can now do on a test; behaviour, what they do differently in the work; and results, what changed for the organisation.[2][3] The model is taught in every instructional design programme in the world. It is also, in practice, a ladder whose first rung takes almost all the traffic.

ASTD, now ATD, surveyed the profession on exactly this and found that 92 percent of organisations evaluate at least at level 1, reaction, while use of the model drops off dramatically at every level after that, with only 18 percent reaching return on investment.[4] The interesting number is not that one. It is the second one, which almost nobody quotes: when the same practitioners were asked how valuable each level actually is, level 1 came last. Only 36 percent rated reaction data as high or very high value, while 75 percent said that of behaviour and of results.[4]

Read those two findings together and you get a clean statement of the problem. The profession measures the level it trusts least, and trusts the levels it does not measure. That is not ignorance. It is what happens when the cost of an instrument, rather than the value of the claim it supports, decides what gets instrumented. Every AI tutor dashboard shipping in 2026 inherits that shape.

Fig. 1 · Interactive · What gets measured, what gets trustedDrag the slider
MEASURED Level 1 reaction · 92% Level 5 return on investment · 18% RATED HIGH VALUE BY PRACTITIONERS Level 1 reaction · 36% Levels 3 and 4 · 75% the most measured level is the least trusted one
10 400 50 programmes

Shares reported by ASTD with i4cp from a survey of the training profession [4]. The slider applies those shares to a portfolio of the size you choose. It is an illustration of proportion, not a forecast for a specific organisation.
Check yourself · 10 secondsRetrieval beats re-reading
A vendor shows you a chart of weekly active learners rising 40 percent. Which claim does that chart support?

Three claims, three instruments

Strip the dashboard back and there are only three families of evidence an AI tutor can generate, and each one licenses exactly one sentence. Usage telemetry licenses “it was used”. Assessment inside the learning flow licenses “they could do it, at that moment”. Observed change in the work afterwards licenses “it taught”. Taught is a level 3 word, and it is the only one that anyone outside the learning team is buying.

Fig. 2 · Three signals, three sentencesSame platform, different evidence
Usage telemetry sessions, minutes, completions “It was used.” Assessment in the flow items tied to the objective “They could do it.” Behaviour in the work what changes afterwards “It taught.” the reporting shortcut most dashboards take Usage is logged. It is not learning. It must never be reported as learning.
Three instruments, three sentences. The dotted line is the move to watch for in any vendor deck: evidence collected in the left lane, presented in the language of the right one.

Attribution is the second half of the problem. Even a change in the work does not belong to the tutor unless something rules out the alternatives: the people who chose to use it were the motivated ones, the team also got a new manager, the quarter was easier. The instrument that rules those out is not a bigger sample. It is assignment the learner does not control, which is why the design question, who gets the tutor and who does not, has to be settled before the first user logs in. After launch it is too late; there is no comparison to construct.

Interactive · The strongest claim you can defendSelect the evidence you actually hold
Select the evidence you hold and this will write the strongest sentence that evidence supports.
The logic is the evaluation model itself: each instrument licenses one level of claim, and a comparison group is what turns a difference into an attribution.

The levels are not a staircase

There is a tempting shortcut here, and the literature closed it thirty years ago. If the levels are a ladder, surely a good score on the bottom rung predicts the higher ones, so a happy sheet can stand in for evidence you did not gather. George Alliger and Elizabeth Janak examined that assumption directly and found that the model’s causal and correlational claims had simply never been established.[5] The later meta-analysis by Alliger and colleagues put numbers on it: the correlations among training criteria are weak, and reaction measures in particular are poor predictors of whether anyone learned anything.[6] People can enjoy a course that taught them nothing, and resent one that changed how they work.

So the levels are not a staircase you can climb by inference. They are four separate questions, each needing its own instrument, and skipping one means you do not have its answer. That is the discipline, and the strongest evidence in the field respects it. When Kestin and colleagues ran a randomised controlled trial of an AI tutor against active learning in a Harvard physics course, they reported that students learned significantly more in less time with the tutor, with a median of 49 minutes on task.[7] It is a careful, well designed study and it is exactly the level 2 evidence its instrument can support: the learning was measured by tests taken immediately after each lesson. The authors do not claim more than that. The vendors quoting them routinely do.

A tutor is judged by what the learner does once the tutor is gone. Every instrument that stops before that point is measuring something else, and should say so.
Check yourself · 10 secondsSection 4 of 6
Your tutor scores 4.7 out of 5 on satisfaction. What does the evidence say that predicts about learning?

What I instrument instead

This is the measurement design I have built into ALI, the adaptive learning platform behind my doctoral work at the University of Cyprus, and it is deliberately boring. Three signal classes are kept separate and never merged into a single score. Usage telemetry is logged in full and is never reported as learning. Assessment evidence is generated inside the learning flow, on items written backwards from the behaviour each tutorial is meant to change, which is constructive alignment doing its ordinary job.[8] Behaviour is the change in how people work afterwards, and it is the only signal permitted to carry the word taught.

Two design rules make the third one obtainable rather than aspirational. First, arms are randomised, so that a difference has somewhere to come from; the comparison is built before launch, because it cannot be recovered later. Second, the platform hosts the teaching and the subsequent work, so the behavioural evidence is a record of what people actually did rather than a survey asking them to recall it. None of this is technically hard. It is a decision that has to be made early, and the reason so few dashboards can answer the question in the title is that the decision was never made at all.

Exercise · Read a real dashboard

Below are four lines from AI tutor reporting decks. For each, decide what the evidence licenses: usage, learning, or behaviour. Then take the three question audit into your next vendor call.

Classify each claimUsage · Learning · Behaviour
“Learners completed 12,400 tutoring sessions this quarter, up 38 percent, with an average satisfaction of 4.6.”
“Learners scored 31 percent higher on the end of module test than a randomly assigned group taught the same content in class.”
“Six weeks after the tutorial, the trained group’s campaign briefs contain acceptance criteria in 78 percent of cases, against 24 percent in the untrained group.”
“92 percent of learners said they felt more confident applying the skill at work.”
The audit: 1) Which of your numbers would change if the tutor taught nobody anything? 2) What is your comparison group, and was it decided before launch? 3) Name the behaviour in the work that should change, and say how you would see it.
What this evidence does not prove
The ASTD figures are self reported practice from the corporate training profession and they are old.[4] That the default AI tutor dashboard has the same shape today is an inference from the products, not a measured finding, and you should treat it as one. The Kestin trial is a single course at one university, with tests taken immediately after each lesson and no measure of retention weeks later or transfer to other work.[7] It shows an AI tutor teaching more in less time under those conditions, which is a real and useful result, and it is not a behaviour claim. Nothing here shows that AI tutors fail at level 3. It shows that almost nobody has looked, which is a different and more fixable problem.
Key takeaways
  1. Usage, test score and changed behaviour are three claims with three instruments. Only the third supports the word taught [2][3].
  2. The profession measures the level it trusts least: 92 percent evaluate reaction, only 36 percent think reaction data is valuable, and behaviour and results are rated most valuable by 75 percent [4].
  3. The levels are not a staircase. Reaction predicts learning poorly, so a satisfaction score cannot stand in for evidence you did not collect [5][6].
  4. Attribution is a design decision, not an analysis technique. The comparison group has to exist before the first learner logs in.

The follow-up quiz arrives in three days.

Spaced retrieval is the best evidenced way to keep what you just read. Leave an email and you will get three recall questions on this essay in three days, plus the next essay when it publishes. That is the whole mechanism; unsubscribe any time.

References

Primary and peer reviewed sources only. Links go to publisher pages.

  1. Gašević, D., Dawson, S., and Siemens, G. (2015). Let’s not forget: Learning analytics are about learning. TechTrends, 59(1), 64-71. link.springer.com/article/10.1007/s11528-014-0822-x
  2. Kirkpatrick, D. L. (1959). Techniques for evaluating training programs. Journal of the American Society of Training Directors, 13.
  3. Kirkpatrick, J. D., and Kirkpatrick, W. K. (2016). Kirkpatrick’s Four Levels of Training Evaluation. ATD Press. td.org
  4. ASTD with i4cp (2009). The Value of Evaluation: Making Training Evaluations More Effective. Findings reported by ASTD. td.org
  5. Alliger, G. M., and Janak, E. A. (1989). Kirkpatrick’s levels of training criteria: thirty years later. Personnel Psychology, 42(2), 331-342. onlinelibrary.wiley.com
  6. Alliger, G. M., Tannenbaum, S. I., Bennett, W., Traver, H., and Shotland, A. (1997). A meta-analysis of the relations among training criteria. Personnel Psychology, 50(2), 341-358. onlinelibrary.wiley.com
  7. Kestin, G., Miller, K., Klales, A., Milbourne, T., and Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. nature.com/articles/s41598-025-97652-6
  8. Biggs, J., and Tang, C. (2011). Teaching for Quality Learning at University. Open University Press.
MC

Manolis Charalampous

PhD candidate and teaching assistant at the University of Cyprus, researching AI driven marketing performance and adaptive learning. He leads group marketing across 160+ companies in 23+ locations, with twelve years across maritime, telecommunications, legal and finance, travel and gaming. He writes about applied AI, learning science and marketing performance at m4no5.com.

COPIED