Insights  /  Learning Science × AI  /  Adaptive Learning Is Not Personalisation
Learning Science × AIEssay 01INTERACTIVE

Adaptive Learning Is Not Personalisation

Netflix personalises. It does not teach. The distinction sounds pedantic until you watch an organisation buy a recommendation engine and expect a tutor. This essay shows you the evidence, lets you handle the statistics yourself, and ends with a vendor audit you can run this week.

Before you read · 30 seconds3 questions · adapts this page
1. The defining feature of a genuinely adaptive learning system is that it…
2. A course completion rate primarily measures…
3. Which condition tends to produce the most durable learning?

What just happened, and why it matters: three retrieval questions estimated your prior knowledge and this page branched on the result. A rule that small is still a learner model: assessment first, then adaptation to competence. If we had asked what you enjoy reading and reordered the page to suit, that would have been personalisation. You have just experienced the entire argument of this essay in miniature.

Netflix personalises. It does not teach. Its recommender system, by its engineers’ own account, is worth over one billion dollars a year in retention[1], and it earns that by showing you more of what you already like. A tutoring system does close to the opposite: it works out what you cannot yet do, and holds you there, deliberately, until you can. Both adapt to the individual. Only one of them produces learning.

The conflation

In enterprise learning and development, these two mechanisms are routinely sold under one label. Adaptive platforms are marketed in the language of streaming: content you will love, journeys tailored to you, engagement as the headline metric. The purchasing logic transfers from consumer software because the buyers live in consumer software. The result is a market where the word adaptive can mean either a recommendation engine or an intelligent tutor, two technologies with forty years of divergent evidence behind them.

Foundations · What is an effect size, in one minute

An effect size (Cohen’s d) expresses how far an intervention moves a group, measured in standard deviations of the outcome. It lets studies with different tests be compared on one scale. Rough benchmarks: 0.2 is small, 0.5 is medium, 0.8 is large. In education research, interventions above 0.6 that replicate across dozens of studies are rare, which is why the tutoring numbers below matter. The chart in Figure 1 lets you feel what these numbers mean rather than take them on faith.

Handle the evidence yourself

VanLehn’s landmark review found intelligent tutoring systems produced an effect size of 0.79, statistically indistinguishable from the 0.76 measured for expert human tutors[2]. Kulik and Fletcher’s meta analysis across fifty evaluations put the median at 0.66[3]. Numbers like these are usually cited and skimmed. Instead, drag the slider: the curves show a comparison group and a tutored group, and the readout translates the gap into plain language.

Fig. 1 · Interactive · What an effect size feels likeDrag the slider
comparison group tutored group
d = 0 1.2 d = 0.79

Normal distributions of equal spread, separated by d standard deviations. The percentile translation uses Cohen’s U₃, the share of the comparison group the average tutored learner exceeds. Assumes normality and equal variance; real study distributions vary.
Check yourself · 10 secondsRetrieval beats re-reading [6]
A vendor’s case study reports learners at the 50th percentile moving to roughly the 66th after adoption. Which effect size is that closest to?

Two loops, two goals

The mechanisms differ at the level of the feedback loop. A recommender infers what you like from what you consume, then serves more of it; its success condition is that you stay. A tutor infers what you know from your attempts, then serves what you cannot quite do yet; its success condition is that its own predictions become obsolete because you improved. Corbett and Anderson formalised that second loop as knowledge tracing three decades ago[5], and it remains the working core of every serious tutoring system.

Fig. 2 · The two loopsSame word, opposite objectives
RECOMMENDER LOOP You watch something System infers preference Serves more of the same optimises: you never leave TUTOR LOOP You attempt a task System updates knowledge model Targets the edge of competence optimises: you outgrow it
Both loops adapt to the individual. The recommender’s objective function rewards continued consumption; the tutor’s rewards competence gain, formalised as knowledge tracing [5]. Everything else in this essay follows from this difference.
A tutor’s job is to find the edge of your competence and keep you working at it. A recommender’s job is to make sure you never leave.
Foundations · Why preference and learning diverge

Pashler and colleagues, commissioned to audit the learning styles literature, found no adequate evidence that matching instruction to stated preferences improves outcomes[4]. The Bjork laboratory’s work on desirable difficulties goes further: conditions learners prefer, such as smooth, easy, familiar practice, are frequently the conditions that produce the least durable learning, while effortful conditions like spacing, interleaving and retrieval feel worse and work better[6]. Preference is a real signal; it is simply a signal about enjoyment, not about growth.

Where personalisation fails as a learning strategy

The failure is visible in the numbers the industry itself publishes. LinkedIn’s Workplace Learning Report puts course completion around 42 percent[7], presented as an engagement benchmark. Read as an instructional statistic, completion is a consumption metric: it records that content was watched, not that capability changed. MOOCs made the same substitution a decade earlier; their single digit completion rates[8] became the canonical cautionary tale, yet the successor metric is still watching. When the measure is consumption, the optimisation target becomes watchability, and the loop on the left of Figure 2 quietly replaces the one on the right.

Check yourself · 10 secondsSection 4 of 6
Your platform reports 42% completion and 4.6/5 satisfaction. What do you now know about learning?

Exercise · The vendor demo

The skill this essay wants to leave you with is separation: hearing a claim and knowing which loop it belongs to. Below are four statements assembled from real vendor language. Classify each one, then take the checklist into your next procurement meeting.

Classify each claimTutor signal or personalisation signal
“Learners love it. Average session time is up 40% in the first month, and weekly active usage keeps climbing.”
“Here is the mastery dashboard. Each skill carries a probability estimate of whether the learner has it, updated after every attempt.”
“Our AI recommends each person’s next course based on what similar learners in similar roles enjoyed and finished.”
“Item difficulty adjusts so each learner keeps succeeding on roughly 8 in 10 attempts, right at the edge of their current ability.”
The checklist: 1) What happens when the learner gets something right easily? 2) Show me the model of what the learner currently knows. 3) In your best case study, which moved: time on platform, or verified capability?

What this means in practice

None of this argues against personalisation, which does exactly what it was built to do. It argues against buying one mechanism and expecting the outcomes of the other. An organisation that wants engagement should buy the recommender and measure watching. An organisation that wants capability should demand the learner model, accept that effective practice will sometimes feel harder than watching[6], and measure what people can do afterwards. The research permits both purchases. It does not permit confusing them.

What this evidence does not prove
Effect sizes from the tutoring literature come mostly from STEM domains with well-structured problems; transfer to soft-skill training is plausible but less studied. VanLehn’s and Kulik and Fletcher’s reviews aggregate heterogeneous comparison conditions, and meta analyses inherit any publication bias in their fields. Pashler’s review establishes an absence of adequate evidence for preference matching, which is not the same as proof of zero effect. Strong claims deserve these footnotes; the direction of the evidence, however, is not close.
Key takeaways
  1. Recommendation engines optimise preference; tutoring systems optimise competence. Both adapt; the objective functions are opposed.
  2. Intelligent tutoring matches human tutors in rigorous reviews (d ≈ 0.79 vs 0.76); preference matching has no comparable evidence base.
  3. Completion, satisfaction and session time are consumption metrics. To claim learning, demand a capability observation.
  4. In any demo, ask the three checklist questions. The answer to the third usually ends the meeting.

The follow-up quiz arrives in three days.

Spaced retrieval is the best-evidenced way to keep what you just read. Leave an email and you will get three recall questions on this essay in three days, plus the next essay when it publishes. That is the whole mechanism; unsubscribe any time.

References

Primary and peer reviewed sources only. Links go to publisher pages.

  1. Gomez-Uribe, C. A., & Hunt, N. (2016). The Netflix Recommender System: Algorithms, Business Value, and Innovation. ACM Transactions on Management Information Systems, 6(4). dl.acm.org/doi/10.1145/2843948
  2. VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist, 46(4), 197–221. tandfonline.com/doi/abs/10.1080/00461520.2011.611369
  3. Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research, 86(1), 42–78. journals.sagepub.com/doi/10.3102/0034654315581420
  4. Pashler, H., McDaniel, M., Rohrer, D., & Bjork, R. (2008). Learning Styles: Concepts and Evidence. Psychological Science in the Public Interest, 9(3), 105–119. journals.sagepub.com/doi/10.1111/j.1539-6053.2009.01038.x
  5. Corbett, A. T., & Anderson, J. R. (1994). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4, 253–278. link.springer.com/article/10.1007/BF01099821
  6. Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the Real World. Worth Publishers. bjorklab.psych.ucla.edu/research
  7. LinkedIn Learning (2025). Workplace Learning Report 2025. learning.linkedin.com/resources/workplace-learning-report
  8. Jordan, K. (2014). Initial trends in enrolment and completion of massive open online courses. The International Review of Research in Open and Distributed Learning, 15(1). irrodl.org/index.php/irrodl/article/view/1651
MC

Manolis Charalampous

Group Marketing Manager at Fameline Holding Group, directing marketing across 160+ companies in 23+ locations. PhD candidate and teaching assistant at the University of Cyprus, researching AI driven marketing performance and adaptive learning. Twelve years across maritime, telecommunications, legal & finance, travel and gaming.

COPIED