Insights  /  Learning Science and AI  /  The Job Moved from Prompting to Delegating
Learning Science and AIEssayINTERACTIVE

The Job Moved from Prompting to Delegating

By May 2026 a quarter of Codex users had handed the agent a single task estimated at more than a full working day of human work. Most corporate AI training still teaches the chat window. This essay shows the evidence, lets you translate it to your own people, and ends with a training audit you can run this week.

Before you read · 30 seconds3 questions · adapts this page
1. The instructional design method for identifying the tasks a role actually performs, so that training targets them, is called…
2. When you delegate a task of several hours to an agent, the competency that most decides the result is…
3. That a person has actually learned a skill is best evidenced by…

What just happened, and why it matters: three retrieval questions estimated what you already know and this page branched on the result. A rule that small is still a model of the learner: assessment first, then adaptation to competence. That is exactly the discipline delegation demands of you as well. A good delegator holds a model of what the worker, human or agent, can reliably do, and hands over work at the edge of it. You have just done to this page what this essay argues you must do to an agent.

For two years the argument about AI at work was an argument about prompting. Write the request well, give it context, iterate in the chat, and the model answers. Then, quietly, the unit of work changed. In June 2026 OpenAI published usage data on Codex, its coding agent, and the shape of the change is unambiguous: people have stopped asking and started delegating.[1]

The number that resets the plan

By May 2026, among a sampled set of individual Codex users, 80.6 percent had made at least one request estimated to correspond to more than thirty minutes of human work, 70.2 percent at least one exceeding an hour, and 25.6 percent at least one exceeding eight hours.[1] Eight hours is a full working day, handed over in a single instruction. The paper behind the announcement frames the shift plainly: agents change the unit of knowledge work from short interactions to delegated, long horizon tasks.[2] The eight hour share grew the fastest, from a low base, in the six months to May.[1]

Two further figures matter for anyone planning training. Adoption grew fastest among people who are not developers: individual non developer users rose 137 times from August 2025 to early June 2026.[1] And inside OpenAI itself, where adoption ran ahead of the market, Legal, Finance and Recruiting crossed over to the agent as their primary tool around April 2026, and Codex now accounts for 99.8 percent of weekly output tokens generated across the company.[1] This is no longer an engineering story. It is a knowledge work story.

Foundations · What delegation means here

A chatbot interaction is short and self contained: you ask, it answers, you ask again. An agent operates for minutes or hours on its own, calling tools and iterating toward a result you specified but did not supervise step by step. The difference is not the model, it is the unit of work. Prompting produces an answer you read immediately. Delegation produces an output you have to receive, inspect and accept or reject. Everything in this essay follows from that shift.

Translate it to your people

Percentages hide their own scale. The figures below are the reported shares of sampled Codex users who crossed each task horizon at least once. Set the number of people adopting agents like these and read what the shares mean as headcount rather than as decimals.

Fig. 1 · Interactive · The shares as peopleDrag the slider
81 70 26 30 min · 80.6% 1 hour · 70.2% 8 hours · 25.6%
10 500 100 people

Shares are the reported proportion of sampled individual users crossing each horizon at least once by May 2026 [1]. Headcount is that share applied to the group size you set: an illustration of scale, not a forecast for a specific team.
Check yourself · 10 secondsRetrieval beats re-reading
A manager says their staff are “great at AI” because they write clever prompts. In delegation terms, what has not yet been shown?

Two ways to work with a model

The two modes are different at the level of the loop. In the conversation loop you ask, the model answers, and you judge the answer in front of you; the cost of a bad answer is small because you see it at once. In the delegation loop you specify, the agent works unsupervised, and you meet the result only at the end; the cost of a bad result is larger and later, and it lands on whoever has to catch it. Training that rehearses only the first loop leaves the expensive part of the second loop unpractised.

Fig. 2 · The two loopsSame tool, different job
CONVERSATION LOOP You ask a question Model answers at once You judge it in front of you skill: phrasing the ask DELEGATION LOOP You brief and set criteria Agent works unsupervised You verify what comes back skill: briefing and verifying
Both loops use the same model. The conversation loop rewards a well phrased ask judged on the spot; the delegation loop rewards a precise brief and disciplined verification of work done out of sight.
In the conversation loop a bad answer costs a second glance. In the delegation loop it costs whatever the person who signs off fails to catch.

The residual skill is verification

There is a forty year old warning that fits this moment exactly. Lisa Bainbridge, writing about industrial automation in 1983, described the ironies of automation: when a system takes over the doing, the human is left with the harder residual job of monitoring and correcting work they did not perform, a task for which the old training no longer prepares them.[3] Human factors research since has repeatedly found that people over trust automated output and are poor at catching its rare but consequential errors, a pattern sometimes called automation complacency.[4] Delegation to an agent recreates both problems at knowledge work speed. The output arrives polished, which makes it feel finished, which makes verification feel optional. It is not optional; it is the job.

Foundations · Why verification is the hard part

Producing work and checking work are different competencies. Checking a multi step output you did not create means reconstructing what good looks like, locating where it could plausibly be wrong, and testing those points, all without the context that the producer built up along the way. That is why acceptance criteria matter: written before the work starts, they turn a vague act of inspection into a concrete list of things that must be true. Without them, verification collapses into a quick read of something that already looks right.

Training for a task that is vanishing

Instructional design begins with a job and task analysis: you study the task as it is actually performed, then build training for that task, not for a task that used to exist.[5] Measured against that discipline, much corporate AI training is aligned to the wrong task. It teaches the conversation, prompt craft, context windows, talking to the model like a colleague, at the moment the work is moving to delegation. The curriculum is not wrong so much as out of date: it rehearses the loop on the left of Figure 2 while the value migrates to the loop on the right. The gap that results is not a gap in tool access, which is nearly universal. It is a competency model that lags the workflow.

Check yourself · 10 secondsSection 5 of 6
You automate the production step of a task but keep the same training. What does Bainbridge’s argument predict becomes the person’s real job?

Exercise · Audit one training module

The skill this essay wants to leave you with is separation: hearing what a training module rewards and knowing which loop it belongs to. Below are four statements drawn from real AI training. Classify each, then take the checklist into your next learning review.

Classify each elementConversation era or delegation era
“The module teaches staff to write clear prompts and keep refining them in chat until the answer looks right.”
“Staff write an acceptance checklist before the agent starts, then grade the returned output against it point by point.”
“Assessment is a timed quiz on prompt syntax and model settings.”
“Learners must take an agent’s multi step output, find the two errors it introduced, and justify sign off or rejection.”
The checklist: 1) Does the module teach staff to set acceptance criteria before the work starts? 2) Does it assess whether they can verify an output they did not watch being produced? 3) Does the final task look like the job, delegate and check, or like the tool, chat?
What this evidence does not prove
The Codex figures come from one company’s tool and a 0.1 percent sample of individual users, and the task horizons were estimated by a model acting as judge, so treat the thresholds as directional rather than exact.[1][2] Adoption at a frontier lab runs ahead of the average organisation, and these numbers cannot tell you the pace inside your own. The direction of travel, from conversation to delegation, is well supported. The claim here is about which competency that shift makes valuable, not about a precise timeline for any one team.
Key takeaways
  1. The unit of AI work is shifting from the short conversation to the delegated, long horizon task. By May 2026, 25.6 percent of sampled Codex users had delegated a task estimated at over eight hours of human work [1].
  2. The fastest growth is among people who are not developers, so this is a knowledge work change, not an engineering one [1].
  3. Delegation makes verification the residual human skill, and verification is harder than the doing it replaces [3][4].
  4. Most AI training still teaches the conversation. Job and task analysis says to train the task as performed, which is now delegate and verify [5].

The follow-up quiz arrives in three days.

Spaced retrieval is the best evidenced way to keep what you just read. Leave an email and you will get three recall questions on this essay in three days, plus the next essay when it publishes. That is the whole mechanism; unsubscribe any time.

References

Primary and peer reviewed sources only. Links go to publisher pages.

  1. OpenAI (2026). How agents are transforming work. openai.com/index/how-agents-are-transforming-work
  2. OpenAI (2026). The Shift to Agentic AI: Evidence from Codex. arXiv:2606.26959. arxiv.org/abs/2606.26959
  3. Bainbridge, L. (1983). Ironies of Automation. Automatica, 19(6), 775-779. sciencedirect.com/science/article/abs/pii/0005109883900468
  4. Parasuraman, R., and Riley, V. (1997). Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors, 39(2), 230-253. journals.sagepub.com/doi/10.1518/001872097778543886
  5. Jonassen, D. H., Tessmer, M., and Hannum, W. H. (1999). Task Analysis Methods for Instructional Design. Routledge. routledge.com
MC

Manolis Charalampous

PhD candidate and teaching assistant at the University of Cyprus, researching AI driven marketing performance and adaptive learning. He leads group marketing across 160+ companies in 23+ locations, with twelve years across maritime, telecommunications, legal and finance, travel and gaming. He writes about applied AI, learning science and marketing performance at m4no5.com.

COPIED