·

AI Job Design: How to Make Agents Earn Their Place

AI job design is the step most enterprises skip: no role, no scope, no decision rights, no named owner. What an agent needs before it starts work.

Title card reading AI job design, listing the four things an agent needs before it starts work: a statement of work, an authority line, a probation record and a named consequence owner.
In this article9 min read

The licence is provisioned on a Monday. By Thursday the tool is summarising incident calls, and by the following sprint it is drafting the release notes that go to the change board.

Nobody wrote down what it is allowed to do. Nobody wrote down what it must never do. There is no owner, because a licence does not come with one.

That gap is what AI job design describes, and it is the difference between a pilot that impresses a steering group and a system that survives contact with a value stream.

An AI cannot be accountable, because an AI cannot be fired.

AI job design is the step between the licence and the value stream

Hiring a person into a delivery organisation involves a sequence nobody skips. A role is defined. Boundaries are set around what the person may authorise. A manager is named. A probationary period tests whether they have understood the context rather than just the task.

Deploying an AI tool involves none of that. A licence is provisioned, access is granted, and teams are told to be more productive.

That would be defensible if the tool were passive. Moving from one version of a spreadsheet to a newer one requires no job design, because a spreadsheet does not act.

Modern systems act. They summarise meetings, write code, draft compliance reports and increasingly execute multi-step workflows against enterprise systems. An actor placed in a value stream without a defined scope or decision rights is an ambiguous actor, and ambiguous work does not flow.

AI job design comparison showing the six steps taken when hiring a person against the three taken when provisioning a tool, with a dashed line marking everything left undefined.
Ambiguous work does not flow. An ambiguous actor is rejected by the system that receives it.

What exactly would be written on this system’s statement of work?

Most organisations cannot answer that for a tool that has been running for six months.


The pilots are not failing on capability

MIT’s Project NANDA published The GenAI Divide: State of AI in Business 2025 in July 2025, and one number from it has been repeated more than any other: around 95% of organisations were seeing no measurable return on an estimated $30–40 billion of generative AI investment.

The more useful figures sit underneath it. Roughly 80% of organisations had investigated general-purpose tools and 60% had evaluated task-specific enterprise ones, but only about 20% reached a pilot and around 5% reached production.

The report also found that externally sourced tools reached deployment about 66% of the time against roughly 33% for internal builds. That gap says less about engineering quality than about the obligations that come with owning a system rather than renting one. It also found that some 90% of employees were using personal AI tools for work while only about 40% of their employers had bought official subscriptions.

That last pair is the interesting one. The capability was already in the building. What was missing was a sanctioned role for it.

This is not a peer-reviewed study, and the headline figure is routinely overstated. It rests on 52 structured interviews, 153 survey responses gathered at industry conferences, and analysis of 300-odd public initiatives, over a six-month window. “Zero return” means no measurable profit-and-loss impact attributed by the respondent, which is a judgement rather than a measurement. It does not show that 95% of AI work is worthless, and it was never designed to.

What it does describe, consistently, is a failure at the point of integration rather than at the point of capability. The report’s own barriers are a learning gap, a workflow integration gap and a tool-choice mismatch. None of those is a model problem.

20%

of organisations reached a pilot with a task-specific enterprise tool.

5%

reached production with it.

66% vs 33%

deployment rate for externally sourced tools against internal builds.

Agentic changes the stakes, and the vendors know it

Gartner predicted in June 2025 that more than 40% of agentic AI projects would be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.

Anushree Verma, a senior director analyst there, put it plainly: “Most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.”

The same analysis estimated that of the thousands of vendors marketing agentic capability, roughly 130 genuinely had it. Gartner calls the rest agent washing: assistants, robotic process automation and chatbots rebranded without the underlying capability.

A Gartner prediction is analyst judgement rather than measurement, and the firm has commercial reasons to be interesting. Treat the 40% as a considered opinion about direction.

The direction is what matters here. A system that only retrieves information can be governed loosely, because a wrong answer stops with the person reading it. A system that triggers a workflow, provisions access or queries a production database has already acted by the time anyone reviews it.

Regulation arrives later than the risk

The EU AI Act is widely cited as the forcing function. Its implementation timeline suggests otherwise for most delivery organisations: it entered into force in August 2024, prohibitions applied from February 2025 and general-purpose model obligations from August 2025, but the high-risk system obligations phase in from December 2027 and August 2028.

Anyone waiting for a compliance deadline to force the design work has several years to wait, and will spend them accumulating undefined systems.

The text is nonetheless worth reading early, because Article 26 states the requirement more precisely than most internal policies manage: “Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support.”

Natural persons. Competence, training and authority. That is a job description, written by a regulator.


Four things an agent needs before it starts work

Each of these is a decision somebody has to make and write down. None requires new technology.

1. Write the statement of work. A defined purpose and explicit limits, not a capability. An agent authorised to review code against an internal security repository and flag deviations is not the same as an agent authorised to commit to the main branch, and the difference has to exist somewhere other than in the prompt.

2. Draw the authority line. Name where the system’s authority ends and human accountability begins, on the least-privilege principle. An agent may draft a candidate rejection; a recruiter presses send. The test to apply is not typical behaviour but worst case: the maximum damage the system can do in a single unsupervised action.

3. Run the probation. New hires are reviewed heavily for three months and the review produces changes. AI outputs need the same: when a summary is wrong, the team logs the error type and adjusts the context or the prompt, rather than quietly correcting the text. Without that, the system cannot structurally learn, and neither can the organisation. This is the review capacity problem described in the verification gap. The constraint is rarely generation.

4. Name the consequence owner. Every active agent needs a named human who holds the risk for its outputs. If a tool fabricates a figure in a regulatory report, the accountable person is the one who owns the report, not the vendor who supplied the model. The failure mode this prevents is the diffusion of responsibility described in the Boeing accountability argument: a plausible system explanation can make an organisational choice look inevitable, and nobody’s name is attached to it.

For every agent running in the organisation today, name the individual whose judgement is on the line if it is wrong.

If that name does not exist, the agent is not governed—it is merely running.

Diagram placing an agent between an upstream supplier and a downstream consumer, with an authority line above it, a named consequence owner below it and a probation loop capturing error types.
The upstream supplier and downstream consumer are usually the parts nobody assigns.

The failure modes worth watching for

Undefined systems produce a recognisable set of symptoms, and they show up in behaviour rather than dashboards.

– People quote AI-generated summaries in meetings but cannot explain the reasoning or point at the source data. – Background agents sort tickets or manage calendars with no named human checking their accuracy. – Projects succeed in a sandbox and stall at production, because the human workflow was never redesigned around the new participant. – Task completion metrics improve while delivery confidence among engineers falls.

The one that costs most is slower and harder to see. If the routine work is what teaches a junior engineer how the system actually behaves, automating all of it produces polished output and no accumulated judgement. That is an inference rather than a measured finding, but it follows the same logic as deciding who genuinely needs depth: capability forms where people practise, and removing the practice removes the capability.

Download the AI Job Description Pack

The free pack below turns the four elements into something a delivery lead can complete for one agent in a single session, before the licence is renewed.

It contains a statement of work template with explicit prohibitions; an authority line with a maximum-single-action damage assessment; a probation record for logging error types rather than corrections; a named consequence owner with the escalation path; a value stream placement sheet naming the upstream supplier and downstream consumer of the agent’s output; and a fallback rehearsal covering what shipping looks like if the tool is unavailable.

Nothing in it produces a maturity score or a compliance claim. AI job design is a management artefact, not a certification.

What the research can—and cannot—tell us

The MIT NANDA report is a working paper from a research initiative, not a peer-reviewed study, and its headline figure has been misquoted more often than it has been read. Its sample is small, self-selected and weighted towards conference attendees, and its central measure is self-reported.

The Gartner figure is a prediction. It has not happened, and it may not.

Neither source tested job design as an intervention. No study establishes that organisations which write statements of work for their agents reach production more often than those which do not. The correlation the argument leans on—that failures cluster at integration rather than capability—is consistent with the evidence but is not proof of the remedy.

On the regulation, Article 26 is definitive about what deployers of high-risk systems must do. But most tools in a typical delivery organisation are not high-risk systems under the Act, so quoting it at them is a borrowed standard rather than a legal obligation.

Apprenticeship erosion is reasoning from how expertise forms, not a measured finding about AI specifically.

The four elements, the pack and the argument that job design is the missing step are Beta Tester Life interpretations of this evidence. They are not validated, and they are not a substitute for legal, regulatory, HR or professional advice.

Sources used

  1. Aditya Challapally, Chris Pease, Ramesh Raskar and Pradyumna Chari, The GenAI Divide: State of AI in Business 2025, MIT Project NANDA, July 2025.
  2. Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, 25 June 2025.
  3. European Union, AI Act, Article 26: Obligations of Deployers of High-Risk AI Systems.
  4. European Union, AI Act implementation timeline.
  5. DORA, 2025 State of AI-assisted Software Development Report, Google Cloud, September 2025.

DORA’s finding on the same question is the shortest version of the argument: AI does not fix a team, it amplifies what is already there. An organisation that cannot say what its people are allowed to decide will not suddenly become clear when it adds a participant that never asks.

The missing artefact is not a better model. It is a job description with a human name at the bottom of it.

Kevin Campbell, writer behind Beta Tester Life

Behind the notebook

Written by Kevin Campbell

Thirty years of technology, delivery and organisational change—translated into practical thinking for people doing the work.

Continue the journey

One thought leads to another.

Scroll to explore

Conversation

Add to the thinking

Questions, experience and thoughtful disagreement are welcome.

Leave a Reply