Most discussions about AI and organisation structure start with the org chart. Will middle management shrink? Will a manager supervise agents instead of people?
An org chart cannot answer these questions, because it shows who reports to whom. It does not show how work moves from the request that starts it to the result that completes it. When an agent takes over steps in that work, the workflow changes first. Reporting lines change later, if at all.
This week I ran a class exercise that starts from the workflow and reaches the org chart last. This post explains the idea behind it, and paid subscribers get the full method and the worksheet.
Two companies divide the same work in two different ways
Stripe and OpenAI have both published how one of their engineering teams writes code with agents. The work is the same in both cases: turn a task into a merged change. The division of responsibility is different.
My reconstruction from the companies' published accounts. Neither company endorsed it.
At Stripe, an engineer starts a background job, the agent prepares a proposed change, the same engineer inspects it and a second engineer reviews it before it merges. Stripe reports more than 1,300 of these human-reviewed changes merged each week. In one OpenAI product team, engineers set priorities and acceptance criteria, other agents review the code, and human review is optional in that one repository.
Neither company publishes a full organisation chart or a ratio of agents to engineers. What the diagram shows is that three things which used to sit inside one job now separate. Someone owns the outcome and is accountable for it. A separate team maintains the agent service, its tools and its permissions. An agent may assign a task to another agent for as long as that task runs. An org chart has no place to show any of these.
What this could look like outside engineering
I drew the same idea for a recruitment function in a large company. This is a forecast of one plausible design, and it does not describe any company's current structure.
On the left, people in separate teams source candidates, coordinate assessments and arrange interviews, and the work moves between them. On the right, one person owns the hiring outcome and directs agent services that a separate team maintains. The hiring manager still decides. The layer that mainly passed work between teams is the part that changes.
A picture like this is easy to draw for someone else's company. The useful question is what the equivalent picture looks like for your own team, and the answer depends on one decision you have to make for every step of the work.
The decision that controls everything else
For each step in a workflow, you have to decide whether it still needs a person. When I ran the first version of this worksheet, my students marked almost every step as human judgement. Approving leave, checking the balance and informing the team all went to a person.
The problem was my definition. I had given them the example of allowing leave beyond entitlement, and they read it as "any exception needs a human." An exception can follow an agreed rule, though. If extra leave is allowed when cover is arranged, an agent can check that condition and decide.
The revised worksheet marks a step as human judgement only when it passes two tests.
The first test asks whether the choice is still unresolved after you try to write guidance for it. The second asks whether a wrong choice would have a serious cost that you can name. Two urgent leave requests that conflict, where the rules cannot decide whose need comes first and either choice leaves a critical service unstaffed, pass both. Checking a leave balance passes neither.
When a group applies both tests honestly, far fewer steps stay with people than they first expected. That changes how many people each position needs, which changes the reporting lines. The org chart comes out of those decisions, and it is the last thing you draw. I have not yet run the two-test version with a class, so treat it as my current definition.
The full exercise has five parts: drawing today's hierarchy, writing every step, applying the tests, deciding what remains for each position once an agent is added, and drawing the proposed hierarchy. The rest of this post walks through how to run it with your team, what usually changes when you do, and what to check before acting on the result.
[Insert the paywall break here, then delete this line.]
Running the exercise with your team
Use one workflow your group knows well, and a group of people who actually do that work. No AI account is needed. The worksheet uses a leave request as its example throughout.
Draw today's hierarchy. Name what starts the workflow and what completes it. List everyone involved with head counts, and draw who reports to whom for this team only.
Write every step. One action per row, numbered T1, T2 and onwards. Include checks, handovers and rework, because those steps often take the most time.
Mark each step. J for judgement under the two tests, S for securing someone's cooperation, P for physical work. Most steps get no mark. Every J must state the choice, why guidance cannot settle it and what a wrong choice would cost.
Assume the company adds an agent. Agree what it can do in this workflow. Then fill in one card for every current position. Each card asks what the agent would do, what work still needs this person and why, what other work the person could do, and what skills, access or authority they would need. Fill in a card even when a position has no work left in this workflow, and do not invent work to keep someone occupied.
Draw the proposed hierarchy. Draw what the remaining human work needs, show reduced counts such as 3 to 1, and list removed positions separately. Then compare people and reporting levels before and after.
Here is what the whole sequence produces for a simple leave-request workflow.
Five of the six steps go to the agent, because written guidance can settle them. Only the conflict between two urgent requests passes both tests. The leave officers who remain handle that step and every case the agent refers to them, which is why the count falls from five to three and does not fall to one. Your own numbers will depend on how many of those cases arrive in a normal month.
What usually changes
A layer whose main job is allocating work and passing status upward becomes less necessary once an agent does that. A layer that also handles judgement, people development or legal obligations usually stays, with different work.
A manager's span becomes two numbers: human direct reports and agent runs. More agent runs do not mean the manager can support more people. The limit is how many difficult exceptions arrive.
Decision rights need to be written for each action. For each consequential step, name who may act, what evidence they need and who can stop it. An agent repeats every ambiguity each time it runs.
Check these before acting on the result
One workflow does not justify redesigning a department. A person who loses most of their work in this process may have other duties the exercise never examined. Test the proposed structure against your busiest month and a week when two people are absent. Count how many exceptions actually arrive each month, because the human work that remains is mostly exception handling, and that count sets the head count.
Get the worksheet and the AI version of the exercise
Everything for this exercise is in one GitHub repository:
github.com/Shivak11/human-agent-org-design
It contains the ten-page fillable worksheet, which you can print or complete in a PDF app with your team. It also contains a Claude skill that runs the same exercise as a conversation: it interviews a manager about one workflow, one question at a time, and produces the current and proposed structures as diagrams. Three fictional examples are included, from a small support team to a CHRO's whole function, so you can compare your answers after you have made your own.






