A field guide that takes your firm from scattered personal use to AI running real work reliably. You don't read it front to back: you find where you actually are and start there.
Which recurring tasks actually drain your week? You have the basics; the value so far is real but scattered. That is the norm, not a failing, and this altitude is where it stops being scattered.
Most people meet AI as a curiosity: a poem on demand, a quick answer to a trivia question. The working version is less flashy and worth far more. The rule of thumb: if the work is made of language, a model can carry part of it today. Reading, drafting, comparing, extracting, summarizing. What stays yours is the judgment: decisions, advice, the final word. The model prepares those, gathering and comparing so you decide with better material, but it does not make them. Most professional jobs are a braid of the two, which is why the gains land everywhere once you look for them.
A model is strongest at volume: the reading and comparing that costs a person hours, done without tiring on page two hundred. It is weakest where a single exact fact or figure must be right without checking. Both of those facts shape how you use it, and the next article turns them into working habits.
The tasks are one person's week at one firm. The marking is the part to copy: language work gets handed over, judgment stays, and most items turn out to be a braid of both.
Go through what you actually did this week and sort it. The test is what the work is made of, not how important it feels.
Language work can be handed overReading, drafting, comparing, extracting, summarizing. If you could hand the task to a capable temp with written instructions and check the result, a model can carry it.
Judgment stays, preparedDecisions, advice, the final word. The model gathers, drafts and compares so you decide with better material. It prepares the decision; it does not make it.
The same tool lands on every desk differently. Six patterns cover most of what firms find in the first month.
These are patterns, not products. None of them needs special software: the mainstream AI tools can do a version of all six today, in a browser, from a plain-language instruction. Your own week supplies the real ones.
Keep your task list for one full week and mark each item as it happens. Marking real work beats brainstorming about work, for the same reason the later stages record processes instead of remembering them.
Three marks are enoughHand over, hand over the draft, stays with me. Most knowledge work turns out to be the middle category: the model produces, you certify.
Note what repeatsThe items that show up every week are worth the most, because whatever you learn about them pays back weekly. They are also the seeds of the library this stage ends with.
Pick one recurring item from the marked-up week and delegate it properly. The weekly update is a good first choice because the material already exists and you can judge the result at a glance.
Give it a real instructionGather your notes and sent messages from the week, and ask for a summary of what was achieved, in your voice, addressed to the person who actually reads it.
Review, fix, sendThe first draft will be close. Your corrections are not overhead; they are what makes next week's draft better, and they are the beginning of the review habit the next article builds.
This needs no setup, no connected systems, and no permission from anyone. The account you already have is enough, and free plans are enough to feel the gain.
If the work is made of language, a model can carry part of it today, across every desk in the firm. Start with one recurring task this week; the account you already have is enough.
The difference between a party trick and a working tool is how you talk to it. A question asks the model to produce an answer from what it already knows, and you get something generic, because it knows nothing about your situation beyond the sentence you typed. A brief gives it the situation, the material and the standard to hit, and you get something usable. Most disappointing first experiences with AI are questions that should have been briefs.
Any of the mainstream tools works, in a browser, and a free account is enough to learn on. The tool is not the skill. The skill is the loop this article teaches: instruct, review, improve.
State the task, hand over the material, name the audience and the standard to hit. What you would tell a colleague taking over the task is exactly what the model needs.
Attach the material, don't describe itThe thread, the notes, the document. The model works far better from your actual material than from your summary of it.
The structure is usually sound. The specific facts, names and numbers are where you check.
Make "look it up" your reflexWhen a fact matters, have the model search and cite its source rather than answer from memory, and open the source yourself when the stakes are real.
AI executes, you certify. Nothing the model produces goes to a client, a colleague or a decision without your read. That rule is what makes delegating safe, here and at every later stage.
Say what is wrong and ask for the revision. The second pass with real feedback is where the quality jumps.
Feedback beats retyping"The rate is wrong, use the attached" gets a better result than silently fixing it yourself, because the correction improves this draft and teaches the next brief.
Keep the briefs that workedA brief that produced a usable draft is worth saving as written. The last article of this stage turns those into a library.
Two habits keep the loop honest, and both are about what the model knows.
Give it your context deliberatelyWithin a chat it knows only what you put there. Paste the relevant background instead of assuming it remembers your business. The next article makes this permanent.
Start fresh when a thread driftsA long chat full of corrections drags old misunderstandings forward. Take the good output and open a clean session.
Brief it like a colleague, review it like a junior's draft, and make "look it up" your reflex. The loop is the skill; the tool is just where you run it.
A model that knows nothing about your business starts every chat as a stranger. This article shows you how to give it your world once. Context is what it knows before the chat starts: standing instructions, project spaces, memory. Connections are what it can see and do, live, inside the access you granted: your calendar, your files, your inbox. Set up what it should always know, connect what it may see, and keep both under your control.
Control is the other half of the setup. You stay accountable for what the system reads and does. Narrow scopes, per-use approvals, and an occasional look at what memory has stored keep that accountability real rather than theoretical.
Who you are, what the firm does, how you like output. Written once, applied everywhere.
Keep it short and currentA page at most: the firm's services, your role, the tone rules, the formats you want by default. Stale instructions quietly degrade every chat, so re-read them when something feels off.
One space per ongoing engagement, loaded with its documents, so every chat inside it starts informed.
Load the documents, not descriptions of themThe engagement letter, the notes, the drafts. A space that holds the real material turns "let me explain the background" into a sentence you stop writing.
Connect the systems the work needs, one at a time, each with the smallest access that does the job.
Grant access like a key cardTo named rooms, not the whole building. Approve actions per use rather than choosing "always allow", and disconnect what you stop using.
Let the work pull the connectionConnect the calendar when the Friday summary needs it, not because connecting things feels like progress. Every connection you don't have is one you don't have to think about.
A short standing review of your own setup, the personal version of what Stage 4 runs for the whole firm.
One professional habit before client work goes in: know what your tool does with what you paste.
The data settings tell you whether conversations are used for training and how long they are kept. Check once, choose deliberately, and the question is settled. Client material deserves the same care in an AI tool as in any other system, and when colleagues start following you onto the tool, business terms are the natural next step. The later stages pick that up.
Give it your world once: standing instructions, project spaces, and the narrowest connections that do the job. What it knows and what it may touch stay decisions you make, not defaults you inherit.
Questions save minutes; delegated tasks save hours. Dabbling uses the model for fragments of work you are already doing, a sentence here and an idea there, so the task's shape and its hours stay yours. Delegating hands the model a whole task with a brief and makes you the reviewer: it produces the full first version, and your time moves from producing to directing and certifying. The real hours live on the delegating side of that line.
Delegation changes where your time goes, not how much care the work gets. A real brief plus a serious review still puts your judgment on the page. What disappears is the hours of production in between.
Produce the full first draft of the proposal for the new enquiry, ready for my review. Not an outline and not suggestions: the complete draft.
The enquiry notes from Outlook, our last two proposals for similar work from SharePoint, and the outline I sketched. Work from these, not from general knowledge.
The client's operations director, who reads quickly and dislikes filler. Our standard tone: direct, concrete, no superlatives.
Three pages at most. The scope section mirrors their language from the enquiry. Anything the notes don't cover is flagged as an open question, not invented.
I check every fact, name and number against the material, judge the structure, and send back one round of pointed corrections before anything leaves my desk.
The systems and the task are one firm's. The five sections are the part to copy; they fit any task with a known good shape.
Proposals, summaries, reports, first-pass reviews: you can describe what good looks like, so you can check what comes back.
First versions, not final callsDrafts, options and analyses to react to. Decisions stay with you.
Stay inside its strengthsLanguage-shaped work first. Hold back the tasks where you cannot verify the output until your verification habits are solid.
Material, audience, tone, and what good looks like, stated up front. The brief is where your judgment enters the work.
Say what must not happen"Anything the notes don't cover is flagged, not invented" belongs in the brief. The model follows the standard you state, not the one you assume.
Facts, names and numbers checked; structure judged; nothing forwarded unread. Because it does carry your name.
One revision with pointed feedback is part of the job, not a failure of the tool.
Brief, review, iterate: the same loop as the last article, on a bigger unit of work. The habits carry over unchanged; what grows is the size of what you hand over.
Tasks you delegate well once are worth making repeatable. That is the last article of this stage.
Delegate whole tasks with a real brief, and review like the work still carries your name, because it does. The hours come back on the delegating side of the line.
The first delegation is an experiment; the fifth should be a routine. A good result that disappears into the chat history gets rebuilt from memory next month, a little differently, with a little less of what made it work. A captured asset is that same working brief, saved as a reusable instruction that runs the same way every time and can be handed to a colleague. Capture is the difference between being personally faster and building something the firm keeps.
This is also the point where paying for the tool earns its keep. A free plan carries you through the working loop. The machinery of capture and context, project spaces, connections, saved skills and scheduled tasks, mostly lives on the paid plans. You are not paying for a smarter chatbot; you are paying for the features that turn a chatbot into a delegation system.
The entries are one person's. The shape is the target: your most-repeated tasks captured and named, the most-delegated one written as a skill, one piece of recurring work arriving on its own.
Your best instructions, kept where you can rerun them, named so you can find them.
Capture at the moment it worksThe brief that produced the best draft so far is in this week's chats right now. Saving it takes a moment; rebuilding it from memory next month loses a little of what made it work.
A skill is a multi-step instruction the model applies on demand: how your firm writes summaries, structures reviews, formats deliverables.
Reusable, versionable, shareableA skill can be handed to a colleague and improved in one place. It is the first asset you build that belongs to the firm rather than to a chat.
Recurring work set to run on a rhythm: the Monday pipeline summary now arrives instead of being produced.
Review what arrivesThe certify rule still holds; what changes is when your attention enters. Improve the instruction, not each output, and the quality compounds.
One agreed place, improved in the master copy, reviewed now and then for what has gone stale.
Scattered private variants are how capture quietly dies: three colleagues each running a slightly different version of the same brief, none of them the best one. The master copy is the asset; everything else is a copy of it.
The starter library, concretely: five briefs that work, one skill, one scheduled task. That completes what this stage set out to build.
When it works twice, capture it: templates, skills, scheduled tasks. Personal speed fades with the person; a library compounds.
Most readers should run at this altitude for a while, and staying here is fine. Operating well at Stage 1 looks like: the setup is trusted, the delegation habit is real, and the starter library grows on its own. Stage 2 is there for the day a shared workflow needs more than personal leverage, not because a next stage exists.
This is where AI stops being a personal trick and becomes something the firm does: one shared workflow, with a quality bar that survives the enthusiast leaving.
Stage 2 starts with a decision: which workflow gets your firm's first serious AI effort. Candidates are not the problem; any working firm has plenty of tasks that could be automated. The hard part is picking one you can actually win. The workflow that would look best in a demo is usually heavy on judgment, full of exceptions, and wired into half your systems, and that is exactly why it fails as a first build. The right first pick is less glamorous: it runs often, it has a clear start and finish, and a mistake is easy to catch. This article takes you from a long list to one commitment.
One thing to know before you start looking: the workflows worth winning are not always the ones people complain about. A step everyone grumbles about may cost the firm very little, while a quiet step nobody mentions holds up every job that passes through it. What matters is where work sits and waits, and where your most qualified people lose their week.
Proposal drafting. It passes all seven tests, the hours at stake are real and weekly, and we can check quality by reading the drafts. Next step: map the current process and count one week as the baseline.
A sample, not a prescription: the candidates and systems are this firm's. A software business might recommend support replies; an accounting practice the seasonal crunch. The shape of the document is the part to copy.
The inventory is a structured look at the firm's day-to-day work, and it deserves real attention: this list decides where the first months of effort go. Three questions organize it.
Where does work sit and wait?Go through the queues and handoffs: quotes waiting to be written, tickets waiting to be answered, invoices waiting to be checked. Shared inboxes and handoff points hold more of this than any job description shows.
Where do people retype and reconcile?Every place someone copies data between systems, cross-checks two lists, or re-enters something that already exists somewhere is a candidate.
Where does the same work repeat?The reports, replies, summaries and first drafts that get produced to the same pattern every week.
If the firm has been through Stage 1, part of the list writes itself. The tasks people already delegate to a model on their own initiative are strong signals for what the firm should automate properly.
Ask what people do, not what they think"What do you already hand to AI?" gets concrete answers. "What should we automate?" gets opinions. In many firms the proposal queue shows up here on its own, because associates are already drafting from intake notes with a model and tidying the result by hand.
Every candidate faces the same seven questions. One clear no is usually reason enough to set a candidate aside for now.
Proposal drafting passes all seven at most firms that live on proposals: weekly volume, a clear start and finish, real hours at stake, an easy quality check, and inputs that already sit in the firm's systems. Whatever passes at your firm moves on to the ranking.
Score each surviving candidate on two dimensions, and take your first workflow from the high-impact, low-effort corner. Impact means the hours it drains, how often it runs, what delay costs, what an error costs. Effort means the systems it touches, the exceptions it produces, the judgment it needs.
The candidates placed are the sample shortlist's. Place your own survivors the same way, and the choice usually makes itself.
The commitment is a single bounded task, not a department. If proposals are the bottleneck, the first win is drafting the first version of one, not reinventing how the firm sells.
Put the commitment in writingOne line: the task, where it starts and ends, and the person who owns the win. That line becomes the top of the map you build next.
Where firms typically start: proposal drafting at an agency, first-line support replies at a software business, the seasonal crunch at an accounting practice, after-hours enquiries at a local service firm. Treat these as places to look, and test what you find against the seven tests.
Candidates are everywhere; choosing is the skill. List broadly, filter with the seven tests, rank by impact against effort, and commit to one bounded task.
Every workflow exists twice. The official version lives in handbooks and job descriptions, written from memory, often by someone who no longer does the work. The real version runs this week, full of workarounds, exceptions and quiet fixes people have built over time. Builds fail because they are made from the official version and run against the real one. This article's job: a map of the real process, made in about a week, with the number it actually costs attached.
One part of the real process is nearly invisible: people repair bad data without noticing they do it. They spot the stale price, remember the address that one system gets wrong, fill the field the form left empty. That silent repair disappears the moment a machine runs the step. And a model handed wrong data does not stall the way a script would; it produces something plausible from it. So the map has to capture the data behind the work, not just the steps.
The systems are this firm's, not a requirement — yours might read Gmail for Outlook, Salesforce for HubSpot, Google Drive for SharePoint. What matters is that every step names its real ones.
The person who does the work records one real pass, screen and voice, the way they would show a new joiner. Recording produces the real process; writing from memory produces the official one.
Use the recorder already on the machineA solo video call with recording on, a screen-recording tool, or the system's built-in recorder, whichever is already at hand. The transcript matters more than the video.
Record a real, untidied item, not a staged oneNarrate the off-screen parts as you go ("here I'd normally ring the client back") and don't clean anything up: the workarounds are the most valuable footage.
If several people touch the workflow, each records their partFor the proposal workflow: one recording from the associate, one from the partner. The handoff between them is the most important footage.
Hand the transcript to a model with clear instructions and it returns a structured first draft of the map. The setup decides whether this works once or becomes a habit.
Mapping one workflow? A chat is enoughUpload the transcript with an instruction describing the output: steps in order, who does each, the systems read and written, exceptions marked as such, the decision points. One page, plain language.
Mapping more than one? Make it a standing setupA dedicated project for process mapping, with that instruction saved once. Every transcript comes back in the same shape, and the space keeps past maps, so a new map's handoffs connect to workflows the model has already seen.
That growing collection is the beginning of the firm's process library, the one Stage 4 makes official. The fifth map costs a fraction of the first.
The person who does the work corrects the draft against the recording. A person certifies the map, not the model.
Fix what the recording missedWrong order, skipped steps, and the exceptions that didn't occur in the recorded pass. Mark every decision point; the next article chooses the route on exactly those forks.
Hold it to the barOne page, plain language, and a new colleague could run the workflow from it. If they couldn't, it isn't done.
An automation is only as good as the data it reads, and the map is where you find out what that data really is.
The baseline is one real week of the workflow, counted while it happens, not remembered afterwards and not asked about, because people's sense of time saved is not a reliable measure.
Put a tally where the work happensA shared sheet is enough. Count every pass that week and time each one door to door, arrival to finished, including the waiting. Delay usually hides in the handoffs, not the doing.
This number is what the pilot will be judged against; without it nobody can say whether the new way is better, including you. No baseline, no build.
A single person owns the map, the baseline, and the decisions the next moves ask for. Named now, not after something goes wrong. When the pilot needs its go or no-go, this is whose name is on it.
The map is an asset, not admin. Record the real process, check the data it runs on, count a real week, name one owner; everything you build next stands on it.
The map from the last article lists the workflow step by step, and every step is one of two kinds. Mechanical steps run the same way every time: the same situation always gets the same action. Judgment steps depend on what is in front of you: something has to be read, weighed or worded before the step can happen. Mark each step as one or the other and most of the routing work is done, because the rule of this article is simple: give every step the simplest tool that can handle it.
The tools come in three sizes. A fixed rule does the same thing every time: if this happens, do that. An AI step takes on a single judgment task with your instructions and guidelines. It drafts the reply in your tone or pulls the key figures from the contract, then it stops. An agent gets a goal, boundaries and sign-off points, and carries out several steps in a row, choosing the order to fit the case in front of it. Every size up buys flexibility and gives up predictability. That is why the simplest tool that fits is the right choice, not the most capable one.
Go through the map with its owner and mark each step as one or the other. The decision points the map already carries show you where the judgment sits.
The test is a sentenceIf you can honestly finish the sentence "when this happens, we always do that", the step is mechanical. If the sentence needs "it depends", it is a judgment step.
Disagreement is informationWhen two people who know the workflow disagree about a step, it is usually two steps written as one. Split it on the map before you route it.
Rules are predictable and fast, and they should carry as much of the workflow as possible.
Start with what your systems already offerMost systems on the map, from the inbox to the client record to the file store, have automation built in for exactly these moves: when an enquiry arrives, create the record; when a proposal is sent, file it and log it. Where two systems don't reach each other, the automation platforms bridge them.
Keep each rule small enough to say out loudOne rule per line of the map. If nobody can say from memory what a rule does, nobody will trust it when the workflow misbehaves.
An AI step is a single task, handed over with instructions and guidelines. It does that task and stops. What happens before and after stays with you.
For occasional use, a saved instruction is enoughThe associate opens the standing instruction: draft the first version from these intake notes, in the firm's tone, at this length, and flag anything the notes don't cover. Attach the notes, get a draft back.
For every pass, wire it into the flowThe same instruction, saved once inside the automation that carries the workflow, so every new enquiry produces its draft without anyone opening a chat. That is the AI-step node in the build above. The instruction gets written once and improved deliberately, instead of being retyped a little differently each time.
Either way, the instructions are the real asset. The tone, the structure and the guidelines survive even if the tool underneath changes.
The steps that stay with people become configured stops in the flow: the workflow assigns the task, waits for the answer, and cannot continue without it.
Decision points become assigned tasksThe partner's scoping call arrives as a task with the enquiry attached, and the pass waits until it is answered. Built this way, the decision point cannot be skipped on a busy day.
Approvals become required sign-offsWhich results need a person's approval before they count is the firm's own written choice. Many firms draw the line at whatever a client sees. Wherever your firm draws it, the stop belongs in the build, not in anyone's memory.
If the steps in your map hold for nearly every case, automate them as written. If the rules keep producing exceptions until nobody can hold them all, that stretch is agent territory.
What handing over a stretch meansThe agent gets a goal, the boundaries it works within, and the points where it pauses for your sign-off. It chooses the order of the steps to fit the case. You still decide which systems it may use and which rules it follows.
The big automation platforms now offer agents as a built-in option, so getting one is not the hard part. Deciding how much of the route to hand over is. On the proposal build, nothing needs one; the work that varies case by case arrives with Stage 3's support inbox.
Someone has to build the route, and the choice follows the map. It is a question of fit, not of ambition.
Ready-made: the workflow as a finished productA vendor has already built it, the way bookkeeping tools read receipts on their own. A product has one shape, so the adapting is done by your process, not by the software. That trade is worth making when your workflow looks like everyone else's.
Custom-made: built around the map you drewThe map stays the map instead of bending to a product's shape. This used to mean a software project. With standard connectors, models that can operate tools, and mature automation platforms, a custom workflow is now assembled from proven parts and hardened in weeks rather than built from scratch in months.
Where it runs is a separate decision: a vendor's cloud, or systems you run yourself. Data sensitivity, what you already run on, cost at your volume, and who can maintain it all weigh in. And before anything gets built, three lines go on the map in writing: what it may see, what it may do alone, and who owns it — an alert when it fails, and one named person who responds.
Rules for the steps that never change, AI for the judgment steps, an agent for the stretches that vary. The simplest combination that covers your map is the right one.
The route is built. Now it has to prove itself. A hard switch, where the old way is dropped on day one, puts everything on a setup nobody has watched under real load and leaves the way back improvised. A parallel run tries the new way while the old way keeps producing the real output, so nothing breaks while you learn. Pilot in parallel, and decide before the first run what result means go.
The referee is the baseline from the map, the counted week. The pilot's numbers are compared against that count, not against how the old way felt. And the pass mark gets written before the first run for a simple reason: a pass mark chosen after the results are in will bend toward whatever result you got.
A pass has to take less time than the baseline week's count, door to door, and the share of drafts the partner accepts without rework has to beat the share we counted in the baseline week. Both reference numbers come from the map. Nothing here is estimated.
The build drafts every proposal from the intake notes in Outlook. The associate still writes theirs as usual. Both are timed, and the partner's review decides which drafts would have gone out.
The associates who draft daily, and the partner who signs off. The person who built the setup supports the pilot and does not judge it.
Every pass, timed door to door, both ways. Every case the new way got wrong, and why. The difficult cases from the map (the urgent request that arrives by text, the client record with the stale rate) get pushed through deliberately.
The systems named are this firm's; swap in your own. What transfers is that everything except the counts is written before the first run.
The old route keeps producing the real output for the whole pilot. The comparison is the point of the exercise.
Same passes, both waysThe build drafts every proposal, and for this month the associate still writes theirs as usual. Both are timed, and the partner's review says which drafts would have gone out.
Clients see the old way's output until the pilot passesThe new way's drafts are compared, not shipped. That is what makes an honest pilot affordable: a bad week costs learning, not clients.
Two or three numbers on one page, each set against the baseline week's count.
Use measures the baseline already hasThe time a pass takes door to door, and the share of drafts accepted without rework. If the pass mark needs a number the baseline week didn't capture, go back and extend the baseline first. A comparison without a referee turns into an argument.
This is the discipline the whole article hangs on: go and no-go are defined before anyone has seen a result.
A pilot gets a fixed end date, and the people who do the work daily judge it, not the person who built the setup.
A month is typical for a weekly workflowLong enough for the workflow to run often enough to judge. Short enough that a decision has to come. A fixed date beats an open-ended trial that nobody ever closes.
The builder supports, the team judgesA demo run by the builder always looks good. The real test is an ordinary Tuesday under real load, in the hands of the people who will live with the workflow afterwards.
If nothing broke, the pilot was probably too gentle. Put the difficult cases through on purpose.
The map's exceptions are the test setThe workarounds and exceptions the map recorded, like the urgent request that arrives by text or the client record with the stale rate, are exactly the cases to push through the new route while the old way is still there to catch them.
Every miss goes in the log with a line about why. That log becomes the work list for the rollout, not a mark against the pilot.
On the end date the owner reads the counts against the pass mark and makes the call: go, no-go, or one extension in writing with its own end date.
A pilot that ends in a clear no has done its job. It cost you weeks. A workflow kept alive on hope costs quarters.
Run the new way beside the old, judged by the people who do the work against the counted baseline, with the pass mark set before you start. Stopping is a legitimate result.
The pilot passed. Now the new way has to become the way the work is done, and this is where many rollouts quietly fail: the new tool gets added on top of the old routine, every old step survives, new checks grow around it, and the workflow ends up with more steps than before. The alternative is to rebuild: remove the steps that only existed because a person did the work by hand. A workflow counts as rolled out when the new way is the default, not an option.
Expect the first weeks after rollout to run slower than the pilot suggested. The whole team is now learning what the pilot group already knew. Budget for that dip instead of reading it as failure. It is the cost of the last mile, not a sign the workflow is wrong.
From a named date, the new way is how the workflow runs. The old way remains only as the documented fallback.
Close the choiceAs long as both routes stay open, the busy week picks the familiar one. The announced date closes the choice: from then on, a new case goes down the new route without anyone deciding that it should.
Walk the map once more and ask of each step: does this still need to exist now that the drafting runs on its own?
Coordination built around waiting goes firstAssignment meetings, chasing lists, status check-ins. These existed to coordinate manual work, and each one removed gives time back to the team.
Name the steps that changed in kindThe associate's hour moves from assembling to reviewing. Say that clearly during the rollout. People should hear that their step changed shape, not that it disappeared.
One person on the team who uses the workflow daily, answers colleagues' questions, and collects what annoys them.
A lighter job than it soundsA few questions a week, not a new role. It is how new routines survive a busy month: small questions get answered right away instead of turning into reasons to go back to the old way.
Champion and owner are different jobsThe champion lives inside the workflow and speaks for the team. The owner holds the map, the numbers and the decisions. In a small firm one person may do both, but the two jobs still exist.
If people are still judged by numbers that reward the old way of working, the old way comes back quietly.
Judge the workflow by its measured outcome, the number the next article tracks, and check that nobody is scored on something the rollout deliberately removed. A team praised for the volume of hand-written drafts will keep writing drafts by hand.
Someone checks the workflow's output on a schedule, and the owner can switch back to the documented old way at any time, without calling a meeting.
Pass all three and the rollout is real. Fail one and the workflow is still an experiment with a rollout's name.
Rolled out means the new way is the default, someone owns its health, and the way back exists. Rebuild the workflow around what runs on its own; don't bolt the new tool onto the old routine.
The workflow runs. The last move decides what it earned. The tempting measure is adoption, how much the new way gets used, because it is easy to count and comfortable to report. But what counts is the verified outcome: a result that was accepted and used, like the draft that went to the client at the quality the firm actually ships. Count verified outcomes against the baseline, and let that number make the keep, revise or retire call.
Ask the team how it feels and you will get an answer, just not a measurement. People's sense of saved time is unreliable in both directions. The decision comes from the counted numbers. The team's experience still matters, as context for the revise decision, not as the measure.
The workflow is this firm's; the form transfers as it is. Whatever your firm automated, these six entries are the whole file.
Verified outcomes over a set period, taken from the workflow's own records, not from memory.
Define accepted before you countA draft counts when it passed the person who signs off, at the quality the firm actually ships. Drafts that go out with light edits are wins. Drafts rewritten from scratch are counted separately, and their reasons feed the revise pile.
The cost is the tools' price for the period plus the human time still inside the workflow.
The time hides in three placesPreparing what the AI step reads, reviewing what it produces, and correcting what it got wrong. The review hour is part of what a result costs. Leaving it out flatters every workflow equally.
Cost divided by successes, set against the same number from the baseline week. The workflow wins if a good result now costs less, in money or in scarce hours.
Both sides of the comparison come from counting: the baseline week priced the old way of working, and the measured period prices the new. If the new number is not clearly better, that is what the revise decision is for. The log from the pilot usually says where the cost is hiding.
Keep, revise or retire, decided on a date that went into the calendar when the rollout ended, not whenever someone remembers to ask.
One page: the before and after, the decision with its date, and the lines about what you would do differently, written while the experience is fresh.
With the page filed, the firm has run the whole loop once: picked, mapped, routed, piloted, rolled out, measured. Everything the loop produced, from the map to this page, is reusable the next time it runs.
Operating well at this level is a destination in itself. Stage 3 is there when the next constraint genuinely needs machines acting unattended, not because a next stage exists.
Count verified outcomes, not usage, and price them against the baseline. The number, not the mood, makes the keep, revise, or retire call.
Many firms should operate here, and some should stay. Operating well at Stage 2 looks like: the workflow runs as the default, its dossier is current, its number gets checked monthly, and no heroics are involved. Stage 3 is there when the next constraint genuinely needs machines acting unattended, not because a next stage exists.
You have won once. Now the loop runs again and again, and the route ladder extends up into automation and agents, with the rule that saves the most pain: rules before agents.
Stage 3 lets machines act without someone watching each run, and what they act on is your systems and your data. At Stage 1 a connection acted as you, inside your own accounts, and the blast radius was your own inbox. What runs unattended needs more: its own login and its own narrow permissions, so what it can touch is decided, recorded and revocable. The difference is not the tool; it is the accountability around it. This article gets both ready: grant the least access that does the job, and fix unfit data before you build on it.
The Stage 2 map already did half the work here. It marked every point where a person quietly repairs bad records on the way through. At this stage that marking becomes law, because a machine repeats the error faithfully, at speed, in every run.
Each automation gets its own login rather than borrowing whoever built it, so the access survives that person's departure and can be switched off without touching anyone's account.
The proposal build shows whyIt reads the CRM and the template folder under its own login, and nothing else. When its builder changes jobs a year later, nothing about the automation changes, and nobody has to remember what was connected through her account.
The invoice workflow reads invoices, not the whole accounting system. Grant access like it will be audited, because one day it may be.
Never let one system hold all threeAccess to private data, exposure to content outsiders can write, and a way to send data out. Drop any one of those three legs and the worst case collapses. It is the single most useful scoping rule at this stage, and it costs nothing to apply.
Clean, improvable, unfit, as in the figure. Nothing runs on the unfit pile.
Start from the map's quiet fixesEvery point where the Stage 2 map recorded a person correcting data on the way through is a finding: the stale rate, the wrong address, the empty field. Each one names a dataset and tells you which pile it belongs in.
Stage 1 gave your personal AI its world; this stage gives the firm one. The firm-scale version is deliberately low-tech: a small set of plain documents the firm owns, in a folder the firm controls.
What goes inWhat the firm does and for whom, how it writes, its service conventions, the decisions people keep re-explaining. One routing file on top tells any tool where to look first. A "how we write" file, a services file, one page of conventions per major client.
Why plain documentsBecause the files are ordinary documents rather than any vendor's memory, the asset survives every change of tool and model underneath it. Every workflow in this stage starts informed instead of briefed from scratch.
Context ages the way data does, so it gets the same discipline as everything that runs unattended: a scheduled job folds each week's new decisions in and retires what has expired, a periodic check catches duplicates and contradictions, and one named person approves what changes. Without that maintenance, a shared context layer becomes a contradictory, expensive mess within a quarter.
Four questions, asked before anything gets connected. An afternoon spent here saves the weeks that follow a build on sand.
Least access that does the job, an identity of its own, no builds on unfit data, and a context layer the firm owns. Machines act at machine speed; readiness is what makes that a feature.
The fastest wins in this stage are not the cleverest ones. The steps that never change go first, for three reasons: the rule already exists, on a checklist or in someone's head, so automating it is transcription rather than invention; they fail loudly and fix cheaply, because a file in the wrong folder is visible and reversible in a way a wrong judgment is not; and every step a plain rule handles costs almost nothing per run, which starts to matter once the platforms begin metering your volume. Automate the steps that never change before the ones that think.
A large delivery company automated one boring internal task: account-unlock requests. Around eight hundred a month, each dropping from thirty-five minutes of a person's time to twenty, roughly two hundred hours a month recovered, and the build took five hours on a standard automation platform. No cleverness anywhere: plain rules, with manager approval kept exactly where it was.
The proposal workflow from Stage 2 is surrounded by predictable work that never made it into the first build: the filing, the logging, the chasing, the weekly digest.
Transcribe the checklist, don't inventThese rules already exist, on a checklist or in someone's head. Automating them is writing down what the firm already does, which is why they go first and why they rarely surprise anyone.
The platforms meter volume in different units, and the unit matters more than the sticker price.
The same workflow can be cheap on one meter and expensive on another. Model a month of your real volume on each before committing.
The bands in the figure set the rules: read-only freely, internal records with the harness, client-facing through a parallel run first.
The band is about failure, not importanceA small client-facing automation outranks a large internal one, because the question is what happens if it breaks, not how much it does.
Two practitioners' rules of thumb, useful precisely because they are unglamorous.
Let the machinery change, not the screenIf the team lives in a spreadsheet, let the new workflow read and write that same spreadsheet instead of forcing a new screen. The machinery changes underneath while the habit survives.
Know where chat-with-review stopsWorking through a list inside a chat tool, with a person reviewing, is comfortable up to somewhere around one or two hundred items a run, at real usage cost. Past that, or the moment nobody reviews each item, the job belongs on an automation platform built for volume. Rules of thumb, not physics.
The safe first taste of generative AI inside automation is production work: one approved brief becomes many drafts, variants or formats, produced by a model inside the workflow.
What makes it safe is the brand gate: a person or a hard checklist that everything passes before it leaves the building. The shape is predictable even though the content varies, which is why it belongs with the predictable work.
You have outgrown a platform when you spend more time working around it than in it: workflows split in two to fit its limits, exception handling that fights the tool, costs that climb faster than volume.
That is a normal milestone, not a mistake, and the map you keep from Stage 2 moves with you.
Automate the steps that never change first: they are specified, cheap to run, and cheap to fix. Know your billing unit before the bill teaches it to you, and size every build by what happens if it breaks.
An agent is the first thing you will run that chooses its own path through the work. One word, two meanings, so let's separate them: the popular beginner courses use "agent" for a chat persona with standing instructions, and if you have built one of those, that was real work; keep it. This stage means something with more consequence: software that acts in your systems, toward a goal you set, choosing the order of its steps within your rules, without you watching each pass. The first one gets a narrow job, real boundaries, and a suggest mode to prove itself in.
The support inbox has been waiting for this since Stage 2: it sat on the shortlist marked "kept for later", because the right reply varies too much for rules. That varying shape is exactly what an agent is for, and inbox triage is the classic first agent job: constant volume, quick to check, cheap to correct.
Read each incoming ticket, label it by category, and propose the reply. Route anything it cannot place to a person. One kind of work in one place: triaging this inbox, not "handling support".
Reads: the support inbox, the client list, and the answer library in SharePoint. Nothing else. Scoped in the platform itself, not just described in the prompt.
The firm's own written list. This firm routes complaints, anything mentioning a deadline or legal step, and every sender not on the client list. However confident the agent is.
Ten real tickets, labeled and answered the way the team would, attached as examples. The agent is graded against these, not against a feeling.
Suggest: it labels each ticket and proposes the reply; the team sends it or fixes it. Acting on its own is earned later, category by category.
A running tally of right, wrong, and escalated, kept per category. This becomes the promotion file the next two articles run on.
The systems and the routing list are this firm's; swap in your own. What transfers is that every section is written down before the agent exists.
The platforms you already use offer agents as a built-in option, so capability is not the question. The first design question is whether this work should be delegated to an agent at all.
Two kinds of work fail the gateWork you cannot check, because you would never know whether it was done right. And work a fixed rule already handles, because an agent adds variability where none was needed.
One kind of work in one place. Agents with sprawling toolkits choose badly; small and focused wins.
If you built a chat persona in Stage 1, it makes a useful rehearsal space, at zero stakes.
Give it a mandatory stop phrase and make it end every session with a short report of what it did and where it struggled. You have just rehearsed ceilings, off-switches and logging, the whole grammar of the next article, before anything real depends on them.
The goal, the boundaries, the sign-off points, and examples of right answers, as in the figure. The definition is the work.
Boundaries live in the platform, not the promptThe systems it may use and the data it may see are scoped in the platform itself. A prompt describes intentions; a scope enforces them.
Attach real examplesTen tickets labeled and answered the way the team would. Right answers you can point at beat qualities you can only describe.
In suggest mode the agent drafts the decision and a person confirms it: it does the reading and the routing work, while every outcome still passes a human hand.
Act mode is earned, not assumedA category moves up when its suggestions have stopped needing correction. The path from suggest to act, and the doors that never open, are the next two articles.
Tally from day oneRight, wrong, escalated, per category. The promotion decisions ahead need this record, and it cannot be reconstructed later.
First agent: one narrow, frequent, checkable, low-stakes job, launched in suggest mode with its job description written first. Capability is easy to get; the definition is the work.
Everything in this stage that runs unattended needs a structure that keeps it dependable. Errors in autonomous runs compound: a bad night, unwatched, is an open-ended cost unless something bounds it. The structure is the harness, and it has four parts: ceilings, an off-switch, a logbook, and one page of instructions for the colleague who is not you. Give every unattended run a ceiling, a logbook, and an off-switch.
Ceilings, logs and off-switches are what turn something that works in a demo into something the firm can lean on. None of it needs a security team. It needs an afternoon of configuration and the habits this article describes.
A spend it cannot exceed, a time limit per run, a cap on steps and retries. Ceilings turn a bad night into a bounded cost.
One known way to stop it now, that works, that more than one person can reach. And one page of instructions, written for the colleague who is not you: what this is, what normal looks like, what to do when it is not, who owns it.
Have the agent produce structured output: a table or a labeled list an owner can audit line by line, rather than prose that can hide a wrong call inside a nice paragraph.
Obstacles reported, not improvised around"Could not find the May invoice" is a good result. A guessed invoice number is not. Require the report, and read it.
Every tool the agent holds is something it can use at 2 a.m. with nobody watching. Give it only what the job description needs.
Shared skills and plugins are software from the internetRead what one does before installing it. Skills can run tools and execute code; treat them with the care you would give any software the firm adopts.
An agent that operates a computer directly, screen, mouse and keyboard, gets its own fence: isolate it like a contractor on day one.
Its own user account, a machine that holds nothing else, access granted surface by surface, and a hard payment ceiling if a card goes anywhere near it.
Someone runs this machinery, and deciding who is part of hardening. On a managed platform the vendor keeps it alive while you own the definitions and the checks; running it yourself buys full control of your data and of cost at volume, and it is a real operating job: updates, backups, monitoring.
Data sensitivity, what you already run, volume, and who would honestly maintain it make the call. The full where-should-it-run question gets its own treatment in Stage 4.
Ceilings, off-switch, logbook, one page of instructions, and a scheduled look at what it is actually doing. The harness is what makes unattended work dependable.
How much should the system be allowed to do on its own? Most firms answer by feel: a good week builds confidence, a bad morning revokes it, which means the answer is really a mood. This article replaces the mood with a ladder: autonomy promoted category by category, on performance you can show, plus a set of doors you decide, in advance, to keep human.
Every new capability starts at suggest, where the agent proposes and a person decides, however good the demo looked. A category that keeps earning it climbs to act with review and then to act and report; if corrections rise after a promotion, it moves back down. And some doors never enter the ladder at all: not because the numbers are bad, but because you decided in advance they stay human. Which doors is a business judgment, and it is yours. Many firms choose what clients see, movements of money, and the irreversible. Your list may be shorter or longer; what matters is that it is written before the first promotion, so it never has to be argued case by case.
A category climbs on a low, stable correction rate — measured against a bar written before it started.
The doors you chose to keep human never enter it — exempt by your decision, not by performance.
Password-reset tickets ran six weeks in suggest mode with corrections near zero, earned a month of act-with-review, and now act and report. Refund requests live in the same inbox and score just as well, and they stay in suggest mode anyway, because this firm put movements of money behind a human door. That is not a judgment on the technology; it is a choice about the business.
The ladder works per category, not per agent, so the first job is deciding what a category is.
Group by "same kind of request, same kind of answer"In a support inbox: password resets, renewal questions, refund requests. Each narrow enough that one correction number means something, and each holds its own rung.
Too broad shows up as noiseIf a category's corrections swing week to week for no visible reason, it is probably two categories wearing one label; split it.
Both are set in advance, in writing, so promotion day is a check against a page, not a debate about a feeling.
The ladder runs on the correction rate, so the review step has to leave a countable trace.
Run suggest mode through a visible review stepThe person approves, edits, or rejects each suggestion; that action is the log. Most platforms record it, and where one doesn't, a per-category tally does the job.
Keep the count per categoryA blended number can hide one category quietly failing inside three that are fine. This is the scorecard from your first agent, continued: the promotion file.
Promotions happen on a schedule, not on impulse: a short, recurring look at the numbers, category by category.
Then promote one rung, one category at a time, and log the decision with a date. The list Stage 4 keeps of everything that runs is its natural home.
Permanent doors are configured, not remembered: built into the workflow as required human steps that no promotion can remove.
Each door becomes a required approvalWhatever your firm put on the list routes to a person at every rung. Sent under your name means certified by a person.
If the platform can't enforce the stopThen the category isn't eligible for promotion at all; a door only counts if it cannot be walked around.
When corrections rise, the category moves back down a rung for a set period: no meeting, no post-mortem, no blame. Demotion is the ladder working, not failing; a system that can only ever gain autonomy isn't earning trust, it's accumulating risk.
Autonomy is earned per category, against a correction-rate bar set in advance, and some doors stay human by design. Demotion is maintenance, not drama.
Expansion is not one big program; it is the same loop, run again on the next constraint. The map from Stage 2 is the brief for every build in this stage: record the work as it really happens, count a real week, name one owner. Two sentences carry the whole discipline: map the real process, and no baseline, no build.
Growth changes the failure mode. One workflow that breaks is an incident; ten automations without monitoring are ten quiet degradations racing to be discovered by a client. The monitoring habit from the harness article is what makes ten manageable.
Five entries at one firm, a few months into the stage. The board is one page. When keeping it current becomes its own job, Stage 4 has arrived.
The next workflow is chosen the way the first one was: where work waits now. The shortlist from the first win is still the list.
Wins compound; scattershot doesn'tOne constraint at a time keeps each build standing on the last one's assets: the map habit, the context files, the platform experience. Scattershot builds compound maintenance instead.
Route, pilot against the baseline, harden, measure. The steps that feel skippable on workflow five are the ones that fail on workflow six.
Some workflows a standard platform covers, some arrive as finished products, and for some the right answer is a specialist who builds and hardens it to fit. The route logic from Stage 2 decides, workflow by workflow, never as a firm-wide policy.
Unmonitored automations degrade quietly, and the cheap ones degrade quietest, because nobody is looking.
The Friday pipeline digest broke in March and nobody noticed until May, because a missing nice-to-have looks exactly like a quiet week. The fix took an hour. The two months of decisions made without it did not.
An alert or a scheduled look, sized to its stakesThe client-facing build gets an alert the moment it fails. The digest gets a monthly glance. What nothing gets is no watcher at all.
The portfolio board from the figure, kept honest: no orphans, nothing unwatched, nothing that neither pays for itself nor gets retired.
When enough runs that keeping the list current is itself a job, you have arrived at the operating engine. That is Stage 4.
Grow one constraint at a time, run the whole loop each time, and watch everything you leave running. The current, owned list of what runs is the door into Stage 4.
Operating well here is a complete destination. It looks like: several workflows running under real harnesses, an autonomy log that says what earned trust, and monitoring that catches drift before clients do. Stage 4 is there when keeping the whole picture current becomes its own job, not because a next stage exists.
What changes when AI stops being a project and becomes something you rely on. What this stage builds is light: documents and routines, not platforms, and far lighter than the governance genre implies. Most firms have none of this, which is exactly the opportunity.
If you self-located straight to this stage, here is what the engine gathers up: the personal collections of prompts and instructions become one shared library, the notes on each automation become one list, the guardrails around each build become one page of rules, and the measuring of each workflow becomes one review of the whole. Each earlier stage teaches its piece in full; this stage only assumes they exist somewhere, however informally.
The operating engine starts with two short documents. The policy says what your firm does and does not do with AI, in six short sections. The register is the single list of everything that runs, one entry per automation. The policy is the rules; the register is the reality. Together they answer the two questions every serious client, insurer or regulator asks: what do you allow, and what actually runs?
Most firms have none of this, and not only small ones: ask a leadership team anywhere what AI actually runs in their business today, and few can answer with confidence. The governance literature makes the fix sound like a program with a steering committee. It is an afternoon: six sections on one page, one list, one review date. Light is the point, because heavy governance dies of neglect.
The entries are this firm's; the labeled fields are the part to copy, exactly as they appear here. Keep the list wherever the firm keeps its working documents; what matters is that there is exactly one.
What your firm does and does not do with AI, in six short sections.
One entry per automation, seven fields each, as in the figure. No automation without an entry, and new builds start by adding one.
One list, not severalKeep it wherever the firm keeps its working documents. The place matters less than there being exactly one, because two lists disagree the week after they are created.
Every automation needs one more decision: which parts of the work it does by itself, and which parts it hands to a person. The answer goes on its register entry.
There are two easy settings, and both waste the automation. If it does everything by itself, the firm carries risk nobody agreed to. If a person has to approve every single thing it produces, that person becomes the bottleneck, and most of the time saved is lost again. The setting that works for most work sits in between: the automation handles the routine cases by itself and hands the unusual ones to the person named on its entry.
The line moves as trust growsStart careful. When the reviews show corrections are rare, let it handle more by itself. Move the line at the review, based on the numbers, and update the entry so the register always says where the line sits today.
The support inbox triage from Stage 3 began by proposing replies for a person to send. Because corrections stayed rare, it now sends the routine replies itself, and still hands complaints, legal questions and anything it has not seen before to the support lead.
An automation nobody is responsible for runs until it breaks, and then fixing it is nobody's job, which in practice means it stays broken or quietly keeps producing the wrong thing.
One person, who knows the workflow and answers for itWhen that person leaves, the entry gets a new name the same week or the automation gets retired. No third option.
The documents your automations rely on, like a price list or a description of your services, get the same discipline: name who may change them, because a quiet edit there changes what everything built on them does.
The "how we price" document feeds three automations, and the register names the operations lead as its owner. When a partner wants a rate changed, the change goes through her, and all three automations change behavior on the same deliberate day instead of drifting one by one.
The biggest software vendors have started building this same kind of inventory into their own products, because even they found that nobody could reliably say what was running. Yours can be far simpler and do the same job.
One page of rules, one list of what runs, both owned and current. This is an afternoon's work, and it puts you ahead of most firms of any size.
Every AI tool your firm uses does its work somewhere: on a provider's computers, or on your own. Your email almost certainly runs at a provider; your files might, or they might sit on a machine in the office. Nobody agonized over those choices; each followed a real need. AI is the same decision in new clothing, and by this stage the choices add up to how much control, cost and dependency the firm carries. The rule: for everything that runs, you can say where it runs and why.
Most tools are services you rent: the provider runs everything, you start through an account, and your protection is the contract. Some can run on computers you own: real effort to set up, someone has to keep it healthy, and in return nothing leaves machines you control. Between the two sit real middle options, and most firms end up using several at once. The point is not to pick a camp; it is to match each automation to the protection it actually needs. Paying for maximum control everywhere means paying for a constraint you do not have.
The placements are this firm's; the columns are the part to copy: where, why, whose job, and the way out. Most firms' honest page looks like this one: nearly everything as a service, one or two hard constraints handled deliberately.
Everyday work as a service; a hard constraint earns a more controlled option. The right answer depends on the work, not on fashion.
The middle options are realA private space at a large cloud provider, where the software runs on their machines inside a walled-off area only your firm can reach. A provider's tool configured so your data stays in your own systems, or in your own country. Most firms end up using several options at once, deliberately.
The documents your tools read before doing any work deserve their own decision, because every tool depends on them.
Ordinary files, in a folder the firm controlsOrdinary files can be read by any tool, moved anywhere, and survive every change of provider. Keeping that knowledge inside one tool instead, in its settings or workspace, is more convenient day to day, and it quietly ties you to that tool: switch providers, and the firm's accumulated knowledge has to be rebuilt from scratch.
For a service, the provider's. For anything you run yourself, someone at your firm, and that someone is named before the breakage, not during it.
Providers raise prices, retire products, and get bought. Leaving calmly takes three things you can arrange today.
Pick any automation on the register and the page answers: where does it run, could you leave, could you change calmly?
When a provider updates or retires something, the register tells you exactly what to re-check. That is what makes provider news an item on a list instead of a small crisis.
You will hear more and more about systems where many agents pass work between themselves. The idea is real, and most firms should not start there.
A fleet multiplies everything this playbook taught you to keep in check: what the system can reach, what it can spend, and what can go wrong while nobody watches. The honest readiness test: your individual agents run reliably, they rarely need correcting, and your list of what runs is current. Until then, several simple agents doing one job each beat one clever fleet.
Match each automation to the protection it actually needs, keep the firm's knowledge in files you own, and know for everything that runs where it runs, who cares for it, and how you would leave. These are decisions, not projects.
The engine is documents and routines, and documents and routines only stay alive when people look after them. The people plan is deliberately small: give every standing piece one named person, and lead adoption by example rather than by announcement. The reason for the names is simple: a responsibility held by "the team" is held by nobody.
Adoption is cultural before it is technical. Announced tools get quietly ignored; demonstrated tools get copied. The habits in this article are deliberately cheap, because the expensive version, a rollout program with mandatory training, mostly produces attendance.
The list needs its reviews to happen, the shared collection needs occasional pruning, and people's questions need somewhere to go. Each gets a name, folded into an existing job.
Adoption follows what leaders visibly do, not what they approve. A leader who uses the tools in the open gives everyone else permission by example.
Five minutes on Monday beats the training budgetA partner opens the Monday meeting with the brief she used to draft a client memo, including the two corrections she had to make before it was right. The opposite is just as visible: a leader who announces AI and never touches it teaches the firm that the tools are for other people.
Each person automates the task they like least, self-chosen, so nobody's work is automated at them.
Whatever proves itself goes into the shared collection with the builder's name on it. The credit is not decoration; it is why the next build day has volunteers.
"Show me how you work with AI" says more than any line on a resume, and it tells candidates what kind of firm they are joining.
Starting with onboarding, because it is the people process every firm has and few maintain.
The firm's onboarding manual, converted by a model into short modules with a quick check after each one and a first-day practical task marked against a checklist. A new hire's gaps show up as flagged answers to talk through together, and the manual finally gets read.
The people side is working when three questions get answered without a pause.
A named person for every standing piece, leaders who visibly use the tools, and people who learn by building. The people side costs hours a month, and without it the documents go stale.
Earlier in the playbook you measured one workflow; the engine measures all of them. The discipline is the same one that proved your first workflow: compare what the automation delivers against what the work cost before, and decide from the number, not the mood. What changes at this stage is only the scope: everything the firm runs, together, on a schedule it cannot skip.
These numbers are what earn the engine its budget. Measured wins argue for the next build better than any enthusiasm, and the retired automations are what make the wins credible: a list where everything supposedly works convinces nobody.
Decisions dated, written into the register, retirements included. Quarterly works for most firms; look monthly at anything new or busy, and twice a year is enough for the stable and quiet.
Results and costs since last time, taken from the automation's own records. The responsible person brings the numbers; the meeting reads them.
Automations fade as inputs, tools and models shift around them. The review is where fading gets caught on a schedule instead of by accident.
Flat is a finding tooAn automation that merely matches its launch quarter while volume grew has quietly gotten worse. The comparison against its best stretch is what shows it.
Keep it, fix it, or retire it, dated, in the register, retirements included. Retiring things is normal.
The review is not just a scorecard; each of its three outcomes leaves the firm with something useful.
Three plain questions per automation: what did it deliver, what did it cost, is that worth keeping. Set the dates in advance, and let nothing skip its turn.
Sized to the automation, not to ambitionQuarterly for most. Monthly for anything new or busy. Twice a year for the stable and quiet. If the review takes more than a morning, the measuring has grown too clever.
Three signs the review is honest: nothing gets skipped, fading gets noticed, and the decisions are written down, dated, retirements included.
On a fixed schedule, for every automation: what did it deliver, what did it cost, keep it or not. A firm that never retires anything is admiring its automations, not measuring them.
Sooner or later someone will ask what your firm does with AI: a client before they sign, an insurer at renewal, eventually a rule. A clear, unhesitating answer signals exactly what the person asking wants to know: that someone here is in charge of this. And because most firms cannot yet answer the question well, answering it plainly sets you apart more than the tools themselves do. The rule: whoever asks, the honest answer already exists in writing.
If your people picked up the habit early of noting what a model touched in each piece of work and who checked it, none of this is new effort. The engagement letter simply says at firm level what the work already says internally, and the honest sentence writes itself from practice.
The questions are typical of the lists clients and insurers now send; the answers are one firm's. What transfers: every answer names the standing document it came from, and none required a meeting.
Clients are told, in the letter, which tools touch their work and what they are used for. A vague "we may use AI" line protects nobody and reads like it is hiding something.
Name the tool, name the use, state the training answer"Drafting and document review are assisted by [named tool] under our review; your data is not used to train it." One sentence, and the most common question is answered before it is asked.
Clients and insurers increasingly send question lists: what do you use, what does it touch, who checks the output. The register already holds those answers.
A short summary, kept currentOne page distilled from the register, refreshed when the register changes. It turns a week of scrambling into a same-day reply, as in the figure.
Less than the headlines suggest. For a typical services firm using general-purpose tools, the current rules come down to two things: your people know how to use the tools responsibly, and nobody is misled about dealing with AI.
Chatbots say they are chatbots, and AI-made content is labeled as such. The heavy obligations target uses like automated hiring decisions, and they sit mostly with the companies that build the models, not with firms that use them. A short internal training and the one-page policy from the start of this stage cover most of what is asked.
This is the one part of the playbook that touches law, and law moves. Treat this as orientation as of mid-2026, not as advice; your counsel has the final word for your firm.
Readiness is checkable, and all three checks read from documents that already exist.
An honest letter, ready answers, and people who can say what they use. Firms rarely lose work for using AI; they lose it for being vague about it.
The engine is a steady state, and it is deliberately light. Operating well here looks like: the list of what runs is current, the numbers are honest, the answers are ready, the library keeps growing, and none of it needs a committee. This is the destination for firms with enough running; there is nothing above it to climb to.
The first three stages run the same six-move loop at rising scale: on your own recurring work, then on one shared workflow, then across many workflows with automation and agents. The fourth stage is different in kind: the operating engine, the light set of standing rules and shared documents that keeps everything you've built running and compounding. Six principles sit underneath all of it.
What actually breaks first, or drains the most hours. One bounded task, not a whole function.
Watch the process as it really happens, capture a baseline, and name one owner.
Pick the simplest option that clears the bar; move up a rung only when the one below runs out.
Run it beside the current way, with the go/no-go rule decided in advance and a time limit.
Add the guardrails, name the owner, and make it the default way the work gets done.
Track verified outcomes, not usage. Retire whatever stops paying its way.
Start with the bottleneck you actually want to clear, and let that choose the tool, not the other way around.
An AI that runs on its own is only as reliable as the process behind it, so define the workflow as clear steps and rules first, then let it take over.
Let the system do the work, but a person reviews and signs off on anything a client sees or that would be costly to get wrong.
What makes an automation safe to trust is the structure around it: spending limits, clear boundaries, an off-switch, and a record of what it did.
When a model gives you an answer, make it cite and check its sources, because you stay accountable for the result.
Once something works, capture it as a repeatable process, so the next time is faster and someone else can run it without you.
Bring the workflow that bothers you most. We map where the work leaks, tell you what we would run the loop on first, and you keep the notes either way. No deck, no obligation.