How to find where AI can help in operations
Most of the work AI could take is invisible from a leadership meeting. It is spread thin: 20 minutes here, an hour every Friday, a report 4 people each rebuild from the same exports. None of it belongs to anyone in particular, so none of it comes up.
The work that does come up, the drafts nobody wants to write and the tickets somebody sorts by hand, is usually right about the work. It is a genuine fit. It is usually wrong about where to start, because work visible enough to be named in a meeting is work someone has already built a workaround for.
Where AI fits tests one candidate at a time. This chapter produces the candidates to test.
Write down the work
Section titled “Write down the work”Starting from the AI side produces a generic list. “What could AI do for us” comes back as drafts, classification and summaries: true of every company, and it names no actual work inside this one.
Starting from the work produces something usable: where the day actually goes, and which of it is dull, slow, or done twice. 4 ways in, each turning up a different kind of candidate. A thorough search uses all 4.
By function
Section titled “By function”Walk the org chart. For each function, 3 to 5 sentences on what takes the time. Not job titles, work shapes. The unit is “the team spends X hours a week on Y.”
Most of what AI can plausibly help with falls into a few shapes that repeat across companies:
- Sales. First-draft outreach, prospect research, call notes turned into CRM fields.
- Marketing. Copy drafts, audience research pulled together, brand-voice rewrites.
- Operations. Process write-ups, meeting notes distilled, ticket routing.
- Customer support. First-response drafts, sorting (urgent or routine, billing or product), long threads summarised.
- Finance and admin. Expense categorisation, invoice extraction, contract clause checks.
- HR. Job descriptions, screening summaries, onboarding packets.
- Engineering and IT. Code drafts, log analysis, internal tools, documentation.
A real company’s version will hold shapes that are not on that list and miss several that are. The function pass exists so that no whole part of the business goes unlooked at, not to be complete on its own.
By recurring meeting
Section titled “By recurring meeting”A standing meeting marks work the company has decided is worth talking about repeatedly. In the weekly pipeline review, someone spends an hour preparing a status that takes 10 minutes to deliver. In the monthly ops review, the same numbers get compiled fresh from the same 4 systems. In the quarterly review, 2 slides always take a week.
A meeting that exists to recover information already sitting in a system is usually covering for a candidate.
By the questions leaders keep getting asked
Section titled “By the questions leaders keep getting asked”Every function lead answers the same handful of questions. What is the status of X. How are we tracking on Y. What did the customer say about Z. Asked every week, those questions are a candidate: usually a summary, or a small tool that answers from the system holding the answer and shows where it got it (see tools and memory).
This pass catches what the function walk misses, because the work happens to the lead rather than inside the function.
By the bottlenecks people complain about
Section titled “By the bottlenecks people complain about”Friction is signal. Work that piles up, handoffs that stall, the person who says “I lose half of Monday to this” about a part of Monday they do not enjoy.
This pass tends to turn up the candidates with most to gain, because the pain is already identified and nobody has framed it as an AI candidate yet. It also catches handoff work, the converting a person does between 2 teams, which sits on no org chart and so never shows up in the function walk.
What comes out is a list of 20 to 40 work shapes, in the company’s own words. Most will not be AI candidates. That is fine.
Screen the list
Section titled “Screen the list”The screen is the 4 properties: does the work happen often, is it made of words, is an approximate answer useful, does nobody enjoy it. Run each item past all 4 and mark the ones that come back a no.
The pass here is coarse on purpose. Its job is to get from 30 items down to a handful, not to be right about any single one. The careful version happens later, one candidate at a time. Whether an approximate answer is useful is the property that gets waved through at this stage and the one that sinks candidates later, so an item that passes it easily is worth marking for a closer look.
The same chapter names 5 patterns that sink candidates which passed all 4: safety-critical work with nobody checking, exact-number work, work where the same input must always give the same output, output nobody can verify, and knowledge nobody has written down. An item that trips one is marked harder than it looks, not struck off. Some of them work once the right checking is around them. Some do not.
Most of the list falls away here.
Rank what survives
Section titled “Rank what survives”The temptation now is to pick the highest-value candidate. That is usually the wrong move.
The first pilot’s job is not to earn the most. It is to teach the company how AI lands inside its own work: what changes, how the team responds, what has to be checked, what breaks. The candidate that teaches most for the least risk is the right first one, even when it is not the biggest prize on the list.
4 things to weigh:
- Value. Time, money or quality if it works. Order of magnitude is enough. Precision is wasted at this stage.
- Effort. What it takes to stand up: an off-the-shelf product, assembled tools, or a custom build. Time to first output is the proxy.
- Risk. What happens if the AI is wrong and nobody catches it. An internal draft is low. A customer-facing auto-reply is high. An input to a financial decision is higher still. (Building trust in AI covers what an uncaught mistake costs and the checking that fits each level.)
- Reversibility. Whether the plug can be pulled. Tools that sit alongside current work usually can be switched off. Tools that replace a process usually cannot.
There is no formula here. A score out of 10 would be false precision on 4 judgement calls. What the ranking looks for is workable value, low effort, a cheap mistake and an easy way out. That combination is what a company can run a real experiment on without breaking anything.
Pick one
Section titled “Pick one”3 or 4 candidates cluster at the top. One gets picked.
The question is not which has the best case on paper. It is which produces a clear yes or no fastest. A pilot that runs 6 months and ends in “it sort of helped, hard to tell” costs more than the time and the money. It costs the attempts after it, because an unclear result is hard to get funded a second time.
Picks that answer clearly share 3 things. Somebody measured the before-state: current time per unit, throughput, or the quality bar without AI (measuring AI results covers what to measure). The scope is one team and one shape of work, so not “AI for support” but “first-draft replies to tier-1 tickets in the SaaS line”. And a wrong answer gets caught: internal drafts, sorting with a review step, summaries with the source attached.
What is left at the top of the list is the queue behind it. Not abandoned. Sequenced.
What comes out of it
Section titled “What comes out of it”A ranked list, a first pilot, and an order for the rest. Not an answer.
The list finds candidates. It does not clear them. The candidate that scores 4 for 4 on the properties is the one the next chapter takes apart, and what it fails on is invisible from an afternoon of asking people where their day goes.
The search itself is cheap: an afternoon of asking, an hour of screening, one table. Companies skip it because the visible candidate is already sitting there and already looks like a decision.