← Back to all field notes

Which Repetitive Work Should an AI Agent Take First? Use These Five Tests

Start from a task rather than a job title, then score business value, frequency, knowledge readiness, controllable risk, and process ownership.

Companies increasingly ask whether AI can enter a real workflow and remain useful, not whether someone can build another demo. The choice of pilot matters more than the choice of tool. A poorly selected pilot turns even a strong model into internal presentation material.

Start from a task, not an entire job. Look for a bounded group of repetitive actions where people currently spend time gathering information, applying rules, preparing a decision, or transferring context.

Three promising process types

1. Knowledge-intensive work with scattered rules

Examples include standard operating procedures, exception handling, customer answers, and cross-team handoff instructions. These processes quickly reveal whether company knowledge is actually readable and maintainable.

If naming, versions, and ownership are inconsistent, AI will amplify the disorder. The pilot may still be valuable because it forces the team to build a better knowledge structure.

2. High-frequency coordination

Status updates, reminders, material preparation, and repeated information transfer are often good starting points. They produce enough samples for a short pilot and their outputs are usually easy to inspect.

3. Decision preparation with a human final call

An agent can extract fields, retrieve rules, show conflicts, and prepare a recommendation while a person keeps the final authority. This is often safer and more informative than attempting full automation immediately.

Score the candidate on five dimensions

Use a 0–2 score for each dimension:

Dimension012
Business valueoutcome is unclearimproves a local experienceaffects an observable business result
Frequencyoccasionalweeklydaily or every business batch
Knowledge readinessmostly tacitdocuments exist but versions are messyrules and exceptions are traceable
Controllable riskfailure is hard to reversea person can reviewthe workflow can stop and roll back safely
Process ownerno one decidessomeone coordinatesa named owner controls rules and outcomes

The highest score is not automatically the best choice. A zero in controllable risk or ownership is a warning that should be repaired before development.

Define the smallest useful loop

A first pilot should fit into two to four weeks and answer one practical question. For example:

Can the system receive a freight inquiry, identify missing fields, retrieve the relevant rules, prepare a review package, and hand exceptions to the correct person?

This is much more testable than “automate sales.”

Define:

  • the input and sample;
  • the output and acceptance rule;
  • the human baseline;
  • the actions the agent may take;
  • the conditions that stop the workflow;
  • the owner who reviews exceptions;
  • the evidence collected after each run.

Avoid the easiest-looking process

The easiest demo is not always the best pilot. A task may be easy for a model but too infrequent to evaluate, too far from business value, or owned by no one.

Choose a process that is both bounded and consequential enough for the team to care about the result.

Use the free AI pilot priority scorecard to compare up to three candidates. Then continue through the pilot design and evaluation topic.

Continue reading
Pilot Design & Evaluation
Is the AI Employee Worth It? Calculate Hours Saved Against Money Spent Do not judge an agent by feeling. Three numbers — human hours saved, money spent, task coverage — plus one J-curve decide whether to keep it or cut it. Choosing AI Tools and Models Is a Decision Framework, Not a Shopping Trip Do not pick the most expensive model, and do not pick the most hyped tool. Grade tasks, match capabilities, and calculate total cost. Enterprise AI selection is a repeatable decision framework. The AI Employee Works, a Human Supervises: How Agents Upgrade Themselves An agent should not be written once and frozen forever. Humans review its work log, feedback turns into improvements automatically, and the agent gets better over time. Self-iteration is how agent value keeps flowing.