By the end of 2027, Gartner expects more than 40% of agentic AI projects to be cancelled — not because the technology failed, but because of rising costs, uncertain business value, or insufficient risk management. That number lands before the argument does, and it points to something I have seen repeatedly: organisations pick a tool before they pick a task.
Someone decides to 'automate customer onboarding.' They choose a platform, build a workflow, and then discover that the onboarding process involves four judgment calls, two approval steps, and a compliance review that nobody wrote down. The automation handles the easy parts; the human still has to check every output. The promised time savings vanish into rework.
Microsoft's Work Trend Index reports that 62% of employees spend a significant portion of their day searching, communicating, and coordinating — a huge reservoir of potential automation. Yet only 11% of UK SMEs use AI extensively to automate operations. The gap is not a technology problem. It is a selection problem.

Success starts with picking the right task, not the right tool. The framework I will walk through is a 7-factor decision matrix that separates high-ROI automation candidates from tasks better left augmented or manual. It was shaped by patterns from over 50 real-world AI-enhanced workflows, and I have tested it against my own failures. It works only when you are honest about each factor.
The seven questions that matter
The framework asks seven questions about a task. Answer each one honestly, and it tells you whether to automate, augment, or keep manual. The questions come from Teresa Torres's work on task triage. I have refined the wording for a knowledge worker's context.
- Frequency — Do you do this daily, weekly, or rarely? High frequency favours automation; a one-off task is not worth the setup.
- Enjoyment — Do you enjoy it, or would you pay someone else to do it? Low enjoyment means you will happily hand it off; high enjoyment suggests you should keep it.
- Complexity — How many steps, branches, and exceptions exist? Low complexity is easier to automate; high complexity often requires augmentation.
- Articulability — Can you describe the steps in a clear, repeatable process? The more clearly you can articulate it, the better a candidate for full automation.
- Judgment need — Does the task require human discretion, context, or nuanced decision-making? High judgment pushes the task toward augmentation or manual.
- Success definition — Can you clearly define what 'done well' looks like? If success is fuzzy, automation will miss the target.
- Risk — What are the consequences of an error? Low-risk tasks (e.g., draft a status update) are safe to automate; high-risk tasks (e.g., approve a loan) need human oversight.
Each factor uses a simple 1–5 scale. A 5 on frequency means 'multiple times a day.' High scores generally favour automation — except for judgment and risk, where high scores push toward manual.
I have seen people skip the articulability question because they think they know the steps. Then they realise there's an unspoken exception — a flag they only notice when the output is wrong. That is where the framework earns its keep: it forces you to surface the unknowns before you code anything.
Four tasks from a busy week
Let us apply the matrix to real tasks a knowledge worker might face. I will score each on the seven factors and show what zone they land in.
| Factor | Email triage | Meeting notes | Research synthesis | Reschedule meeting |
|---|---|---|---|---|
| Frequency | 5 (dozens/day) | 3–4/day | 2–3/week | 1–2/week |
| Enjoyment | 1 (dread) | 2 | 4 (I like it) | 1 |
| Complexity | 2 (mostly pattern-based) | 3 (speakers, topics, action items) | 4 (multiple sources, integration) | 2 (rules-based) |
| Articulability | 5 (clear rules: flag, label, archive) | 3 (partial: transcription is clear, but prioritisation is not) | 2 (hard to describe synthesis process) | 5 (straightforward calendar logic) |
| Judgment need | 2 (few judgment calls) | 3 (which action items matter?) | 5 (interpretation and relevance) | 1 (no judgment) |
| Success definition | 5 (inbox zero, correct flags) | 4 (good enough summary, key actions) | 2 (quality is subjective) | 5 (new time works for both) |
| Risk | 2 (low: missed message can be caught) | 1 (low) | 1 (low: no immediate consequence) | 2 (low) |
Email triage and rescheduling meetings are clear automation candidates: high frequency, low enjoyment, high articulability, low judgment, clear success, low risk. Meeting notes falls into augmentation — the transcription part can be automated, but deciding which action items matter still needs a human. Research synthesis stays manual because the high judgment need and fuzzy success definition make it a poor fit. The matrix saves you from automating something you will end up rewriting from scratch.
The numbers that get left out of the press release
The headline figures sound impressive. According to a McKinsey/Slack survey, knowledge workers using AI agents recover a median of 6.4 hours per week per seat. A Federal Reserve study pegs the saving at about 2.2 hours in a 40-hour week. I do not buy either number as a guarantee. The same sources — Forrester TEI — report that unmeasured human rework absorbs 22–38% of those time savings in mature deployments and over 50% in early-stage ones. That is the part that gets left out of the press release.
Here is what the rework figure actually means. When you automate a task, the system produces an output. Someone still has to review it, fix edge cases, and handle exceptions. If the task was poorly chosen — say, a process with fuzzy success criteria — the review can take longer than doing the task manually. The matrix lowers this risk by flagging tasks with unclear success definitions or high judgment needs before you invest in automation.
Only 41% of agent rollouts cross positive ROI within 12 months, and 19% never reach payback — primarily because of evaluation drift and unmeasured rework, not agent capability. Customer service is the exception — 63% reach payback within a year, with a median payback of 4.1 months. That makes sense: customer service queries are highly structured, success is binary (resolved/not resolved), and the cost differential is enormous — $0.46 per ticket for AI versus $4.18 for human handling. The matrix works because those tasks score well on almost every factor.
Automate, augment, or keep manual
- Automate — High frequency, low enjoyment, low complexity, high articulability, low judgment, clear success, low risk. Hand it to a machine.
- Augment — Moderate scores: some judgment needed, success partially clear. Use AI to produce a draft, but keep a human in the loop.
- Manual — High judgment, fuzzy success, or high risk. Keep doing it yourself. The matrix saves you from creating more work.
Once you have identified which tasks to automate, the next step is choosing the right tool. See our category-by-category comparison of the best AI productivity apps to pick the tool that matches your task profile.
A reference card and a quick-start template

Use the image above as a printable reference. Here is the quick-start:
- List 5–10 recurring tasks from your week.
- Score each on the seven factors using a simple 1–5 scale.
- Map the scores to the three outcome zones using the criteria above.
Start with the ones that score high on frequency and low on enjoyment — those will give you the biggest motivation to follow through. The matrix only works if you are honest about judgment need and risk. I have seen people overestimate articulability and underestimate risk, and every time the rework came back to bite them.
For a broader view on combining tools and workflows after you have identified the tasks, check out our guide on building an AI productivity stack.
Comments
Join the discussion with an anonymous comment.