Enterprises have spent an estimated $30–40 billion on generative AI, and 95% are getting zero return, according to MIT's Project NANDA. The same report holds the clue to the 5%: the clearest returns came from back-office automation — boring, measurable, close to structured data — while roughly half of AI budgets flowed to sales and marketing projects that convert the least.
Use-case selection is doing more work than model selection. For a first agent, it is nearly everything.
The five-part filter
Run every candidate workflow through five tests. A first agent should pass all of them.
- Narrow — one workflow, one team, one definition of done. "Automate finance" fails; "match invoices to POs and route exceptions" passes.
- Repetitive — high volume of similar cases. Repetition is what turns agent accuracy into money.
- Data-adjacent — the records it needs already exist in systems you can reach and, ideally, trust. The closer to clean data, the shorter the foundation phase.
- Measurable — a number exists today: cost per invoice, hours per reconciliation, exception rate. If you cannot measure the before, you cannot prove the after.
- Reversible — early agents should draft, queue, and recommend. Irreversible actions come after accuracy is earned.
What passes, what fails
| Candidate workflow | Verdict | Why | |---|---|---| | Invoice extraction + PO matching | Strong first agent | Narrow, repetitive, measurable, reversible via review queue | | Contract triage and clause extraction | Strong first agent | Document-heavy, bounded, easy to audit | | Support ticket classification + routing | Good first agent | High volume, clear metric, low blast radius | | Open-ended customer support agent | Avoid first | Unbounded inputs, brand risk, hard to measure | | "AI assistant for the whole sales org" | Avoid first | No single workflow, no metric, no owner | | Cross-ERP reconciliation with no source of truth | Fix data first | The agent would inherit the conflict, not resolve it |
The pattern: first agents succeed where the work is structured and the blast radius is small. They fail where the use case is really a strategy deck wearing a workflow costume.
Build, buy, or wait
Three honest options, and the filter tells you which applies.
Buy when the workflow lives inside one vendor's ecosystem and runs on data you already trust. Off-the-shelf agents are real leverage there.
Build when the workflow spans systems no vendor owns, or when accuracy depends on your specific sources of truth and definitions. That is most consequential enterprise workflows — and it is where a documented invoice-agent engagement went from seven stalled months of agent tuning to a two-week production deployment, once one month of data-foundation work came first.
Wait when the filter fails on data-adjacency — critical records conflict and no source of truth exists. Waiting does not mean idling. It means doing the readiness work that converts a bad first use case into a good second one. The ten readiness questions will tell you which situation you are in.
Key points
- Use-case selection predicts outcomes better than model selection — the 5% pick boring, measurable, back-office work.
- Five tests: narrow, repetitive, data-adjacent, measurable, reversible. First agents pass all five.
- Buy inside one ecosystem on trusted data; build across systems; fix data when records conflict.
- Judge the first agent by what it unlocks — the second agent should ship faster on the same foundation.
The bottom line
The right first agent is unglamorous: a bounded workflow, a number you can move, data you can reach, and actions you can undo. Pick that one and the pilot-to-production funnel inverts. Pick the impressive one and you become next year's cancellation statistic.
Sources: MIT Project NANDA — The GenAI Divide: State of AI in Business 2025 (July 2025); Gartner — Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 (June 25, 2025).


