In practice
One month of data engineering beat seven months of agent tuning.
A Fortune 100 fintech tried to automate invoice OCR and accounting entries with AI agents. Results stayed mediocre — until the data underneath was rebuilt.
- Client
- Fortune 100 fintech
- Domain
- Accounts payable — invoice OCR and posting
The timeline
Prior attempt
7 monthsAgents automating invoice OCR and posting. Mediocre accuracy, no path to production trust.
Data foundation
1 monthWe organized the data estate: clean datasets and sources of truth the agents could rely on.
Agent workflow
2 weeksThe full workflow deployed — near-perfect production accuracy, outperforming the original attempt.
Since then
Every 2 weeksA new agent ships on the same foundation. The acceleration compounds.
The full account
A Fortune 100 fintech came to us after seven months of trying to make AI agents do real accounts-payable work: read incoming invoices, extract the fields, and apply the correct accounting transaction. The models were capable. The results were mediocre — accuracy that never got good enough to trust in production.
The diagnosis was the one we make on this site every day. The agents were being asked to reason over data that was scattered and inconsistent. No amount of prompt engineering fixes a missing source of truth.
So we didn’t start with the agents. We spent one month organizing the data: building clean datasets and establishing sources of truth for the records the workflow depended on. Unglamorous work. It changed everything.
With the foundation in place, deploying the full agent workflow took two weeks. It outperformed the seven-month attempt immediately, reaching near-perfect production accuracy — because the agents were finally operating on facts, not fragments.
That’s the part most firms miss: the foundation isn’t just for the first agent. Since that deployment, a new agent has shipped roughly every two weeks — each one standing on the same clean data, each one faster to build than the last.
Questions we hear
Production invoice agents, straight answers.
What went wrong with the prior invoice OCR and posting agents?
A Fortune 100 fintech spent seven months automating invoice OCR and accounting entries with AI agents. The models were capable, but results stayed mediocre — accuracy that never reached production trust — because the agents were reasoning over scattered, inconsistent data.
Why did data engineering come before more agent tuning?
No amount of prompt engineering fixes a missing source of truth. One month organizing the data estate — clean datasets and sources of truth the workflow depended on — changed the outcome. Unglamorous work; it made the agents operable.
How long did the production invoice agent take once the foundation was ready?
With the foundation in place, deploying the full agent workflow took two weeks. It outperformed the seven-month attempt immediately, reaching near-perfect production accuracy on facts instead of fragments.
What happened after the first accounts-payable agent shipped?
The foundation was not just for the first agent. Since that deployment, a new agent has shipped roughly every two weeks — each one standing on the same clean data, each one faster to build than the last.
The foundation comes first. Then the agents compound.
A readiness assessment maps your silos, your gaps, and the shortest path to a platform agents can work on.

