Document Extraction: Templates, Trained Models, and Agentic Pipelines
How the three generations of document extraction actually work — zone templates, trained IDP models, and agentic LLM loops — and which failure modes each one carries.
The useful distinction between extraction approaches is not which one uses AI, but what each one does when it is wrong — templates fail silently, trained models fail confidently, and agentic pipelines are built to catch their own errors before you see them.
Most comparisons of document extraction stop at "templates are rigid, AI is flexible." That was a fair summary in 2022. It no longer describes what the current systems actually do, and it leaves out the mechanism that matters most when you are deciding what to buy or build: error handling.
Here is how each generation works, in enough detail to evaluate one.
Generation one: zone templates
A zone template is a coordinate map. You open a representative document, draw rectangles over the fields you care about, and bind each rectangle to a target field. On every subsequent document, the engine runs OCR, then reads whatever text falls inside those coordinates.
The model has no concept of what a purchase order number is. It knows that on this vendor's document, the PO number lives 4.2 inches from the left edge and 1.8 inches from the top.
This works remarkably well under one condition: the documents are genuinely identical. Government forms, standardized claim documents, and machine-generated statements from a single system all qualify.
The failure mode is silence. When a layout shifts, the zone still returns text — just the wrong text. A template expecting a PO number reads the quote number that moved into that position. Nothing errors. The field is populated, the confidence score refers only to OCR character certainty, and a wrong value flows downstream looking exactly like a right one.
The maintenance burden compounds with a second-order effect people underestimate: you need a template per layout, and you need to know which template to apply before you can extract anything. That means a document classification step, which is itself a thing that can be wrong.
Generation two: trained IDP models
Intelligent document processing replaced coordinates with learned features. Instead of "read the box at 4.2 inches," the model learns statistical associations — that a value near text reading "P.O. #" or "Order No." and formatted as six to ten alphanumerics is probably a purchase order number.
This is a real improvement. The model generalizes to layouts it has never seen, within the document class it was trained on. Commercial pretrained models for invoices, receipts, and IDs work well out of the box because those classes are visually consistent across the world.
The failure mode is confident generalization. A model trained on invoices applied to purchase orders will still return an answer. It will find something PO-number-shaped and something table-shaped and report reasonable confidence, because confidence is calibrated on the training distribution, not on how far your document sits outside it.
This is worth stating plainly because it is an easy and expensive mistake: running a prebuilt-invoice model against purchase orders is not a fair test of that model, and it is also not an unusual thing for a team to do by accident. Document class matters more than vendor choice at this tier.
The other constraint is table extraction. Header fields are largely solved. Line items are not. Multi-page tables, rows that wrap, merged cells for long descriptions, subtotal rows interleaved with item rows, and continuation markers are where second-generation systems lose data — often by returning a partial table rather than failing.
Generation three: agentic extraction
"Agentic" is doing real work as a term here, though it is often used loosely. The distinction is not that a language model is involved. It is that the model runs inside a control loop with verification steps, rather than being called once and trusted.
A pipeline that sends a page image to a vision model and parses the JSON that comes back is a single-pass LLM extraction. Useful, but it inherits the second-generation failure mode with worse calibration — language models are fluent, and fluency reads as confidence.
An agentic pipeline adds four things:
1. Constrained decoding. The model emits into a schema rather than into free text. Quantity is typed as a number, unit of measure is an enum, dates are ISO-8601. Structurally invalid output cannot be produced in the first place, which eliminates a whole class of parse errors that single-pass pipelines handle with retries and regex.
2. Grounding. Every extracted value carries the region of the page it came from — a page index and a bounding box. This is the load-bearing feature, for two reasons. It makes human review fast, because a reviewer checks a highlighted region instead of hunting the document. And it makes hallucination detectable: a value with no corresponding region on the page is, by construction, invented.
3. Self-verification. The pipeline checks the extraction against the document's own internal consistency. Do the line extended prices sum to the stated subtotal? Does subtotal plus tax and freight equal the stated total? Does the line count match a stated count? Does every row have a quantity? These checks need no external data and no ground truth, and they catch the dropped-row failures that generation two produces silently.
4. Targeted re-reading. When a check fails, the loop does not discard the result or escalate the whole document. It re-reads the specific region that failed — often at higher resolution, or with the failing constraint stated in the prompt — and re-runs the check. This is where the extra cost and latency of agentic extraction goes, and it is spent only on the documents that need it.
The result is a system whose errors are mostly converted into flags. A document where the totals will not reconcile after re-reading gets routed to a human with the specific failing line marked, rather than posting a quietly wrong order.
What agentic extraction does not solve
Grounding and self-verification address whether the extraction faithfully represents the page. They say nothing about whether the page maps onto your business.
A line reading CHOC BAR DARK 70% 12CT can be extracted perfectly — correct string, correct quantity, correct bounding box, all arithmetic reconciling — and still be unusable, because nothing in the document says it corresponds to SKU MV-DK70-12 in your catalog, that this customer orders by the case, or that their contract price is 28.50.
That resolution layer is separate from extraction and is usually where B2B order automation actually succeeds or fails. We covered the split in document parsing APIs versus order extraction.
Comparison
| Zone templates | Trained IDP | Agentic | |
|---|---|---|---|
| Generalizes to new layouts | No | Within document class | Yes |
| Setup per layout | Template build | None if class matches | None |
| Evidence for each value | No | Sometimes | Bounding box per field |
| Detects its own errors | No | Weakly, via confidence | Schema, arithmetic, re-read |
| Typical failure | Wrong value, no signal | Plausible value, high confidence | Flagged for review |
| Line-item tables | Fragile | Partial extraction common | Verified against totals |
| Cost per document | Lowest | Low | Higher, variable |
| Latency | Fastest | Fast | Slower, varies by re-reads |
How to evaluate a claim
Extraction accuracy figures are close to meaningless without their denominator. Three questions make a vendor number interpretable:
- Field-level or document-level? A system with 99% field accuracy across 40 fields gets the whole document right about 67% of the time. Vendors quote the first number; your operation experiences the second.
- Are line items included? Header-only accuracy on purchase orders is the easy half. Ask for the rate on documents with more than twenty line items spanning a page break.
- What happened to low-confidence documents? A 99% figure computed after excluding everything the system flagged is measuring a different thing than end-to-end accuracy. Straight-through rate and accuracy-on-processed are both needed; either alone is misleading.
The same questions apply to your own testing. Score against documents that were never used to build or tune the system, and score line items separately from header fields — mixing them hides the failure that costs you money.
Frequently asked questions
Is agentic document extraction different from using GPT or Claude on a PDF?
Yes, though the same underlying model may be involved. Sending a page to a vision model and parsing the response is a single pass with no verification. Agentic extraction wraps that call in schema constraints, per-field grounding, consistency checks, and targeted re-reads when a check fails. The loop is the difference.
Do templates still make sense for anything?
They do, when documents are genuinely identical and high-volume — output from a single upstream system, standardized regulatory forms. Templates are cheap and fast under those conditions. The economics invert as soon as you have many senders with independent layouts.
Why does grounding matter if the extraction is already correct?
Because you cannot tell that it is correct without it. Grounding is what makes review fast enough to be worth doing and what makes an invented value distinguishable from a read one. It is also what lets you audit a decision months later.
Does agentic extraction eliminate human review?
It changes what review is for. Instead of checking every document because any of them might be wrong, reviewers handle the subset the pipeline flagged, with the specific failing field highlighted. Review effort concentrates where the uncertainty actually is.
How should I test extraction on my own documents?
Take a sample that includes your messiest real inputs — scans, multi-page tables, documents from your newest trading partner — and hold it out entirely from setup. You can run POs through the free PO Extractor to see grounded extraction on your own files, and compare against how your current system handles the same set. Related reading: AI order processing versus OCR and document automation for B2B order processing.
How to Answer "Are You EDI Capable?"
A buyer asked, and you need to answer this week. The four things they are checking, what you can say yes to today, and a realistic date for the rest.
- What a buyer is really asking when they ask this
- The minimum set-up that makes the answer yes
- What you can answer today versus what needs building
- Rough timelines, so you can give a date rather than a maybe
One email with the download. Unsubscribe any time.
Stop manually entering orders
OrderSync turns EDI, email, PDF, and fax orders into structured data automatically. See how it works for your business.
Related Articles
Convert PDF Purchase Orders to JSON (or EDI) via API
How to turn PDF purchase orders into structured JSON or compliant EDI through an API, what the response looks like, and how to handle scanned and low-confidence documents.
TechnologyDocument Parsing API for B2B Orders vs Generic OCR
Generic document parsing and OCR APIs read fields off a page. B2B order processing needs catalog matching and validation. Here is where the two diverge and which you need.
TechnologyPurchase Order API: Extract PO Data Programmatically
How a purchase order API turns PDF and email POs into structured JSON your systems can post. What to expect from the endpoint, the response shape, and matching.
TechnologyAI Order Entry Systems: How They Work and When to Use One
AI order entry systems extract purchase order data from any format without templates or manual setup. Here is how they work, where they outperform traditional systems, and where they do not.
TechnologyAI Order Agent vs EDI: Do You Still Need EDI?
How AI order agents compare to traditional EDI for B2B order processing, when you need both, and when an AI agent can replace EDI entirely.
TechnologyAI Order Agent vs Manual Entry Compared
A side-by-side comparison of AI order agents and manual data entry for B2B order processing, with real cost, speed, and accuracy numbers.
TechnologyAI Order Processing vs OCR: Key Differences
How AI-powered order processing compares to traditional OCR and template-based extraction, and why AI handles layout variations that break OCR systems.
TechnologyAI-Powered EDI Processing for Small Teams
EDI is mandatory for major retailers but brutal for small teams. AI-powered EDI processing automates validation, exception handling, and ERP sync.
TechnologyMore from the Blog
FSMA 204 Deadline: July 20, 2028, Not January 2026
The FSMA 204 compliance date moved to July 20, 2028. Here is the full timeline, the statutory basis for the delay, and what did not change when the date did.
IndustryFSMA 204 Shipping CTE on an EDI 856: Segment by Segment
Where the Traceability Lot Code and FSMA 204 shipping KDEs actually sit in an EDI 856 ASN. The hierarchical loop, the LIN and N9 segments, and what breaks in practice.
EDICan an Invoice Carry FSMA 204 Traceability Data?
Why the EDI 810 invoice cannot carry lot-level traceability the way an 856 can, what a PDF invoice is genuinely useful for under FSMA 204, and where the reference document KDE fits.
IndustryFSMA 204 KDEs From Suppliers Who Do Not Send EDI
Half your supplier base will still be emailing PDFs in 2028. How to capture FSMA 204 Key Data Elements from invoices, packing lists, and faxes without hand-keying every one.
Industry