Trust an extracted table when its values, relationships, and business interpretation survive separate checks. OCR accuracy alone cannot establish that a quantity belongs to the correct product or that a number uses the correct unit. Preserve page locations, evaluate complete documents, and route ambiguous rows to review. The useful operating metric is the rate of consequential errors among automatically accepted records, alongside review effort and coverage. This article develops an acceptance method through a hypothetical invoice workflow rather than claiming a universal model accuracy.
A neatly rendered table can conceal a serious extraction error. A model might recognize every digit on a scanned invoice and still connect a price to the wrong product. The result looks more trustworthy than a garbled OCR transcript because its formatting supplies an appearance of order. An operations engineer has to decide whether that order exists in the source or was imposed during reconstruction.
The relevant question is whether a specific extracted record can enter a downstream system without an unacceptable chance of a costly mistake. That decision differs from choosing the model with the highest published benchmark. It combines evidence from the image, knowledge of the document family, constraints from the destination system, and the cost of asking someone to inspect an exception.
The examples below are hypothetical. Their amounts, error rates, and review times are chosen to expose engineering tradeoffs; they are not measurements of a vendor, customer, or production deployment. Sources were checked for this October 10, 2026 article. The cited research is older work that remains useful for understanding table structure, not an announcement of a new release.
Start with a table whose numbers are all correct
Imagine a supplier invoice containing three lines. The image has a narrow description column, a quantity column, a unit price column, and an amount column. The second description wraps onto another line. A handwritten mark crosses the boundary between the second and third rows.
| Source line | Description | Quantity | Unit price | Amount |
|---|---|---|---|---|
| A | Filter assembly | 2 | 140.00 | 280.00 |
| B | Replacement seal kit | 8 | 12.50 | 100.00 |
| C | Inspection service | 1 | 75.00 | 75.00 |
Suppose an extraction joins the second description to the first row and shifts the descriptions downward. Quantities and amounts remain aligned with one another. The subtotal is still 455.00. Arithmetic validation passes. Character recognition can also appear excellent because the correct strings are present somewhere in the output.
The purchase-order matcher may then accept eight inspection services instead of eight seal kits. The accounting amount is unchanged, but inventory and approval become wrong. A test that compares only totals misses the operational error. A test that compares unordered tokens misses it as well. The acceptance criterion must include the relationship between each description and its numerical cells.
This example reveals three different objects: the characters recognized, the grid inferred, and the record the business intends to create. They require different evidence. PubTables-1M explicitly separates table detection, structure recognition, and functional analysis, and addresses inconsistent table annotation through canonicalization. That distinction is a useful starting point for the evaluator, even when the production document is very different from the scientific articles in the dataset. PubTables-1M paper.
The workflow should therefore retain intermediate representations. If a row fails business matching, an engineer should be able to ask whether the problem began with an unreadable character, an incorrect row boundary, or an unjustified interpretation. A flat spreadsheet containing final values cannot answer that question without reopening and manually reconstructing the source.
Keep the image, the grid, and the business record separate
A practical extraction design stores a page image reference, a detected table region, cell coordinates, raw text, normalized values, and links among those objects. This is a proposed engineering representation, not a requirement of a particular API. Its purpose is to make each transformation inspectable.
Consider the string “1,250”. The raw text preserves the punctuation. A normalization stage decides whether the comma marks thousands or a decimal. A business stage then decides whether the value is a count, a price, or an amount in minor currency units. Collapsing these decisions into one language-model response makes a later correction difficult. Changing the locale assumption should not require rerunning page detection.
Likewise, a merged header should have its own structure. If “Current period” spans two columns, each child column needs a relationship to that parent. Repeating the parent label into every cell may be convenient for CSV export, but it should be an explicit export transformation. Otherwise, an evaluator cannot tell whether the model correctly recovered the merge or merely emitted a plausible sequence of column names.
Amazon Textract documents table output containing cells, merged cells, headers, titles, and related block relationships. Microsoft’s Table Transformer repository notes that text extraction is a separate input for its table inference pipeline. These sources illustrate why structure and text should not be treated as interchangeable accomplishments. Textract tables, Table Transformer repository.
There is an immediate storage tradeoff. Retaining source coordinates and intermediate outputs requires more space than keeping accepted records alone. However, the expensive artifact is usually the document image already needed for review. Small structural records can make that image useful to both software and humans. The design choice should be evaluated against debugging time and the ability to revise normalization rules without recreating the whole extraction.
A correction also needs lineage. When a reviewer moves a value from one row to another, preserve the original association and the accepted association. Training on the accepted output without the original error can improve a model; investigating a recurring failure requires both. A record can be correct today while the process that produced it remains poorly understood.
Build an evaluation set around failure surfaces
Randomly selecting a few attractive PDFs is an easy way to overestimate readiness. Production failures often cluster by supplier, scanner, language, document revision, or page position. A representative evaluation set should expose those surfaces deliberately rather than averaging them away.
Start with the operational population. Suppose a department receives 60% digital PDFs, 25% scanner exports, and 15% phone photographs. Evaluating only scans answers an incomplete question if the intended system handles everything. On the other hand, a scan-only pilot can be legitimate if its intake filter genuinely excludes other inputs. The evaluation population and the product boundary must agree.
Within scanned documents, split by conditions likely to alter structure: skew, faint borders, multi-line descriptions, continuation pages, repeated headers, blank cells, footnotes, and handwritten additions. These are candidate stress conditions derived from the workflow, not a claim that every model fails on every listed condition. Include ordinary documents too. An evaluation made entirely of pathological pages cannot estimate routine review volume.
Hold out document families rather than merely pages. If five pages from the same supplier template appear in training and five in testing, template recognition may masquerade as generalization. A second holdout composed of unseen suppliers asks a different question: whether the extractor handles unfamiliar layouts. Report both results when both are relevant to deployment.
Ground truth deserves its own budget. Two annotators may disagree about whether a wrapped description is a new row or a continuation. Resolve the representation convention before scoring. PubTables-1M’s treatment of oversegmentation demonstrates that annotation ambiguity can change apparent model performance. In an invoice system, the equivalent problem is deciding whether a bundled service is one purchasable item or several narrative sublines. That decision comes from the business task.
Finally, maintain a small challenge collection outside the headline score. It should contain consequential mistakes that previously escaped acceptance. Replaying it after a model or normalization change provides a concrete regression signal. Keep its purpose clear: success on known failures is necessary evidence for a repair, but not an independent estimate of future error frequency.
Score what the destination actually uses
An evaluator can compute several scores without pretending they answer the same question. Text accuracy measures recognition. Structural accuracy measures relationships. A complete-record score asks whether every required field in a row is correct. A complete-document score asks whether all consequential records in a document are acceptable.
The gap can be large. If ten required fields each had an independent 99% correctness probability, the probability that all ten were correct would be approximately 90.4%. Real extraction errors are often correlated, so this multiplication is an illustration rather than an estimator. It shows why a high field score does not automatically imply a similarly high complete-record score.
Weighting also matters. Missing a decorative title and swapping a payment amount should not count as equivalent operational failures. Keep an unweighted diagnostic score so the team can compare technical changes, then evaluate consequential errors using a documented business definition. The latter might include wrong currency, wrong purchase-order line, omitted billable row, or an unsupported sign conversion.
A matching algorithm must handle row order without hiding errors. If the destination needs sequence numbers, moving a row changes the record even when the values are present. If order is irrelevant, compare records by meaningful keys rather than by position. Be cautious with fuzzy matching: an overly permissive description match can forgive exactly the association error that the workflow needs to catch.
For table structure, the GriTS work proposes measures involving cell topology, location, and content. Such research metrics help locate a structural weakness; they do not define the invoice department’s acceptable loss. An engineer can use a structural score during model selection and a separate record acceptance score during release. GriTS paper.
The metric denominator should remain visible. “Ninety-nine percent accepted correctly” means something different if only 10% of documents were accepted automatically. Report automatic acceptance coverage, the error rate within that accepted subset, and the review rate together. Otherwise, the team can improve apparent reliability simply by routing almost everything to humans.
Use disagreement to choose review, not to invent certainty
A second extraction can reveal ambiguity, but agreement is not proof. Two models may share training data, preprocessing, or a preference for familiar layouts. Two prompts sent to the same model can repeat the same mistake with different wording. The usefulness of disagreement depends on the independence of the failure paths.
A good candidate pair might use native PDF text when available and image recognition for a separate check. For a scan without a text layer, the alternatives could involve different cropping or structure methods. Evaluate the pair on known errors. If both consistently attach the wrong description to the same quantity, adding agreement as an acceptance rule increases confidence without increasing correctness.
Treat discrepancies as typed observations. A value disagreement differs from a row-count disagreement. An unchanged amount with a changed sign differs from a small punctuation difference. Routing all disagreements to one undifferentiated review queue makes the reviewer rediscover the problem. Show the cells and the competing interpretations directly.
Confidence scores require similar care. A number supplied by an extraction API may describe a local detection or recognition estimate, not the probability that the final business record is correct. A generated phrase such as “high confidence” is even less suitable as a numerical acceptance threshold. Calibration should compare a candidate score with observed correctness in the intended population.
Suppose scores above a chosen threshold accept 700 of 1,000 evaluation documents, with seven consequential errors. The estimated accepted-subset error rate is 1%. Raising the threshold accepts 500 documents with two errors, an estimated 0.4%. That is a coverage-versus-error tradeoff, not an abstract improvement in intelligence. Whether it is worthwhile depends on the additional 200 reviews and the cost of the five avoided mistakes.
Do not extrapolate the small sample beyond its support. If none of 50 rare-format documents fails, that is encouraging but weak evidence about low error rates. Collecting more examples of the relevant format may be more valuable than optimizing another decimal place on the dominant template.
Arithmetic checks are necessary and deliberately incomplete
Arithmetic provides inexpensive constraints. A row quantity multiplied by unit price should often match its line amount within a defined rounding tolerance. Line amounts should reconcile with a subtotal when the document actually defines that relationship. These checks work best as explicit rules tied to document semantics.
For the sample invoice, 2 multiplied by 140.00 equals 280.00, and the three amounts sum to 455.00. A failed check can identify a transcription or interpretation problem. A passed check establishes only that the selected values satisfy that equation. It cannot prove product identity, currency, tax treatment, or that another page was omitted.
Rounding must be defined at the same level as the source. A supplier may round each line before summing, while another may calculate a total from unrounded intermediate values. For 100 low-value lines, a tiny per-line difference can become a visible subtotal discrepancy. Automatically forcing the extracted total to equal a recomputed total would erase evidence of a genuine source difference.
A useful system distinguishes source inconsistency from extraction inconsistency. If the image itself says quantity three, price ten, and amount twenty-nine, the extractor may have done its job correctly. The destination still needs a decision about the invoice. That belongs in an exception type such as “source arithmetic mismatch,” so reviewers do not spend time correcting accurate transcription.
Currency and scale checks should precede arithmetic acceptance. A table labelled “amounts in thousands” can reconcile perfectly after every value has been interpreted at the wrong scale. Preserve the relevant header and footnote as context attached to the table. When the source provides no explicit currency, do not turn a familiar symbol into an unsupported jurisdiction assumption.
There is also a useful distinction between a hard constraint and a warning. A required quantity field may make a record unusable when missing. An unusual price can be legitimate. Hard rejection should follow the destination contract; statistical surprise should invite investigation. Combining them into one pass/fail result makes it difficult to know which evidence actually justified acceptance.
Calculate the review economics before selecting a threshold
Suppose the department receives 10,000 scanned invoices monthly. Full manual entry takes six minutes per invoice, producing 1,000 hours of work. The proposed pipeline automatically accepts 70% and routes 30% to focused review. If a focused review averages two minutes, the visible review workload becomes 100 hours.
That arithmetic is tempting, but it leaves out exception repair. Assume 1% of automatically accepted invoices contain a consequential error. Seventy errors would escape each month. If finding and repairing each takes 30 minutes, the workload gains another 35 hours. The system’s labor saving is still large in this illustration, yet operational quality may be unacceptable if a mistake triggers inventory loss or an incorrect payment.
A stricter policy accepts 50% with a 0.2% accepted-subset error rate. Review takes approximately 167 hours; ten escaped errors add five repair hours. The stricter policy uses about 37 more hours overall but avoids 60 errors. Dividing that extra labor by the avoided errors gives roughly 37 minutes of additional work per avoided error.
The decision now has a concrete shape. If the average total consequence of an escaped error exceeds the value of 37 minutes of work, the stricter rule may be preferable. The consequence should include downstream interruption and delay, not only the repair action. Some errors can carry much larger losses than the average, so a risk-sensitive workflow may need separate policies for payment amounts and low-impact descriptive fields.
These calculations must use actual review times after the interface exists. A theoretically two-minute review becomes longer if the reviewer has to locate the page, zoom repeatedly, and compare an unlabelled spreadsheet. Conversely, a good crop and a highlighted discrepancy can make review much faster. Model selection and interface design affect the same cost equation.
One additional cost is delayed throughput. If the review queue grows near month-end, invoices can miss processing deadlines despite a low average labor requirement. Evaluate the arrival pattern and reviewer availability. The acceptance threshold is an operating policy embedded in a queue, not merely a static model parameter.
Design a review screen that resolves one question at a time
The reviewer should see the source crop, the proposed row, and the reason it was routed. If the issue is a missing quantity, show that cell and its neighboring labels. If the issue is whether a line continues across pages, show both boundaries. A generic “please verify this document” prompt transfers the whole extraction task back to the human.
Keep raw and normalized values visible where normalization caused the uncertainty. A source value of “(75.00)” and a proposed negative amount should be inspectable together. The reviewer should not need to reverse-engineer whether the model saw parentheses or inferred a credit from the document title.
Correction controls should preserve relationships. Moving a cell to another row should update its association rather than copying its text into a new location and silently leaving the original. Splitting a merged row should require a visible decision about which fields belong to each result. The accepted record should reflect what the reviewer actually decided.
Measure review quality as well as speed. A rushed reviewer can confirm a plausible output without checking the source. A small blinded recheck sample can estimate this risk, provided reviewers know the process and the sample is used to improve the workflow. When review itself is imperfect, the final accepted record is not an infallible gold standard.
Some exceptions should return to intake. A cropped photograph missing the total cannot be repaired through interpretation. Ask for the missing page or route the document to a process that can tolerate the missing information. The interface should support “insufficient source” as a valid outcome so that completion pressure does not turn absent evidence into a guessed value.
Finally, feed error categories back into evaluation. If most review time concerns continuation rows, a better page-boundary treatment may have more value than improving average character accuracy. The review queue is a measurement instrument when its outcomes are structured; it is just labor when its corrections disappear into a final spreadsheet.
A further evaluation question concerns duplicate pages. If intake accidentally includes the same page twice, the extractor can produce two perfectly correct copies of every row. A record-level text comparison may score them all as accurate while the destination doubles the bill. Conversely, a deduplication rule based only on identical descriptions can remove legitimate repeated items. Treat document identity and page identity as separate evidence from table recognition.
For the hypothetical invoice workflow, preserve the original page sequence and an intake identifier before extraction. Compare suspicious repeated regions, but require a policy for repeated line items rather than deleting them automatically. Test one document containing an accidental duplicate page and another containing two legitimate identical charges. The expected outcomes differ even though the visible text is similar. This adds a failure surface that improved character recognition cannot resolve: deciding which source occurrences belong to the received business document.
Expand by document family, with a visible exit condition
A sensible initial scope might include two suppliers, one currency, and invoices with one table per page. That scope should be enforced at intake, not merely written in a project plan. Unknown suppliers and ambiguous layouts can still enter the organization; they should take a path whose assumptions match their documents.
Expansion should answer a specific unanswered question. Adding a third supplier tests unfamiliar layout handling. Adding multi-page tables tests continuation logic. Adding a second currency tests contextual normalization. Combining all three in one release makes a failure harder to attribute and a successful aggregate score harder to interpret.
Retain a way to return an affected document family to manual handling. If a supplier revises its layout, automatic acceptance can pause for that supplier while other families continue. This requires intake classification and policy versioning, but it avoids treating the whole pipeline as either completely enabled or completely disabled.
An operating report can remain small: volumes by family, automatic coverage, consequential errors found in accepted records, review time, and new exception types. Changes in these measures have different meanings. Falling coverage can indicate a safer threshold or a new input mix. Rising errors with stable coverage can indicate a regression. Averages should not conceal a deteriorating minority family.
The exit condition for the pilot is a decision, not a presentation. The department should know whether to expand, revise the acceptance policy, or keep the tool as an assisted-entry system. An extractor that saves substantial reviewer effort without qualifying for unattended acceptance can still be valuable. Its output should be presented with the level of authority that the evidence supports.
The durable engineering result is a chain of explainable decisions from image to cell, cell to record, and record to acceptance. Better models can improve that chain. They cannot replace the need to define what a correct record means for the operation that will use it.
Sources
- PubTables-1M: Towards Comprehensive Table Extraction From Unstructured Documents, Smock, Pesala, and Abraham; dataset and annotation foundations.
- Microsoft Table Transformer, official implementation and evaluation documentation.
- Amazon Textract: Tables, documented table and cell relationships.
- GriTS: Grid Table Similarity Metric for Table Structure Recognition, Smock, Pesala, and Abraham; structural evaluation research.