Beyond Reading: How AI Prepares Document Data

October 3, 2026 · 8 min

Fields and rows in a source document are linked to their corresponding values in a structured record

Every character in a document can be read correctly, yet the wrong data can still end up in the system. The date in the header may be the report date, while the work took place earlier. The address in the contact details may belong to the contractor's office, even though the equipment is at a different site. These distinctions are familiar to staff. Automated processing needs them to be defined.

A discussion about AI should therefore start with the record you need: what information should reach the team that plans maintenance, keeps equipment histories or signs off completed work? InBase's article on unstructured documents describes extracting values and filling fields in a document management system; it also notes the need to configure processing for specific document types. Here are the decisions worth making before that configuration begins.

A usable field combines a value with a clear meaning, a source and a review status.

Two dates — two different questions

Imagine an equipment maintenance report. This is an illustrative example, not an account of an actual implementation. The header gives the report date and the contractor's details. Below it are the service date, the equipment location and a list of tasks. Fields labelled simply 'date' and 'address' would each have several plausible answers. Ask more precise questions: when was the report prepared, when was the equipment serviced, and where exactly is it located?

Distinguishing similar values in an illustrative maintenance report
Source textRecord fieldDo not confuse with
Date in the headerReport dateService date
Date in the service descriptionService dateReport date
Office address in the contact detailsContractor's officeEquipment location
Address beside the site descriptionEquipment locationContractor's office

Google Document AI's documentation describes a field name, description, value type and expected number of occurrences as parts of a document schema. Defining that schema does not, by itself, confirm that every value will be extracted correctly.

For each field in your own schema, add a short rule explaining where to find the value and what to do if there are several candidates. For the service date, this might mean looking in the description of completed tasks. If the report covers several days, decide whether you need a date range or a separate date for each row. Otherwise, the ambiguity simply moves from the document into the system of record.

A row must remain a row

Task names may recur in our report: inspection, cleaning, replacement. A list of those words is of little use if their links to the equipment, quantity and unit of measurement have been lost. An entry reading 'replacement — two' does not explain what was replaced or at which site. Here, the error lies in the relationships between cells that were read correctly.

Azure Document Intelligence describes table analysis results in terms of rows, columns, indices and cell spans. Coordinates and text positions help link a result to its location in the document. These provide a basis for checking the output, rather than a guarantee that every table will be reproduced correctly.

For the report, transfer the task list as a set of rows, keeping the equipment, task, quantity and unit together within each one. If an equipment name appears in a heading above several rows, check which tasks it applies to. A blank cell does not always repeat the value in the previous cell.

Keep the source page and a link to the relevant passage alongside the value. Staff can then check an unclear row directly, without searching the whole report. The link also helps explain a correction: which value changed, and which passage supported the change.

Standardise the format without inventing information

After extraction, values often need to be put into the format the system expects. A date written in words can be converted to the required format, and extra spaces can be removed. That is not a licence to fill gaps: a month without a day is not a complete date, and a quantity with no unit cannot simply be assumed to count individual items.

Google Document AI distinguishes between the original extracted value and a normalised representation for supported fields. Its documentation also describes enrichment with external information as a separate operation. A returned value should therefore not automatically be treated as a literal excerpt from the document: you need to know how it was produced.

In our example, keep the original text alongside the prepared value. If an equipment address has been completed using an internal site register, identify that register as the source of the added information. Do not present the completed address as though it appeared in full in the report.

Agree separately on how to handle abbreviations. A task name can be matched to an approved list, but an unfamiliar abbreviation should be referred for clarification. Once the responsible member of staff has explained it, the team can update the rules. The new variant then becomes a controlled addition to the terminology list, rather than an undisclosed guess by the model.

Found does not mean accepted

The system of record may expect a completed form, but the document may not contain all the fields you need. If the report does not clearly identify the equipment location, making the address mandatory will not produce a reliable address. For a pilot, you could agree on a simple set of statuses:

  • Found. The value has been extracted and linked to the source text; checks against the agreed rules are still pending.
  • Needs clarification. There are several possible values, conflicting information or an incomplete entry; a member of staff needs to resolve it.
  • Not found. The expected value is missing from the extracted data; if the omission matters, check the original document.
  • Confirmed. The value has been accepted under an agreed rule or by the responsible member of staff; the method of confirmation is recorded.

This is a suggested way to organise the work, not a status standard for a particular product. Record normalisation and additions from a reference register separately: they describe how a value was processed, but do not in themselves confirm it.

Data transfer needs an agreed rule too. Can you create a draft record with no service date? Should the entire report wait for clarification? The answers may differ by field. The process owner decides before integration, and the team checks that a missing value is not silently replaced with today's date or a default address.

Test the fields that matter to your work

The pilot sample should include awkward cases as well as tidy reports: multiple dates, a table continuing onto another page, a missing unit of measurement or different addresses. The person who will use the results should mark the correct values and acceptable omissions in advance.

Google Document AI evaluates extraction against labelled test documents and provides metrics for individual fields as well as overall performance. The aggregate result takes account of how often fields occur, so it is not enough to assess a rare but important field on its own.

  • Meaning. The service date has not been replaced with the report date, nor the equipment location with the office address.
  • Relationships. The task, equipment, quantity and unit remain together in the correct row.
  • Source. You can open the relevant passage in the report and distinguish its text from information added from a reference register.
  • Transfer. Missing or conflicting information triggers the agreed procedure instead of silently filling in the field.

Start with one report form: list the fields you need, add examples of ambiguous entries and name the person responsible for resolving each one. Take that schema and a sample of documents to a discussion about a pilot with IQusion. You will have a clear description of the record you need and examples against which to test the results.

Frequently Asked Questions

How should you handle one report covering several pieces of equipment?

Define which equipment each task row relates to. If an equipment name appears above a group of rows, check where that group starts and ends. If the relationship is unclear, ask for clarification rather than applying the first equipment name to the entire report.

Does 'not found' mean the value is absent from the document?

No. The value may simply have been missed during extraction. For an important field, check the original and distinguish an extraction omission from information that is actually absent from the report.

Who should clarify an unclear entry?

Assign this role before the pilot, for example to the person who signs off the work or maintains the equipment history. If the required information is missing from the report, you may need to ask its author. Keep a record of the confirmation and the reason for any correction.

Sources