AI Document Classification with Human Review

· Updated · 9 min

NIST AI RMF 1.0 helps an organisation define human roles, measure AI performance, and respond to risk, while ISO/IEC 42001:2023 sets requirements for an AI management system (AIMS). Neither document prescribes a single HITL implementation: thresholds, human-review queues, and decision logs are architectural controls selected for the context and risk.

Document classification should separate three layers: risk-governance requirements, technical uncertainty signals, and the operational review process. A model can automate part of the flow, but handing a document to a person does not by itself guarantee correctness. The system also needs defined operator authority, review-quality controls, reproducible records, and rules for returning work to the automated path.

What NIST AI RMF and ISO/IEC 42001 actually cover

NIST AI RMF 1.0 is a voluntary, sector-agnostic framework rather than a mandatory standard or technical specification. GOVERN 3.2 addresses roles and responsibilities in human–AI configurations, MAP 3.5 covers the definition and assessment of human-oversight processes, MEASURE 2.4 covers monitoring system behaviour in production, and MANAGE 4.1 covers response mechanisms including appeal, override, incident response, recovery, and change management.

ISO/IEC 42001:2023 specifies requirements for establishing, maintaining, and continually improving an AIMS at organisational level. It covers policies, roles, risk assessment and treatment, monitoring, and improvement, but it does not define a universal threshold → queue → operator sequence. An organisation selects the appropriate degree of human oversight according to context, error consequences, and its chosen controls.

ISO/IEC 22989 establishes foundational AI terminology and concepts. A specific HITL/HOTL distinction should not be attributed to it without a precise clause. AI-governance practice uses different human-control configurations, from confirming an individual result to monitoring a system with authority to intervene. As of September 2026, ISO/IEC FDIS 42105, which specifically addresses human oversight, is still a final draft under development.

  • GOVERN 3.2 and MAP 3.5: define human roles, authority, and oversight processes in accordance with organisational policies.
  • MEASURE 2.4 and MANAGE 4.1: connect production monitoring with response, override, incident response, and change management.
  • ISO/IEC 42001:2023: specifies requirements for an organisation's AIMS, not for one classifier pipeline.
  • ISO/IEC 22989 and FDIS 42105: the former is a terminology foundation; the latter develops separate human-oversight guidance and has not yet been published as an International Standard.

Designing a risk-based human-review workflow

A classifier may return a confidence score, margin, or predicted class probability. This number must not automatically be interpreted as the probability of correctness. For example, a value of 0.9 corresponds to roughly 90% correct decisions only when calibration has been demonstrated on a separate, representative dataset for the intended context.

Set a threshold using observed precision, recall, per-class errors, false positive and false negative costs, reversibility of the downstream action, and operator capacity. A single global threshold is often less useful than class- or action-specific thresholds. Revalidate it after changes to the model, document templates, OCR pipeline, or production population.

A low score is only one signal. High-confidence errors can occur under distribution shift or unfamiliar inputs. The workflow should also assess image and OCR quality, out-of-distribution signals, mandatory-field presence, business-rule conflicts, class sensitivity, and the risk of the next automated action. OCR/image-quality scores and classification scores are not interchangeable.

When a decision should be handed to a person

A financial document may be routed to an accountant because of a low class score, poor scan quality, or a missing mandatory field. For a medical document, checking whether narrative text supports a particular code is a separate semantic-validation task rather than simple document-type classification. Depending on the scenario, the system may pause a business action, apply a fallback rule or second model, request more data, perform sampled review, or retain the result for post-audit.

Thresholds, routing, and a reproducible decision log

Risk is determined not by a file's name but by its use: who may be harmed by an error, which action follows automatically, whether it can be reversed, and how quickly an error can be detected. The matrix below is therefore illustrative and is not a NIST or ISO classification.

Illustrative example of risk-based routing
Use case Risk signals Possible response
A payment document triggers a financial transaction Low score, inconsistent details, a large amount, or a new template Pause processing and route to an authorised accountant
A medical document affects coding or routing Conflicting fields, an unfamiliar format, or critical downstream impact Route to a domain specialist; require a second review for critical decisions
Internal correspondence is indexed for search Low score without an irreversible automated action Classify automatically with random sampling or post-audit

The NIST AI RMF Playbook suggests maintaining histories, audit logs, override statistics, and adjudication results, but these suggestions are voluntary and contextual. For practical reproducibility, record the model and OCR-pipeline identifiers and versions, policy and threshold versions, predicted class and top-k scores, routing rule, operator decision and rationale, time, role or user, second-review result, and correlation/request ID. Define access, retention, and personal-data minimisation separately.

UnityBase provides an audit trail for entities with the audit mixin enabled; its documentation lists the user, action time, old and new values, and request ID. AI-specific fields—model version, score, threshold, and correction rationale—must be designed explicitly in the data model and application logic. Scriptum.DMS provides AI-assisted document recognition and classification, Scriptum provides a Camunda-based process layer, and Megapolis.DocNet provides BPMN routing and audit trails. Combining them can support a HITL solution, but it does not create a complete AI audit automatically.

Drift monitoring and controlled feedback

Changes in document formats, sources, or scan quality can shift the input distribution—causing data drift—and may reduce model performance, but that relationship must be confirmed against current ground truth. Monitor shifts in input and prediction distributions, per-class precision and recall, calibration, review rate, queue time, override rate, and observed error frequency separately. NIST places production monitoring under MEASURE 2.4.

Verified operator corrections can become candidates for evaluation, error analysis, rule or prompt changes, model retraining, or replacement. They must not become training data automatically. A low-confidence queue is a biased sample, so it does not replace a separate representative validation/test set and a random or stratified sample of production documents.

Before training, reconcile conflicting labels through second review or adjudication, version the data, and test the new model outside production. Deploy changes under change control with approval, staged rollout, monitoring, and rollback. Sometimes the right response to feedback is not retraining but changing a rule, route, OCR pipeline, or the task definition itself.

Technical HITL checklist supporting an AIMS

This checklist covers only the technical HITL workflow and does not demonstrate conformity with ISO/IEC 42001. It helps gather evidence for a broader AI management system that also covers policies, leadership, roles, competence, risks, suppliers, internal audit, and continual improvement.

  • Context and authority: document error consequences, operator roles, override authority, escalation, and cases requiring mandatory confirmation.
  • Threshold validation: evaluate calibration, precision, recall, and class-specific errors on representative data; document threshold choice and the error/coverage trade-off.
  • Routing signals: consider not only classifier score but also scan and OCR quality, OOD signals, business rules, sensitive classes, and random sampling.
  • Operational capacity: monitor queue size and age, the target review time, workload, operator training, and safe behaviour when the human-review path is unavailable.
  • Reproducibility: log model and policy versions, scores, the triggered control, decision, rationale, user, time, and linkage to the downstream action.
  • Human-decision quality: use gold-standard examples, sampled audit, second review, or adjudication in proportion to risk.
  • Production monitoring: compare metrics with a baseline, track drift, errors, overrides, and incidents, and define rules for raising or lowering automation.
  • Change management: separate review data from the approved training set, version datasets and models, test changes, and support rollback.

Frequently Asked Questions

How can a threshold be configured without overloading operators?

Set the threshold on a representative validation set, accounting for calibration, per-class errors, false-positive and false-negative costs, and queue capacity. Different classes and downstream actions may need different thresholds. Production evidence can justify either raising or lowering them; do not rely on the score's absolute value alone.

Is an audit trail sufficient for ISO/IEC 42001 conformity?

No. A log supports decision reproducibility and can provide technical evidence for an AIMS, but it does not by itself demonstrate conformity with ISO/IEC 42001. The assessment covers the organisation's entire management system: policies, roles, risk assessment and treatment, competence, monitoring, internal audit, and continual improvement.

How should operator errors during manual review be controlled?

An audit trail makes a decision reproducible but does not automatically establish that it was wrong. Critical classes may require a second expert review, adjudication, gold-standard examples, inter-rater agreement measurement, and sampled audits. Only confirmed corrections should enter a training dataset.

Sources