KOURCHAL_
Menu
← All insights

OCR workflow plugins: deterministic rules and human review in OCRAgent

How Ironclad separates document-reading providers from business workflows, and why its rule engine and review controls matter more than an impressive extraction demo.

Two companies can submit similar PDFs and need different decisions. One cares about container references; another compares accepted deliveries with supplier invoices. Hardcoding both policies into an OCR prompt makes the reader and the business policy difficult to change independently. Ironclad OCR separates those concerns through provider, workflow and repository contracts. This is a practical architectural choice behind the Kaliits document-workflow approach.

This article describes public OCRAgent commit 647717e, verified on 2 October 2026. It is a technical implementation note, with fictional examples and explicit pilot boundaries. Kourchal documents the mechanism; Kaliits is the commercial place to discuss the actual process and decision owner.

Three contracts with different jobs

DocumentIntelligenceProvider declares classification, extraction, health checking and capabilities. WorkflowPlugin declares a workflow identity and version, supported document types, allowed schemas, reconciliation and report construction. Repository implementations store jobs, artifacts, facts, decisions and review history. The processor receives these dependencies rather than defining every infrastructure call and business rule itself.

This resembles a ports-and-adapters design: the core processing code talks to declared contracts while concrete adapters handle document intelligence and persistence. It is a useful boundary, not a claim that every module is fully independent or that a new workflow needs no engineering. A new adapter still has to satisfy the evidence model, error contract and workflow's real inputs.

Document schemas are part of the workflow

MoroccoImportDossierWorkflow exposes schema_for for each supported type. A declaration may carry a declaration number, container reference, weights and value. A bill of lading may carry a shipment reference, vessel, ports, packages and gross weight. A freight invoice has a narrower set including freight amount and currency. UNKNOWN has no extractable field schema.

DossierProcessor passes those fields to the provider as allowed_fields. This keeps the domain vocabulary in one visible place. For Kaliits, it also creates a precise pilot brief: which fields are required for this decision, from which document families, and with what supporting evidence? Extracting every possible string is not the same as implementing that brief.

Changing the reader does not change the authority

The default Paddle adapter recognizes pages and builds structured facts; the optional OpenAI-compatible adapter implements the same provider interface. A provider's capabilities declare matters such as PDF input and bounding-box support. Provider identity and model identity are stored with processed results, making the reading component visible in the case record.

The interface does not make providers equivalent. Their classification behavior, evidence support and failure modes require evaluation. It makes the integration boundary explicit so Kaliits can assess a document reader without silently rewriting business policy. Whether a shipment may proceed is still a rule and review question.

Rules are versioned policy, not prompt instructions

The processor loads an active ruleset for the client and workflow. When one exists, evaluate_ruleset generates the decision. When none exists, it invokes the workflow's reconciliation but forces REVIEW_REQUIRED with an unconfirmed ruleset marker. A draft or missing policy therefore cannot act as authority for automatic clearance.

The inspected evaluator implements REQUIRED_DOCUMENT, FIELD_REQUIRED, EXACT_MATCH and AUTHORITY_PRECEDENCE. A rule has an identity, active flag, severity and failure outcome. Missing required evidence or a configured disagreement becomes a structured discrepancy. A failure outcome of BLOCKED sets the blocking result; other supported failures require review.

Unknown templates and numeric or date templates without an implemented configured evaluator remain review-only in this public version. They should not be described as completed tolerance or chronology controls. This distinction prevents a Kaliits pilot proposal from promising a check merely because a rule name can be stored.

An authority rule needs an explicit source

AUTHORITY_PRECEDENCE requires an authoritative document type in its parameters. Without that, the evaluator records a warning and requires review. With an authority configured, it looks for one unambiguous normalized value from that document type. This is a bounded policy: identify the source and check its evidence before treating a value as authoritative.

For example, a fictional packet might contain two weight values because one is gross weight and one is net weight. Declaring one document authoritative does not magically make unlike fields comparable. Kaliits must establish the actual field semantics and units during discovery. Normalization is a software operation; authority is a business decision.

Passing rules still leads to review in this commit

The model includes READY, BLOCKED and REVIEW_REQUIRED. However, this rules evaluator deliberately returns REVIEW_REQUIRED whenever no blocking rule fails. It does not automatically produce READY for a matching packet. The workflow's fallback also disables automatic readiness by default.

That behavior is a control boundary rather than a performance claim. It means the code can prepare a structured finding while leaving final authorization with a person. A later automatic-clearance policy would require confirmed rules, representative tests and a deliberate change. The current Kaliits-facing explanation should describe the review preparation that exists today.

Review protects against duplicate and stale actions

DossierRepository.resolve_review accepts the expected decision version and an idempotency key. It first checks for an existing review with that key, then locks the latest decision and compares versions. A stale action raises a decision-version conflict. Repeating the same accepted request returns its existing result rather than resolving it twice.

APPROVE is restricted to a READY decision. A BLOCKED or REVIEW_REQUIRED result requires rejection or the explicit override route, with its reason handled by the review API. Resolution updates the review and case together with an audit event. Consequently, a matching packet in the current review-only evaluator is not simply approved through a hidden automatic path.

Imagine a reviewer opens decision version three, then a corrected document produces version four. Submitting the older action should fail its version check. That protects a Kaliits operator from resolving a finding they have not actually reviewed. It is a concrete benefit of versioned state, not a vague promise that a human is somewhere in the loop.

Reports use the same decision model

WorkflowPlugin.build_report receives the case context and decision. The import workflow includes workflow and ruleset versions, case identity, external reference, documents and the structured decision. The repository stores that report payload with the decision result. This avoids constructing an unrelated report narrative that can disagree with the reviewable findings.

A useful report answers which field differed, which document supported each value, which rule applied and which decision version was reviewed. Kaliits can use those questions when defining a handoff. A PDF export alone does not establish a complete compliance record or prove every required audit attribute is captured.

The legacy invoice plugin is a separate extension point

ClientPlugin belongs to the earlier invoice architecture. It normalizes vendor names, enriches schemas, reconciles extracted data and builds delivery payloads. SupplyChainPlugin reads purchase-order lines and goods receipts to perform invoice reconciliation. WorkflowPlugin belongs to the dossier path. Sharing an extensibility principle does not make their interfaces interchangeable.

The legacy graph can propose an unfamiliar vendor schema and hold it for a person before reuse. That is distinct from training an OCR model or automatically learning a new business policy. Kaliits should describe schema discovery, extraction and rule changes as separate activities, each with its own acceptance criteria.

A practical extension plan

For a new workflow, define the document families and canonical fields, implement a workflow contract, register it, test missing and conflicting evidence, and verify provider output against real samples. Confirm the rule owner and review actions before enabling a policy. Existing ERP matching or native imports may already handle parts of the process; a custom OCR path should fill a measured gap.

The repository is publicly inspectable and carries a PolyForm Noncommercial license plus a separate commercial-license document. Consult those repository files for permitted use. Public code availability is not a claim of unrestricted commercial reuse. A Kaliits conversation can scope delivery, integration and licensing alongside the workflow.

Inspect the extension points and controls

Provider contract

Workflow and legacy client contracts

Import document schemas and reporting

Deterministic rules evaluator

Versioned review persistence

Repository licensing

Continue reading or discuss an implementation

Evidence-backed OCR architecture

Outbox, retries and Redis Streams

Discuss a bounded workflow with Kaliits