KOURCHAL_
Menu
← All insights

Inside OCRAgent: evidence-backed OCR architecture for document processing

A code-level walkthrough of Ironclad OCR: how PDF packets become classified documents, traceable facts and reviewable decisions in the engineering behind Kaliits.

A PDF can be perfectly readable and still contain the wrong reference. That is the problem behind Ironclad OCR, published in the OCRAgent GitHub repository. The engineering focus is a complete path from a document packet to a decision a person can inspect. Kaliits is the commercial destination for discussing a bounded document workflow; Kourchal explains how the underlying system is built.

This walkthrough examines public commit 647717e, verified on 2 October 2026. It describes implemented code and its limits. The import dossier path is a pilot, and the examples below are fictional. Neither the repository nor these examples establish production accuracy, measured savings or a completed customer deployment.

The architecture in one pass

The dossier path moves through FastAPI intake, PostgreSQL case and job records, a transactional outbox, Redis Streams, a background worker and DossierProcessor. The processor calls a document intelligence provider, obtains the workflow's allowed fields, stores extraction evidence and applies the active ruleset. The Next.js interface reads dossier, review and report APIs. Each stage has a separate job: storing a PDF, recognizing its contents and deciding how to treat a disagreement.

For Kaliits, this separation supports a useful discovery conversation: which documents arrive, which facts matter, which source owns those facts and who may resolve an exception? A model choice alone cannot answer those questions. The workflow must make them explicit.

Start with a case, not an isolated OCR response

DossierRepository.create_case_with_sources stores a case, intake provenance, source artifacts, processing jobs, outbox entries and an ingestion audit event within one database transaction. A source artifact includes its filename, storage location, SHA-256 and page range. A job points back to its case and document artifact. This gives processing a durable identity before a worker starts reading pages.

The distinction between source and derived segment matters when one upload contains several document types. An operator can submit a merged PDF, but the rules still need to compare a declaration with a bill of lading rather than treat every page as one interchangeable document. That is a practical requirement for a Kaliits import-dossier pilot.

Classification must account for every page

DossierProcessor counts the PDF pages and asks the provider to classify supported document types. normalize_segments sorts the proposed ranges and verifies that coverage is continuous from the first page to the last. Gaps, overlaps and invalid ranges cause the packet to be represented as UNKNOWN. Classification below the configured threshold also becomes UNKNOWN.

This is a conservative structural guard. It does not prove that a high-confidence classification is correct. It does ensure that incomplete page coverage cannot quietly remove a page from the workflow. UNKNOWN segments remain stored for review and skip fact extraction, because the processor has no confirmed document schema to apply.

The provider reads; the workflow defines the questions

DocumentIntelligenceProvider is a Python Protocol with classify_pdf, extract_pdf, healthcheck and declared capabilities. The default Paddle provider recognizes PDF pages, derives document segments and extracts facts constrained by allowed_fields. An OpenAI-compatible provider exists behind the same contract. Provider and model identifiers are persisted alongside source results.

WorkflowPlugin defines supported_document_types and schema_for. In the Morocco import workflow, a bill of lading permits fields such as its reference, container number and gross weight; a freight invoice permits freight amount and currency. Restricting extraction to a declared field set keeps the business vocabulary outside the OCR adapter. For Kaliits, changing the document-reading component should not silently change what a business rule means.

Evidence survives extraction and PDF segmentation

ExtractedFact contains a field path, typed value, page, source text, confidence and an optional bounding box. The model validates that pages begin at one, confidence lies between zero and one, and bounding-box coordinates are normalized. EvidenceReference adds the document identity so a discrepancy can refer back to its supporting source.

When the processor extracts a segment beginning on source page four, segment page two is mapped back to source page five. It also rejects an evidence page beyond the segment's end. Without this mapping, a reviewer could be shown the right value on the wrong page. The small arithmetic check protects the connection between the derived fact and the uploaded packet.

A fictional container-number conflict

Imagine a declaration containing container ABCD1234567 and a bill of lading containing ABCD1234568. OCR may read both with high confidence. That confidence measures recognition, not agreement. A configured exact-match rule can produce a blocking finding with both compared values and their evidence. A Kaliits diagnostic would first establish whether these fields should match and which document revisions are in scope.

An unreadable number is a different exception from two readable but conflicting numbers. The first needs better evidence or a correction; the second needs reconciliation. Preserving source text and confidence helps the interface explain that difference rather than reduce both situations to a generic red warning.

Processing, business findings and review are different states

The case lifecycle includes INGESTED, PROCESSING, AWAITING_REVIEW, RESOLVED and FAILED. Decision status includes READY, BLOCKED and REVIEW_REQUIRED. Review resolution includes APPROVED, REJECTED and OVERRIDDEN. A provider timeout is a processing failure or retry condition; it is not proof that the dossier violates a business rule.

At the inspected commit, the active configurable rules evaluator returns BLOCKED when a configured rule finds a blocking discrepancy and otherwise retains REVIEW_REQUIRED. READY exists in the model but is not the normal successful result of this evaluator. That boundary is central to describing the Kaliits pilot honestly: passing the implemented checks does not automatically authorize the operational action.

Where LangGraph fits

The repository also contains an earlier invoice flow built with LangGraph. It fingerprints vendor headers, searches stored schemas, proposes a new schema for human review when a match is absent, extracts with an accepted schema, reconciles and delivers a webhook. The worker routes dossier messages to DossierProcessor and legacy invoice messages to the graph. These are related workflows with different orchestration, not one universal LangGraph dossier agent.

What to evaluate before a real workflow

A useful evaluation set includes unfamiliar layouts, missing pages, weak scans, merged packets, contradictory identifiers and genuine matches. Measure field corrections, evidence quality and reviewer effort separately. A synthetic packet demonstrates a branch of the software; it does not establish representative document accuracy. Kaliits can help scope that evaluation around one document family and a named decision owner.

Inspect the implementation

Dossier processing and page mapping

Typed facts and evidence

Paddle document provider

Legacy LangGraph invoice flow

Continue the architecture series

How the outbox and Redis Streams make processing recoverable

How workflow plugins and rules preserve human authority

Discuss your document workflow with Kaliits