Ten thousand PDFs in. One clean table out.
RCA turns large quantities of PDFs and paper documents into clean, catalogued tables. Extract the key values, report on them, and find the patterns that were stuck in filing cabinets. More capable than standard digitisation software, at a fraction of the price.
Both apps run on your own hardware and work offline. Every extracted value carries a confidence score you can check.

Specimen AUS-TRN-0001, generated by the RCA Synthetic Document Generator
What we build
The catalogue · index of holdingsRCA Document Library
Point it at your PDFs and scans. It reads them, extracts the fields you care about, scores its own confidence on every value, and builds a clean table you can search and export.
- 96.1% field accuracy on OCR-only extraction (V2 benchmark, 76 fields)
- Offline-first: OCR, extraction, search and export without internet
- Confidence built from four visible factors, not a black box
- Optional AI validation layer that never overwrites OCR results
- Costs a fraction of a standard digitisation contract
RCA Document Generator
Realistic synthetic Australian documents for training and evaluating AI, with complete ground truth on every field. No real patient anywhere in the pipeline.
- 45 medical document types across 81 clinical case archetypes
- Deterministic: same seed, byte-identical PDFs and ground truth
- Insurance packs, including red-flag packs with planted inconsistencies
- Fully offline: about 5,000 documents in 5 minutes, no LLM calls
- Valid Medicare and provider formats, NSW addresses, real structure







Numbers from our published benchmarks, misses included.
We publish our misses.
Evidence · live trial 05 Aug 2026Most vendors quote a single accuracy number. We publish the full run log: every field on every document, including the ones the models got wrong. When a model fails a document, it is scored zero and recorded in the results. In this run, two dense flow sheets exceeded the output limit on Haiku 4.5. Both are counted as failures in the table.
Run rca-library-trial-20-w3 · 05 Aug 2026, Sydney time. Opus 5 pass added 21 Aug 2026, same documents and ground truth.
Total API spend across the three model runs: A$4.42
| Field | Sonnet 5 | Haiku 4.5 | Opus 5 | Result |
|---|---|---|---|---|
| Entity | 20/20 | 18/20 | 20/20 | strong |
| Amount | 20/20 | 20/20 | 20/20 | clean |
| Document date | 19/20 | 18/20 | 20/20 | strong |
| Reference number | 19/20 | 16/20 | 16/20 | needs work |
| Document type | 14/20 | 15/20 | 14/20 | needs work |
| Overall field checks | 92/100 | 87/100 | 90/100 | honest |
I started RCA because I had this problem myself and could not find a good solution on the market. Most businesses sit on thousands of PDFs and paper documents they cannot use. RCA extracts the key values, builds a clean table for cataloguing them, and gives you metrics and patterns that were locked in paper. It does what the standard digitisation vendors do, with more capability, at a fraction of the price.
Sydney, NSW
See the quality before you spend a dollar.
Request a free preview pack with real sample documents so your team can evaluate hands on.