Local-first document reader · macOS
Open a confusing government PDF and get back what it obligates you to do — the deadlines, the fields, the binding numbers — with every value quoted from the source and every uncertain reading marked uncertain. The default path never touches the network.
The pipeline
Five steps, and every one of them runs on your machine. The model only ever sees what the deterministic steps below it already found — which is why it cannot invent a deadline that isn't on the page.
The app
A conservation application, opened and read. The document stays on the left; what was read from it stays on the right, next to the page it came from.
network calls: 0, and it is not decoration.Why it isn't a summariser
The documents people least want to upload are the ones they most need read. In the local tier there is one exit from the system, and it is a function that cannot run without a grant you gave for this specific document. When a region is too unclear to read locally, you are shown the exact crop before it goes anywhere.
one egress function · asserted by test
On a machine without a local model, the app falls back to an endpoint you configure — and it says so at import, on the screen where the document lands. Not in settings, not in a tooltip. A fallback you have to discover is the same as a lie, and it makes every other claim on this page worthless.
disclosed at import · before anything is read
The ring is not the model rating itself. It is a composite: format validators first, OCR's low-confidence signal second, homoglyph disagreement third — and the model's own opinion last, as a tiebreaker that can never be the only input. A finding that fails its validator can never render green.
≥1 non-model signal required · breakdown in the inspector
Who it's for
The shared shape: a binding document, a real consequence for misreading one line, and a rule — legal, contractual, or just sensible — against uploading it.
Restricted-entry intervals, PPE requirements, application rates. The obligations are binding, they are buried on page 34, and a wrong number is a violation. Every value comes back with the verbatim line it was read from, so an auditor can check your reading against the label instead of trusting a summary.
EPA labels · SDS · restricted-use conditionsYou have a stack of control narratives, DPAs and audit requests you are not allowed to paste into a chatbot. Local tier makes zero network calls, writes nothing to disk, and keeps no state between sessions — so there is no vendor to add to the subprocessor list and nothing to answer for in the next review.
control evidence · DPAs · vendor questionnairesA born-digital government form carries its own fillable fields. PigeonEye reads them straight out of the file — Schedule F is 89 widgets across 84 fields, exact, in about 40 ms. Nothing is inferred, so nothing can be wrong.
IRS Schedule F · 4835 · state equivalentsConservation, subsidy and eligibility paperwork runs to a hundred fields and cross-references three other forms. Get the field list, the page each one sits on, and the handful of values already printed on the page — before anyone starts filling.
NRCS CPA-1200 · FSA · state cost-sharePrivileged material cannot go to a third-party API, and the document you most need read is exactly the one you least want to upload. When a region is too degraded to read locally, you are shown the precise crop and asked. Declining still produces a result.
client files · discovery · noticesThe original user. A 45-page pesticide label, an IRS farm form and a conservation application in the same week, each one able to cost real money if a single line is missed. Plain-language explanation, the deadlines that actually bind, and a checklist — never advice.
the primary corpus this was tuned onMeasured, not claimed
Everything here came out of running the thing over real EPA labels, IRS forms and deliberately degraded scans. Where it is weak, it is weak in public.
| What was measured | Result |
|---|---|
| Network calls in the local tier | 0 — no cache, no queue, no server, nothing to GDPR |
| Bytes of your document written to disk | 0 — the source file is opened read-only and never copied |
| Form fields read straight from the file | 63 / 89 / 105 — IRS 4835 · Schedule F · NRCS CPA-1200 — exact, no inference |
| 45-page EPA label, render + OCR, end to end | 23.3 s — six pages in flight, debug build |
| OCR lines scored to build the confidence rule | 1 092 — across 18 real degraded scans — measured, not assumed |
| Character error rate on EPA prose pages | 1.5 – 8 % — the number that bears weight, and it is good |
| Numeric recall on the flagship 45-page label | 73.8 % — published because hiding it is the actual failure — see below |
| Third-party dependencies in the app | 0 — no runtime to download, no model to fetch, ~20 MB |
About that 73.8%. On the woody-brush rate table of a 45-page label, OCR returns the plant names and drops the rates — the only number that survives one of those pages is the page number. Raising the resolution does not fix it; the loss is not even monotonic in DPI. It is published here because a reader who is told “the rate table on page 34 did not come through” goes and looks at page 34, and a reader who is shown a clean summary does not. Absence has no confidence — so it gets named instead.
One window, one Open… button, and no account. It reads the file where the file already is.
v1.0.0 · macOS 26 · Apple Silicon · ~20 MB · no account · nothing to sign up for