space ocr update #3 — I need 5 real workflows

by

Last week I ran space ocr against Mistral's Document AI, and the number that stayed with me afterwards was not the one in the headline.

The benchmark. Same document photos, same fields, same scorer, three runs each. Across 463 fields, space ocr came out at 91.1% field accuracy and Mistral at 70.0%, and the repeated runs were considerably more stable on my side. Full write-up with the individual cases and where the differences actually showed up: and

The problem with it. I chose every document in that set. I found them, looked at them and understood them before the test ran. I can keep adding receipts, invoices and awkward layouts, but it stays a set I picked — which is the one thing a benchmark cannot fix by growing. The next useful test is one I do not control.

What I am asking for. Five teams that already process real documents as part of their work. Invoices, delivery slips, receipts, statements, forms, anything where you extract fields and then do something with the result afterwards. One real workflow per team, roughly 20 to 100 documents through space ocr. Redacted is completely fine. I set the fields up with you, and we compare the output against your real values together.

The number I actually want. When the extraction is wrong, does space ocr notice? Every value it returns is checked against separately detected characters and their positions on the original image. If the evidence lines up the value comes back verified, and if something does not line up it is flagged for review. So for each workflow I want three counts: how many real errors happened, how many of those were flagged, and how many wrong values still came back verified.

The third count is the one that decides whether any of this was worth building. If $18,700 comes back as $13,700, the mistake is not the interesting part — OCR makes mistakes. The interesting part is that the wrong number arrives in perfectly normal JSON and everything downstream treats it as a correct one. If someone on your team currently checks every extracted value because of exactly that risk, I also want to see whether that review work actually shrinks in practice rather than in a benchmark.

Who I am hoping to hear from:

  • Teams where extraction is already part of a real process and somebody still does not fully trust the output.

  • Teams that have had a wrong value slip through quietly and cause trouble later. That case is the whole test.

  • Not teams whose pipeline works fine and nobody checks anything afterwards. That is a good place to be, and you probably do not need this.

Comment here or email me at . It helps if you say what your team does, what kind of documents you process, roughly how many, what you use today, and which kind of mistake causes the biggest problem when it gets through.

I do not know whether space ocr holds up on documents I did not pick. That is the entire reason I am asking.

26 views

Add a comment

Replies

Be the first to comment