Every OCR tool I tried turned Arabic into soup. So I built SnapDoc (launching Aug 5)
Quick story, and a question for you.
Every OCR and PDF-to-Word tool I tried did unspeakable things to Arabic: letters divorced from each other, whole lines reading backwards, and tables rearranged into abstract art. Charming in a gallery. A disaster on a contract or a government form.
So I built SnapDoc, an Arabic-first OCR that actually respects the language:
- Keeps right to left flow and reconnects the letters, so words stay words
- Holds the tashkeel where it belongs
- Rebuilds RTL tables instead of scrambling them
- Hands you a real editable Word file, not a screenshot pretending to be a document
Bilingual Arabic and English too, which is basically all real paperwork here in the Gulf.
We launch here on August 5. Follow the page if you want a ping the moment it goes live.
But mainly, I want your horror stories: what is the worst thing an OCR tool has ever done to one of your documents? And if you deal with scanned Arabic paperwork, what breaks the most, the tables, the tashkeel, or the reading order?
Cheers,
AISSAM

Replies