What makes a PDF difficult to convert into usable Markdown?

by

While building , I learned that extracting the words is usually the easy part.

The difficult cases are multi-column reading order, dense tables, formulas, images with captions, scanned pages, and documents that mix different languages. A conversion can look fine at first glance while quietly changing the meaning of a table or putting paragraphs in the wrong order.

I’m curious about the PDFs that cause problems in your own workflow.

Is it research papers, financial reports, scanned books, legal documents, manuals, or something else?

And what matters most in the output: searchable text, clean tables, formulas, images, or reading order?

No need to share anything private. Even describing the type of document would be useful.

6 views

Add a comment

Replies

Be the first to comment