Launched this week

PDF to JSON API
Get structured JSON with extracted content for any PDF
18 followers
Get structured JSON with extracted content for any PDF
18 followers
Submit any PDF, get structured JSON back. Thumbnails and rendered PDFs hosted on our CDN. Automatic document classification: invoices get line items and totals, contracts get parties and clauses, resumes get experience and skills. Full text as searchable chunks with page numbers, ready for RAG. Built-in OCR for scanned documents. Same API handles DOCX, PPTX, images, audio, video, and web pages. Free plan available.




Fabric
Hey Product Hunt 👋
I'm Johnny, founder of the team behind the PDF-to-JSON API.
We built the PDF to JSON API because extracting useful data from PDFs still requires way too much infrastructure. You end up installing Ghostscript or LibreOffice, fighting with scanned documents, writing chunking logic, and maintaining a processing pipeline that breaks every time someone uploads a weird file.
So we made it one API call. Submit a PDF, get structured JSON back.
You get text split into searchable chunks with page numbers (ready for RAG), a thumbnail and rendered PDF hosted on our CDN, and all the document metadata. The API also detects whether it's an invoice, contract, or resume and extracts the structured fields automatically: line items and totals, parties and clauses, experience and skills. No templates, no training.
The same endpoint handles DOCX, PPTX, EPUB, images, audio, video, and web pages. You learn one API and use it for everything. Add ?format=markdown if you want LLM-ready output instead of JSON.
PDF to JSON is part of the broader Exabase platform, so the same API key also gets you deep search, memory, and automation if you ever need to go further. But it works perfectly well on its own.
Free plan available, no credit card required.
Thanks for checking it out. I'll be here all day answering questions!
Johnny