How Swift-Slide lifts text off a slide (OCR → inpainting → PPTX) — ask me anything

by•

I'm Tsuyoshi, the solo dev behind Swift-Slide. We launch on Product Hunt on October 6, and I'd rather answer the technical questions here than bury them in launch-day comments.

The short version of the pipeline, per page:

1. OCR — Azure Document Intelligence pulls text with word-level coordinates

2. Classification — Vertex AI Gemini decides which text is safe to remove (text baked into illustrations is deliberately left alone)

3. Background restoration — self-hosted inpainting on RunPod fills in where the text was

4. Style estimation — colour, size and position are inferred in our own backend, then rebuilt as editable text boxes in PPTX

Pages run in parallel on Cloud Run, so a 111-page deck takes about 4 minutes rather than 40.

Things I'm happy to get into: why I stopped using an LLM for step 4, how the "don't remove this text" decision works, what still breaks, and what it cost to go from a Gemini-only prototype to this.

Things I won't share: exact per-page costs.

If you've tried the no-signup demo () and something looked wrong, post a screenshot here — that's the most useful feedback I can get before launch.

3 views

Add a comment

Replies

Be the first to comment