What should AI extract from procedures, work instructions, manuals, and papers?
I’m building Astra Annotator, initially for procedures (SOPs), work instructions, manuals, and research papers. I plan to broaden the focus to Office document workflows later.
Early-release note: The app has only recently become available. Bugs and rough edges remain, and AI may miss, misread or mislabel content. Please treat results as draft annotations, check them against the original document, and report problems through GitHub Issues. The Product Hunt launch is scheduled for September 18.
The idea is simple: describe what matters, watch AI place editable labels and notes on the page, then extract the useful parts. You can also mark regions yourself or supply exact label names.
Possible ways to use it
Procedures / SOPs: Mark prerequisites, numbered steps, decision points and warnings. Extract relevant passages to help a reviewer prepare a checklist or training handout, while retaining the source page for verification.
Work instructions: Highlight required tools, materials, settings, dimensions and inspection criteria. Label the corresponding diagram regions and add clarification notes, so a team can assemble a focused reference for a particular task.
Manuals: Pull out the installation, maintenance or troubleshooting sections you need. Pair relevant diagrams with nearby instructions and specifications instead of repeatedly searching a long document.
Research papers: Separate titles, body text, figures, tables, equations and captions. Extract table content, crop figures, and review equations as LaTeX for literature notes, comparison sheets or presentations.
These are workflows to explore with the annotation and extraction tools, not guarantees of complete or correct results. Human review is especially important for safety warnings, operating parameters and inspection criteria; the original approved instructions remain the reference.
What comes next: After improving these document workflows, I plan to broaden the focus to Word, Excel and PowerPoint use cases. Basic Office import is already available, but browser previews focus on extracted text and do not preserve every original layout or pagination detail.
The recording below shows a real GPT-6 Astra run on three pages of an OpenAI paper through local Codex App Server. It plays at 4× speed.

Try the desktop demo: Astra Annotator on GitHub Pages. Click “Explore a paper” to inspect saved results without an API key. Fresh AI runs use your own API connection or the local server; API usage is billed by your provider.
Which document do you work with most, and which of these workflows would help you? What output would save you time—an annotated PDF, image crops, CSV, or LaTeX? I’d also appreciate reports of confusing behavior or incorrect annotations. Please describe the steps to reproduce a problem and avoid sharing confidential documents or API keys.

Replies