BeatScope measures a song — beats, transients with band energies, tempo changes, bars, cues, structure — and exports a portable timing package: deterministic runtime, self-check probe, hash-verified files. No audio, no visual style, no task statement, so the picture stays the consumer's decision. Beathi, the same app, turns those measurements into a music video on your machine. Same seed, same film, every time. Windows build, pip wheel, MIT.
What real task does your product handle with GPT-6 Astra?
Maker
The task: turn a song into a music video whose cuts actually land on the music.
Beathi measures the song locally — exact beat times, transients with band energies, tempo changes, structure, and an ordering of which moments are worth reacting to — then exports it as a portable timing package that deliberately ships no visual style.
The model does the part a model should: read the measured facts, ask the human what media and mood they want, author the visual system itself, place every cut and accent on the measured timestamps, and verify its own work with the package's self-check probe. Our MCP surfaces let it query facts at any instant instead of guessing.
In the showcase, an agent received only the package, a song and a reference image, and built "Dirt / Cold / Rings" — three visual worlds, one piece of music, every cut on a measured beat. None of the look came from us.
Report
Maker
📌
I make music tools, and I kept hitting the same wall: every "audio-reactive" visual starts by guessing at the music — a loudness threshold here, a fake beat grid there — so the result is never quite right, and never reproducible.
So BeatScope went the other way: measure first, then publish the measurements as data. Exact beat times, real transients with band energies, the structure, and an ordering value for choosing which moments are worth reacting to when your budget is limited. What you get is facts plus a small deterministic runtime — no styling, no "here's the look", because a package that pre-decides the visual stops the agent from asking the person what they actually want.
The part I'm proudest of: the showcase film wasn't made by us. We handed the exported package to a coding agent together with a reference image, and it built the visuals itself — the cuts land on our measured timestamps, the look is entirely its own.
Honest limits: downbeat estimation is the weakest part of the analyser (there's a public benchmark and it shows), and the built-in template is deliberately one look, not a style library. Feedback welcome — especially from anyone who hands a song to an agent to see what it makes.