Hey everyone — I'm Jacky, founder of PonyFlash.
A little backstory: over the past year, my team built CrePal AI, a general-purpose AI video creation agent. Eight months into that build, we hit a wall that every agent developer knows too well — integrating dozens of image, video, and audio models, each with different parameters, different request formats, and different ideal use cases, is an absolute nightmare.
Fast forward to today: personal agents like OpenClaw are becoming a genuine part of how people work and create. But here's the problem — every one of those agents still needs to connect to these complex, fragmented models to do anything useful with media. Someone has to do the dirty work.
So we extracted CrePal's entire model-calling layer, cleaned it up, and packaged it as a standalone skill: PonyFlash. Then we open-sourced it.
Plug PonyFlash into your agent and it instantly gains access to dozens of top-tier voice, image, and video models. Pair it with any business-logic skill to direct it, and your agent becomes a fully automated content creation teammate — no glue code, no per-model SDK wrangling.
We built this because we needed it ourselves. Hope it saves you the pain we went through. Happy to answer any questions below 👇
CrePal