Clipto MCP - Let agents source clips from terabytes of your local video

Clipto MCP gives Claude, ChatGPT, and other AI agents the ability to source clips and more from inside the videos, photos, and audio recordings stored on your computer. Instead of manually browsing files, simply describe what you need. For example, turn a script into a video by matching each sentence with your local footage; find every scene where someone mentioned a topic; create rough cuts; or search years of media as if you have a dedicated assistant editor.

Add a comment

Replies

Best

Hi Product Hunt! 👋 Henry here, a few months ago, I launched , a fully local AI media search engine that helps people search terabytes of videos, photos, audio, meetings, and documents using natural language.

Since then, one request kept coming up:

“Can Claude use Clipto?”
“Can ChatGPT search my media library?”
“Can my agent actually edit videos using my own footage?”

Today, we’re excited to launch .

AI agents can already access your files. What they can’t do is understand what’s inside terabytes of media.

That’s what Clipto MCP changes. It gives AI agents semantic understanding of your local media, so instead of manually browsing folders or scrubbing through timelines, you can simply describe what you want.

Some examples:

🎬 Turn a script into a video by matching every sentence with relevant footage.

🔍 Find every clip where someone mentioned a specific topic.

🎤 Search years of meetings for a decision or discussion.

✂️ Generate rough cuts from thousands of hours of video.

Everything runs on the media already stored on your computer. No uploading your library to the cloud.

This is only the beginning. As AI agents become more capable, they’ll need more than file access. They’ll need to understand the content inside our personal media.

To celebrate our launch, we're offering 1 month free to anyone who signs up this week with code PHLNCH.

We’d love to hear what you’d build with Clipto MCP.

We’ll be here all day answering questions and collecting feedback.

Thanks for giving it a try! 🚀

One more experiment we wanted to share because this one surprised us.

We downloaded a bunch of Elon Musk interviews, indexed them in Clipto, connected the library to Claude through Clipto MCP, and basically asked:

“What would be something fun to make with all this footage?”

Claude came up with the idea of turning Elon into a music video using Daft Punk’s Around the World.

From there, the whole thing ran automatically.

Clipto had already analyzed the footage in detail, turning every interview into structured, searchable media with transcripts, speakers, scenes, and precise timestamps. Through Clipto MCP, the agent could understand and retrieve exact moments across the entire library.

Claude analyzed the song, decided where Elon’s words could fit, asked Clipto for the matching moments, picked from the results, and assembled the final video with FFmpeg.

We didn’t manually select the clips or clean up the result. What you see here is the first output:

Raw output:

Want to try it yourself? We packaged up the same Elon interview clips we used, so you can download them and run your own experiment.

Elon Interview Library — the same footage we used for this experiment. Try a different prompt and see what your agent comes up with.

We also put together a B-roll library if you want to try something completely different.

B-roll Starter Library — 1,000 royalty-free clips across people, work, cities, travel, nature, and more. Use them to experiment with your own ideas and workflows.

Or connect Clipto MCP to your own media and see what your agent can make from the footage you already have.

We’d love to see what your agents come up with.

 To make it easier for everyone to try, here’s the prompt we used for the experiment.

Feel free to copy it, remix it, or use it with your own footage. Have fun—and we’d love to see what your agent comes up with! The following is the original text of the prompt:

clipto-editing # Elon Musk × Around the World — Lyric Supercut Prompt Using my authorized media package in Clipto MCP and local video tools, make a 60-second Daft Punk "Around the World" × Elon Musk lyric-replacement supercut, delivered as a directly playable local MP4. Work fully autonomously — don't ask me to approve individual shots.

The Effect I Want

The song's original music video plays continuously as the base. Every time the song sings "around the world", hard-cut to a clip of Elon Musk saying a complete "around the world", with his "around" landing exactly on the real vocal onset of "around" in the song — so it sounds like Musk is singing the lyric in place of the original vocal. The shot count is determined by how many verified "around the world" occurrences actually exist in the chosen 60-second window — cut on every one of them, spread across the full minute with continuous momentum, never feeling like scattered interview inserts. Reference specs: 1280×720, 25fps, 60 seconds, H.264+AAC, ~-14 LUFS. Overlay a title once at the start (ELON MUSK — AROUND THE WORLD in bold white, with a smaller subtitle line below); during each Musk line, show the full caption AROUND THE WORLD at the bottom in bold yellow with a black outline — never covering the mouth, never split. Musk clips are full-frame with the mouth clearly visible; if the aspect ratio doesn't fit, fill with a blurred background from the same source — no stretching, no black bars. Hard cuts throughout, no fancy transitions.

Non-Negotiables

Alignment must be real. Cut points may only come from word-level forced alignment of the entire song (WhisperX or equivalent), refined by real onset detection of each "around" vocal onset. Never fabricate cut points from BPM grids, average spacing, drumbeats, or interpolation — not even one. If the alignment tool is unavailable, go through the installation flow; never silently degrade. Pronunciation must be complete. Every Musk clip has all three words clear and complete, mouth visible, audio and video in sync, with the initial consonant of "around" and the tail of "world" intact. The best clips may repeat, but never back-to-back; prefer repeating a great clip over padding variety with a mumbled one. Prefer swapping clips over time-stretching that mangles pronunciation (if truly needed, limit to 0.80–1.25×). Audio must be clean. If a qualified instrumental exists, use it as the bed; otherwise smoothly duck the original vocal during Musk's lines so the two never fight. No clipping, clicks, gaps, or level jumps; two-pass normalization to ~-14 LUFS, true peak ≤ -1 dBTP. Failure must be honest. If anything falls short of the bar, auto-fix and retry; if it can't be fixed, report exactly what's blocking — never deliver something substandard as finished, and never pass off a low-precision method as accurate alignment. Everything else — which 60-second window, exact shot sequencing, mix parameters, render pipeline — is your call, guided by the goals above: choose whatever maximizes the "lyric replacement" illusion, and record the reasoning behind key decisions.

Environment Preflight and Guided Installation

Before starting, run minimal live tests on FFmpeg/FFprobe, WhisperX, PyTorch, onset detection tools, Clipto access, and the output directory (actually run them — don't just check that commands exist). If anything is missing, pause, explain what it does and what breaks without it, and propose an installation plan adapted to my actual environment (OS / Apple Silicon / NVIDIA / CPU), noting download size and whether network access or model downloads are needed. Install only after I approve; prefer a project virtual environment, don't touch the system Python, don't upload my media. After installing, resume automatically from where you left off — no need for me to resend the prompt. If I decline, deliver the completed analysis and mark the run dependency_blocked.

Delivery

Create a fresh version directory (never overwrite old ones). Deliver the final MP4 with a matching SRT, plus alignment evidence, the edit plan, and verification results. Wrap up with a few sentences: which segment you chose, how many shots you cut, how many distinct Musk clips you used, the alignment error, and any trade-offs. Then hand me the MP4 directly.

 How can I redeem producthunt promo offer? Sign-up doesn't allow existing coupon code to be removed/replaced with ph offer.

 hi Chintan, when you download and install the app, there is a box for you to type in the invite code: PHLNCH. It will then automatically apply the offer. If you dont see it please DM me the email you used to sign up. I will have the team apply to your account manually. Thanks!

 Congrats on this launch, very cool!

 Thanks Gabe, your support means a lot for us!

Really curious how well the semantic search performs once you throw a huge media library at it.

 That’s exactly the scale we built Clipto for. We’ve tested it on multi-terabyte libraries with years of footage, and keeping search fast and relevant as the library grows has been a big focus for us. Would love to hear how it performs on your library if you give it a try!

TB-scale media libraries are actually one of the core use cases we built Clipto for, so semantic search isn’t limited to a small collection of files.

congrats for the launch! this is really helpful, I can imagine starting a project in Claude and asking it to find all the relevant material before I even open my editor 👀

 Exactly! That’s one of the workflows we’re most excited about. Start with the idea, let the agent understand your project and source the right moments from your entire media library, then bring those clips straight into the editor. Thanks for checking it out!

 This is exactly why we built the MCP connection — so your existing media library can become part of the Claude workflow, instead of something you have to search through separately. Would love to see how you use it!

congratulations! what has been the most surprising use case you have discovered while testing agents with large personal media libraries?

 Avery, thanks! Check out my other comment about our experiment with Elon Musk's interview clips. That's one of the surprising use cases: AI figuring out something creative based on all the media library, even beyond my expectation!

You may check out the raw output here.
. Enjoy it!

Meeting recordings might be the boring use case here, but honestly probably one of the most useful.

 Exactly. Meeting recordings may not be the flashiest use case, but they’re probably one of the most practical. You can import a recording into Clipto for local analysis, or start recording directly in the Mac/Windows app. For sensitive meetings, keeping the recording and analysis on your own Mac(or PC) can make a real difference. We hope this makes things easier for anyone who needs to keep their conversations private. Thanks for calling this out!

Agreed. Sometimes the “boring” use cases are the most valuable 😄 Meeting recordings can pile up incredibly fast, and being able to search across years of conversations and find the exact moment you need can save a ton of time.

Such a smart idea that addresses a real pain point. The hardest part of putting together a demo video is looking for that particular screenshot called "Screenshot 2026-08-19 at 12.04.44 PM" or video named "858551D6-F8A4-4CE0-9A71-DC6EF6DDD2A5" 🤦🏻‍♀️

I'm excited to try Clipto's NPL search for this and curious if the tags work!

Yes because how we as human beings remember things is different from how computer organize files. Our goal is to bridge that gap and make finding things feel as natural as remembering them.

 Thanks! That’s the kind of problem Clipto is built for :) The filename can be completely meaningless—as long as you can describe what was in the screenshot or a particular moment in the video, Clipto can search what’s actually inside the file and find it.

You don’t need to rename or tag everything first. We’d love to hear how it works with your archive!

my agents are happy
Happy agents, happy humans :)

 Coming from Fish Audio, that means a lot :) You’re doing some amazing work with AI voice. Clipto also understands speech and recognizes different speakers across local media, so it feels like there could be some really interesting ways for the two products to work together. We’d love to explore that in the future!

 that'd be fun!

does clipto work with gpt cursor Claude or any agent mcp would work?
Good question. We currently offer one-click setup for these three because they cover many of the tools creators already use. But since MCP is an open standard, Clipto also provides a universal configuration for other MCP-compatible agents. You can copy the JSON configuration from the Clipto app, paste it into your agent’s MCP settings, and save it. As long as you have the Clipto Mac app installed locally, you’re ready to go.
if I connect another agent using json config, will it be having same capabilities as for cursor Claude gpt ?
Not completely—at least for now. Our one-click setup also installs several Skills, including some that help agents produce better results in editing workflows. With the universal configuration, you still get Clipto’s core MCP capabilities, such as searching your media memory, locating files, and retrieving transcripts. We’re working to bring richer workflows to more agents as quickly as we can.

  Are those skills something I could also install manually in another agent or would need to wait for official support?

You can absolutely install them manually. The Skills we provide offer a useful starting point, but you can also find other open-source Skills and combine them to make your agent capable of much richer workflows. You may even get better results from your own setup—we’d love to hear what you try.

Looks cool! How well it handle the large media libraries ? Does performance stay consistent when searching through tb of files ?

 Good question! Once your library has been indexed, search performance stays fast and consistent, even across terabytes of media. The more challenging part is the initial analysis and indexing. We’ve added several processing modes so you can balance speed and system usage, allowing Clipto to work through a large archive in the background while you continue using your Mac/PC normally.

The number of files also matters. Generally speaking, with fewer than 2,000 files, search results are returned within 5 seconds.

Documentary editors sitting on hundreds of hours of footage are probably going to appreciate this one.

 Absolutely. Documentary was actually one of the use cases we had in mind from the beginning. When you’re sitting on hundreds or even thousands of hours of footage, being able to ask an agent for a specific moment instead of scrubbing through it all can completely change the workflow.

No more digging through footage manually. ;)
123
•••
Next