Clipto - Fully local, natural language search over terabytes of media

Like Google Photos, but fully local. Turn the terabytes of video, audio, meetings, and files you work with into searchable memories, without uploading anything to the cloud. Clipto automatically tags people, dialogue, and scenes, so you can instantly find any moment buried in your media just by describing what you're looking for. It's fast too: on a MacBook Pro M5, Clipto indexed 2TB of videos in just 24 hours.

Add a comment

Replies

Best

Just downloaded the Mac app—the UI is surprisingly clean for a local AI tool. How many languages does the transcription support currently?

 Thanks, really glad you like the UI. We currently support transcription in 99+ languages, so it should work well for multilingual audio and video content across different workflows

That is good

But isnt it better to keep your data on cloud no one wants their system to have that much data

 Good question! That’s fair — cloud storage can be convenient, especially if you want everything synced across devices. But for a lot of filmmakers, editors, and creators, the problem is that raw footage is huge and often sensitive. Uploading terabytes of media can be slow, expensive, and not always something people are comfortable with. Clipto is built for the other workflow: your media already lives on your local drives, and the AI helps you make it searchable without uploading everything first. So it’s not really “cloud vs local” for everyone. If you prefer cloud storage, that can still work for you.:)

yeah people can have different use cases

This feels like what the Apple 'Photos' search should have been for professional video files. Super impressed.

 High praise — thank you. 😊

Apple Photos is great for memories. Clipto is built for work: terabytes of raw footage, interviews, production assets — all searchable locally, offline, instantly.

Glad it resonates. Let me know what you find when you try it.

Congrats on the PH launch, & team! 🎉 “Fully local + natural language search” is such a killer combo—especially after that desert story. I’ve wasted hours scrubbing through raw footage myself, so I feel that pain. What I love: the 2TB/day indexing speed on M5 is seriously impressive. And the fact that nothing leaves your drive? 👏 Privacy-first done right. One idea to make it even stickier: allow users to manually name detected faces (e.g., label “Mom” or “Client A”). Right now auto-tagging is great, but custom naming would turn “a person” into your person. Imagine searching “Grandpa’s birthday” and actually finding it. Does Clipto already support that? If not, would love to see it in the roadmap! Congrats again—can’t wait to try it out. 🔥

 Thanks, Zepeng!

Yes, Clipto already supports this today!

You can assign custom names to detected faces, so instead of searching for “a person”, you can search for people that actually matter to you, such as family members, friends, clients, or collaborators.

We’ve found that once people start organizing media around real identities, search becomes much more powerful. Instead of “find a woman speaking on stage,” you can search for things like “Mom’s speech”, “Client A interview”, or “John at the conference.”

We think that’s an important step toward turning media search into a true personal memory system.

Very cool Idea!! If it woks fully in local, you must be using small LLM/VLM on local device. In that case do you see any memory Or CPU issues? How do you fix that ?

 That's a great question.

First, choose a smaller model.

Second, slim it down through optimization.

Finally, schedule tasks flexibly based on how busy your computer is — that is, 'model miniaturization itself, compression optimization, and flexible task scheduling based on the user's machine usage.

the 'store everything but remember nothing' line is the whole thing imo. the part people underrate is that the hard bit was never the search, its doing the indexing on-device without melting the laptop or quietly shipping stuff to a server, which is exactly why most tools just punt it to the cloud. respect for taking the harder path. one thing im curious about: once the first 2TB is indexed, is re-indexing incremental as you add footage, or does it re-chew the whole library? thats kind of the thing that decides whether this stays usable for anyone whose archive keeps growing

 That’s a very insightful observation.

We actually agree with your premise. For local AI products, you have to solve the hardest problem first. If you can’t make large-scale on-device indexing practical, everything else is just a demo.

As for indexing, it’s incremental. Once your library has been processed, Clipto only analyzes newly added or changed files. It doesn’t re-chew the entire archive every time.

A lot of our engineering work has gone into task scheduling, indexing pipelines, and resource management to make sure growing libraries remain practical over time.

That’s ultimately the difference between a product that works for a 50GB library and one that can keep scaling as your archive grows year after year.

One question on the indexing- Most of the examples here are professional, but does this work for a personal or family media archive too? My real-world mess: tens of thousands of files spread across folders, iPhone and Canon EOS naming mixed together, some GoPro footage, some with location metadata and plenty without. Once it's all sitting on a drive or OneDrive, Apple's native location mapping doesn't help anymore, so all I'm left with is filenames and maybe a creation date. Does Clipto index that kind of unstructured personal pile and make it searchable by what's actually in the footage, regardless of filename or missing metadata?

 Yes, absolutely.

While many of our examples come from professional media workflows, this is actually a problem we think about a lot.

In many ways, a family archive is even harder than a professional one. You have photos and videos scattered across phones, cameras, external drives, cloud folders, and years of inconsistent naming conventions.

Clipto is designed to index the content itself, not just filenames or metadata. So even when metadata is missing or incomplete, it can still use visual content, people, dialogue, scenes, objects, and other signals to make the library searchable.

That’s why we’re particularly interested in the “messy archive” use case. Most people don’t have a well-organized media library. They have exactly what you described: tens of thousands of files accumulated over years.

When you’re trying to find something in that archive, what are the searches you most wish you could do today but can’t?

 Exactly the "how would I use this" question I thought about when reading about Clipto.

Last year I tried to put together a photobook for my dad's 83rd. We live abroad so he doesn't get to see the kids much, and I wanted to give him something physical showing them growing up across different settings, school, trips, everyday moments.

The search I wished I had then was something like "one good photo of each of my girls, per year, across these settings."
What I actually did was scroll through years of camera rolls by hand, copy/pasting the useful ones into a shortlist folder, which I then went through a second time to select/crop the ones for the photobook. There was simply nothing I could ask to "show me the kids at this school event" or "find me the photos from trip xyz".

If Clipto can do "find photos of [person] at [kind of moment] over time" on an archive with no consistent file names, that's the feature that would have saved me a whole weekend. Literally.

 This is one of the most compelling use cases I’ve heard so far.

What you’re describing isn’t really a search problem. It’s a memory problem.

The challenge wasn’t that the photos didn’t exist. It was that the context, relationships, and stories connecting them were buried across years of camera rolls and folders.

The query you wanted — “one good photo of each of my daughters, across different stages of their lives and different kinds of moments” — is exactly the kind of experience we think becomes possible when media is organized around people and memories instead of filenames and folders.

Thank you for sharing this story. It’s a great reminder that the most valuable archives are often personal ones!

 happy to share if it can be of help in making tools like Clipto more targeted. Can‘t wait to try it 👍

I’ve been looking for a way to search my local media assets without opening every single folder. This just saved me an hour of digging today.

 That’s exactly the problem we built Clipto for. 😄

Too often we know we have the clip somewhere, but finding it means opening folder after folder. Glad Clipto saved you an hour today!

What languages do you support for dialogue search? Does the search support compound queries such as “Find clips that do X and Y” resulting in three sets of clips: X only, Y only, both X and Y?

 Great questions!

For dialogue search, we support 100+ languages through our speech recognition pipeline, including English, French, Italian, Spanish, Japanese, Chinese, and many others. As long as the language is supported by the underlying ASR models, the dialogue becomes searchable. Accuracy can vary by language, audio quality, accents, and recording conditions, but we’ve found it works very well across most major languages.

For compound queries, yes. We don’t treat search as simple keyword matching. We use semantic retrieval and reranking to understand the intent behind a query. For something like:

“Find clips that contain both X and Y”

clips matching both concepts would typically rank highest, while clips matching only X or only Y may still appear further down the results if they are semantically relevant. In practice, the system tries to optimize for the user’s intent rather than applying strict boolean logic.

We’d love to hear more about the workflows you’re thinking about. This is an area we’re actively improving.

Does it support video transcription search?

 Yes, it does. Clipto supports video transcription search. Beyond transcription, AI can also generate a summary for your video. You can click into any single file to open its detail page, where you’ll see the transcript, summary, and a chat box. From there, you can ask questions about the video, and the AI will answer based on the content of that specific file:)