Clipto MCP - Let agents source clips from terabytes of your local video

Clipto MCP gives Claude, ChatGPT, and other AI agents the ability to source clips and more from inside the videos, photos, and audio recordings stored on your computer. Instead of manually browsing files, simply describe what you need. For example, turn a script into a video by matching each sentence with your local footage; find every scene where someone mentioned a topic; create rough cuts; or search years of media as if you have a dedicated assistant editor.

Add a comment

Replies

Best
Yes - the memory infrastructure that connects your personal context to whichever AI you choose to use.

Love the idea of being able to search years of local media just by describing what you need. This could save creators and editors a ton of time.

Less digging, more creating. Really appreciate the support! 🙌

One of the worst things about recording stuff is later having to spend an annoying amount of time sifting through footage. Absolutely can't wait to try this out, love the idea.

Connecting Clipto directly to ChatGPT, Claude, or Cursor through MCP makes a lot more sense than building yet another separate AI chat.

 Exactly. That’s a big part of the thinking behind MCP for us. We don’t want to build another AI chat interface. We want to make the media understanding Clipto already has available to whichever agent you prefer to work with. Thanks Aharon!

Congrats and the team on the official PH launch of Clipto MCP! 🎉 I’ve been following since your first release, and this MCP upgrade is a game-changer – the idea of letting AI agents "understand" terabytes of local media instead of just accessing files is brilliant. "Turn a script into footage" and "search meeting decisions" are killer use cases for creators and teams alike.

Huge props for keeping everything 100% local (privacy first!) and that 2TB/24hrs indexing speed on M5 is seriously impressive. 💪

One concrete suggestion: since MCP is all about agentic workflows, could you add semantic filters like "emotional tone" (e.g., excited, serious) or "shot type" (close-up, wide) for more nuanced retrieval? Also, custom tag hierarchies (project + client) would make enterprise adoption much stickier.

Quick question: any plans to support shared indexing across NAS or external drives? Or maybe a lightweight API for developers to plug into tools like Notion? I'd love to see how far this ecosystem can go.

Wishing you a huge launch! 🔥 Can't wait to hear your thoughts!

 Thanks Zepeng! Really appreciate you following us since the first launch, and you nailed the distinction we care about most: access to files is relatively easy now, but understanding what’s actually inside years of media is a very different problem.

Love the semantic filter idea. We’re already extracting much richer information than just transcripts, and things like shot type, emotion, people, scenes, and other visual context are exactly the kind of signals we want agents to be able to reason over.

NAS and external-drive workflows are also very much on our radar, especially for professional media libraries where terabytes quickly become tens or hundreds of terabytes. And on the developer side, MCP is really just the beginning. We’d love to make Clipto’s media understanding useful well beyond the Clipto app itself.

Thanks for the thoughtful feedback and for being with us again for this launch! 🙌

Terabytes of local video is exactly the pain – my footage folder is a monster and searching it is hopeless. Does it need everything indexed first or can an agent search cold? Curious how long a first index takes

 That’s the kind of library we built Clipto for. For accurate content-level search, the media does need to be analyzed and indexed first—the Agent can’t search it cold. The good news is that the initial index only happens once; after that, Clipto only analyzes new media as you add it.

The time depends on your Mac and performance settings. In one of our M5 tests, we indexed over 2TB within 24 hours. You can also pause or resume the process and choose between Light, Balance, and Turbo modes depending on how you’re using your Mac.

  Ok 2TB in 24 hours is way better than I expected, I was picturing a full week haha. Does Turbo mode make the Mac unusable while it runs, or can I keep working through it? I'd probably kick it off on a Friday and let it chew overnight

 That Friday-night plan is a perfect use for Turbo :) It prioritizes indexing speed, so you may notice the higher resource usage if you continue working at the same time.

For everyday work, Balance dynamically shares resources between Clipto and your other apps. If you’re doing something more demanding, like rendering a video, you can switch to Light to keep more performance available for your main task!

  Perfect, Balance on weekdays and Turbo on Fridays then. Does it step down on its own when I unplug and go to battery, or is that on me to remember?

The local search part is what caught my attention. Being able to ask an AI tool to find something across your own media library feels genuinely useful.

 Thanks Shirley! That’s exactly the idea. Your media is already there, but finding the right moment across years of footage is still surprisingly hard. We want to make that entire library something you can simply talk to, whether directly in Clipto or through your favorite AI agent.

Your agent has memory now:)

The Claude/Cursor integration is probably the part I’d explore first. I’m already spending a lot of time in AI tools, so giving them access to my own content without constantly uploading individual files feels like a pretty natural next step.

 Exactly. The time you’ve spent getting comfortable with Claude, Cursor, or another AI shouldn’t lock you into a separate workflow. Clipto MCP lets the AI tools you already know connect to your indexed local media. Once connected, you can keep working in the environment you’re familiar with—without uploading files one by one.

Instead of bringing your files to AI one by one, we want to bring your memory to the AI tools you already use.

This sounds useful for large media libraries, but how long does it take to analyze and index several terabytes of footage? Can I continue using my Mac during the process, or will it consume most of the system resources?

Good question. It depends largely on the performance of your Mac. Clipto currently supports Macs with M2 chips or later, and more powerful machines will naturally complete the analysis faster. In one of our tests, an M5 Mac indexed terabytes of media within 24 hours. We’ve also done a lot of work to optimize the process. If your Mac is completely idle, you can let Clipto run at a higher speed overnight and wake up to significant progress the next morning.

 At the same time, we’ve put a lot of work into balancing speed and resource usage, so Clipto can run smoothly across a wider range of supported Macs without getting in the way of your normal work. We’ve also added pause and resume controls, along with three performance modes: Light, Balance, and Turbo.

Turbo is ideal when you’re stepping away from your Mac and want Clipto to process the library as quickly as possible. Balance dynamically manages performance while you continue working. If keeping your Mac responsive is the priority, you can switch to Light mode.

What stands out to me is that Clipto does more than transcribe media—it makes an entire archive searchable and reusable. Finding an exact scene, quote, or moment through natural language could save creators hours of manual work. Keeping everything local makes the idea even more compelling. Congrats on the launch!

Thanks Parsons! You nailed it — the goal is to make years of local media instantly searchable and reusable, without having to dig through files manually. Really appreciate the support! 🙌