I've been building video tooling for a few months and the thing that still gets me is how confidently a model will describe a video it barely saw.
Most pipelines hand it a frame every few seconds plus a transcript. Then you ask about pacing, or camera work, or whether the speaker sounded unsure, and it answers. Fluently. The answer just isn't grounded in anything that was actually passed in.
Hey PH, Leo here from Taiwan.
The backstory. Last month I open-sourced claude-real-video — a CLI that turns any video into scene-aware keyframes and a timestamped transcript, so an LLM can actually read it. It hit the Hacker News front page and now sits at nearly 2,000 GitHub stars. That project stays free and MIT.
The gap it left. The model could tell you what was said — never how it was shot. A slow push-in, a hard cut on the beat, a voice cracking mid-sentence: all invisible to the transcript.
What Pro adds:
- --motion — camera moves, cut rhythm, pacing curve
- --senses — sound events and voice emotion
- --speakers — who is talking, when
- OCR — every frame of on-screen text
- Interactive viewer — audit exactly what the model saw
- --ai-report — the full written breakdown, one command
Privacy. Everything runs on your machine. The only thing that ever leaves is the optional AI report, through your own API key.
Launch offer. $19 through Aug 31 with code PRODUCTHUNT — one-time purchase, keep this version forever, no subscription. Regular price $29.
https://leoaido.lemonsqueezy.com...
Built for agent pipelines, content breakdowns, and footage analysis at scale. Ask me anything — especially where it breaks.
60-second film: https://youtu.be/sw6_8E-57w4
The agent watching on its own: https://youtu.be/xFqPtcju_xo