Reviewers consistently say Clueso saves major time turning rough screen recordings into polished demos, tutorials, onboarding videos, training content, and docs. They repeatedly praise how easy it is for non-experts, with specific mentions of auto-zoom, captions, branded intros, voiceovers, language quality, and PPTX narration. Several teams say work that once took days or multiple people now takes minutes or a single person. Support also gets strong marks for being responsive. The few caveats are minor: some beta features still need refinement.
Upstream
Curious about the multi-frame feedback: does the agent actually get rendered frames back as images after every edit, and how do you keep that from eating the whole context on a 90-second video?
Clueso
@louislecat It actually gets a SINGLE image, that contains a thumbnail of sorts for all frames it's requested for. Because most modern models are good, and a good video tends to have low fidelity (ie, there's not a lot going on screen), this actually works great in helping the model understand multiple frames in one image. This is actually how professional video editors make a storyboard as well: they combine multiple frames into one visualisable frame that you can skim through.
It doesn't get back after every edit, but only when it chooses to take a look post making edits to one clip of the video. Typically, the model will request the MCP for a dense number of frames for a specific animation or a sparse set to get a sense of the whole video.
This is genuinely useful.
Video creation is always one of those things that takes way longer than it should.