It has some of the typical issues you see with AI video generation. You ask it to change 1 thing, and 3 other things may change with it, so you often have to generate several versions and pick the best one rather than continuously refining a single version.
They need tutorials and guides, so users have a direction on how to use it with minimal waste, because there is definitely a learning curve, and currently I use ChatGPT to help me create the prompts. They also need guide around choosing a model. There are many options, but it isn’t always clear which model is best for a particular type of video without trying them. I haven’t tested every model and tend to stick with the ones I’m already familiar with.
The first-frame reference is very useful, but I’d love to see that expanded into a true storyboard feature. Being able to provide several reference frames for how a scene should progress would make it much easier to generate a short video that is close to what you envisioned.
I’ve worked around this by using a storyboard as the first-frame reference for a 15-second generation. It works surprisingly well, but it does waste 1 to 2 secs as it trying to animate the storyboard image itself, before it generates the actual usable video.
It would be great if it also offered a voice selection, so if I do want to use voice, where the avatar speaks, a voice could be selected each time to stay consistent.