Play games created by AI without knowing which model made them. Compare and rate leading AI models based on game concept, visual design, and coding quality. VeilPlays turns AI model evaluation into an interactive, blind benchmark powered by real player feedback.
What could you build with GPT-5.6 that wasn’t practical before?
Maker
GPT-5.6 made it practical for me, as a solo builder, to take VeilPlays from a working prototype to a production-ready service. I used it to review infrastructure and security, identify and fix bugs, prepare deployment, and continue improving the service after launch.
VeilPlays itself is a blind benchmark where users play games created by different AI models from the same prompt in a constrained, single-shot process, then evaluate the results without knowing which model made them. GPT-5.6 helped me handle the breadth of production work that would otherwise have required significantly more time or specialized support.
Report
Maker
📌
Hi Product Hunt!
I built VeilPlays because comparing AI models through benchmark scores alone can feel abstract and disconnected from what users actually experience.
VeilPlays turns AI model evaluation into an interactive blind test. Multiple AI models are given the same prompt and must create a complete playable game in a constrained, single-shot generation process—without follow-up prompts, manual revisions, or additional development requests. Their identities remain hidden while you play.
Not every model successfully completed the game under these constraints, and some outputs were only partially functional. I kept those outcomes as part of the experiment, because the ability to deliver a complete, playable result in a single attempt is itself an important measure of model performance.
After playing, you can compare and rate the results based on gameplay, visual design, creativity, and code quality—without being influenced by model names or brand reputation.
One challenge I’m exploring is how to turn casual gameplay into meaningful human evaluation. Many people enjoy trying the games, but the ratings are what make the blind comparison genuinely useful.
So when you try VeilPlays, I’d really appreciate it if you could complete at least one evaluation after playing. It only takes a moment, and your ratings directly contribute to the comparison results.
I’d also love your feedback on the blind evaluation experience, the rating criteria, and which AI models or game categories you’d like to see next.
Thanks for checking it out!