Reviewers mostly see bolt.new as a fast way to turn prompts into working browser-based apps, especially for mockups, landing pages, MVPs, and simple SaaS ideas without local setup. Several developers praise its full-stack workflow, and founders behind products like and specifically highlight quick scaffolding and smooth browser execution. The main complaints are consistent: token use feels opaque or expensive on larger projects, the AI can loop or rewrite too much when fixing bugs, and some paying users report poor support and billing or token-access problems.
Bolt Forge is a new agent inside Bolt.new that runs entirely on open-source models.
Sitting next to Standard and Max in the agent picker, this one is built only on open models - GLM 5.3 Flash and GLM 5.3 as the main pair, Kimi K3 and DeepSeek v4 Pro as experimental picks.
What makes it different: Every individual Pro plan gets up to 50X more usage on Forge, free, through Oct 14, 2026.
The trade is opt-in: your build sessions (prompts, code, fix traces) get anonymized and go toward training new open-weight models with Arcee AI, a US open-model lab. You can stop sharing anytime by switching back to Standard or Max.
Key features:
One monthly usage bar, no daily limits
Hits 100% → auto-switches to Standard, no overage
Scores 92.2 vs 101 for Bolt's top paid model on their internal benchmark (~91% capability)
Runs on WebContainers in-browser, no server cost
Duplicate your project before testing serious builds, still a research preview
Who it's for: Builders who burn through credits fast in the brainstorm/draft phase and want room to experiment without touching production usage.
P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified → @rohanrecommends
Incredible work from the Bolt team. As far as I know, it's the first of the vibe-coding solutions to really go all in with open-source models, and I commend you!
Question for the makers—right now it handles full-stack in the browser tab really well, but what’s the vision for connecting to existing local microservices or private databases during the prototyping phase?
Dial
"91% of Bolt's top paid model on their internal benchmark" is doing a lot of work in that sentence - internal benchmarks tend to be picked to flatter the new thing. Is there a public eval or a task set outsiders can rerun, or is 91% something we just have to take on faith for now?