Host your own DeepSeek, Kimi, Minimax, MiMo, Step with with one command. iwant hunts across GCP regions for available GPUs, then handles provisioning, model downloads, and vLLM setup.
One command gives you an OpenAI-compatible endpoint, an API key, and SSH access.
Our objective was to set an A/B test and see the results in 5 minutes.
We hit it.
We worked hard to make setting up a benchmark easier by 1. improving model search 2. adding smart model recommendation engine 3. adding support for publicly sharing the results
Run a series of A/B tests for your LLM setup in 15 minutes. Define model, parameters and system prompt and see what's the impact on latency cost and quality. Call it a bespoke benchmark.