Launching today
iwant
Find available GPUs and host open models in one command
10 followers
Find available GPUs and host open models in one command
10 followers
Host your own DeepSeek, Kimi, Minimax, MiMo, Step with with one command. iwant hunts across GCP regions for available GPUs, then handles provisioning, model downloads, and vLLM setup. One command gives you an OpenAI-compatible endpoint, an API key, and SSH access.








Just a couple of commands.
iwant auth checks your cloud credentials and shows setup guidance.
You use your own GCP project, so the deployment lives alongside your existing infrastructure. You’ll need billing enabled, GPU quota, and available capacity. This command helps you check the connection before moving on to a model launch.
iwant up is the starting point. Pick a model and iwant provisions the GPU, sets up the serving environment, and waits for the API to respond.
You get an OpenAI-compatible endpoint, an API key, and SSH access. Everything runs in your own cloud account. GCP is currently supported.
iwant list - see what’s running
When you’re trying several models, it gives you a place to check which clusters exist and find their connection details. You can also use the cluster names from this view with iwant ssh and iwant down.
iwant down - finish the deployment
It removes a cluster when you’re finished.
Run it without a name to choose a deployment interactively, or pass the exact cluster name for scripts. Then use iwant list to confirm it’s gone.
iwant up gpt-oss-20b --infra gcp
GPT OSS 20B is the first-run example in the README. Its recipe requests one L4 GPU which is wildly available. Once the server is ready, use its base URL and API key in your OpenAI-compatible client settings.