AI tool
iwant
iwant deploys open models on GCP by finding available GPUs, provisioning a machine, and setting up an OpenAI-compatible endpoint. It is for developers who want to run models in their own cloud account without managing infrastructure.

About iwant
iwant is a command-line tool that automates the deployment of open AI models on Google Cloud Platform (GCP). It locates available GPU capacity across regions, provisions the infrastructure, downloads model weights, and configures a vLLM serving endpoint. Users receive an OpenAI-compatible API endpoint, an API key, and SSH access - all running within their own cloud account.
Review
iwant tackles a specific friction point: turning a model experiment into a working API often means manually checking GPU availability across zones, handling provisioning scripts, and configuring inference servers. The tool compresses those steps into a few commands. It's a focused utility that doesn't try to be a full MLOps platform.
Key Features
- Automated GPU hunting: iwant scans GCP regions to find available GPU capacity, eliminating manual zone-by-zone checks.
- One-command model deployment: The
iwant upcommand provisions a machine, downloads weights, and sets up vLLM with a single instruction. - Model recipes: Pre-configured recipes abstract hardware requirements and serving settings for specific models like DeepSeek, Kimi, Minimax, MiMo, and Step.
- Cluster management:
iwant listshows running deployments and their connection details, whileiwant downtears down clusters. - Dry-run inspection: The
--dry-runflag lets users review the planned infrastructure before committing resources.
Pricing and Value
iwant itself is free and open source. Users pay only their own GCP infrastructure costs - compute, GPU, and networking charges billed directly by Google Cloud. The tool works with existing GCP credits. No pricing tiers or paid plans are mentioned in the available information.
Pros
- Runs entirely in the user's GCP project, keeping data and infrastructure under their control.
- Reduces multi-step GPU provisioning to a single command.
- Exposes an OpenAI-compatible endpoint, so existing client code works without modification.
- Model recipes handle hardware selection, removing guesswork about GPU type and count.
- Source code is open, allowing inspection and modification.
Cons
- GCP is the only supported cloud provider at launch; AWS and Azure users cannot use the tool.
- Users need pre-approved GPU quota and billing enabled - the tool doesn't help obtain or increase quotas.
- The tool is not well suited for teams running production workloads that require autoscaling, persistent monitoring, or multi-node orchestration beyond what a single vLLM instance provides.
iwant fits developers and researchers who want to test open models quickly without writing Terraform scripts or clicking through cloud consoles. It's most practical for short-lived experiments, demos, and evaluation workflows where spinning up and tearing down GPU instances on demand matters more than long-running service management. Teams with multi-cloud needs or strict production requirements will likely need additional tooling alongside it.











