Grok Bot template · Cloud and DevOps
GPU Kubernetes Operations
Keeps GPU Kubernetes clusters healthy, well-scheduled and cost-efficient for AI workloads.
What it can do
The skills built into this template. Each one tells Grok when to use it, what it needs from you and how to check its work.
- Plan GPU Node Pools
- Install and Verify GPU Operator
- Deploy Standalone Device Plugin
- Configure MIG Partitioning
- Request MIG Slices in Workloads
- Set Up GPU Time-Slicing
- Wire Up DCGM Monitoring
- Build GPU Autoscaling Policies
- Diagnose GPU Scheduling and Health Problems
- Review GPU Cost Efficiency
Apps it works with
Connect these in Grok for the best results. It also works without them: you paste the information in.
Kubernetes cluster access (read-only by default)PrometheusHelmNVIDIA GPU OperatorDCGM exporter
The full template
For members
The complete GPU Kubernetes Operations template: its identity, every skill step by step, its limits and its first-run questions, ready to paste into a new Grok Bot. Members get it, and every other template here.
Jobs this template suits
Our AI checked this template against 500 jobs; these get the most out of it. Each job links to its learning path.