MR
Mayur Rathi
@sickn33
⭐ 47.3k GitHub stars

gpu-kubernetes-operations

gpu-kubernetes-operations is an engineering AI skill with a core value of Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. It helps developers solve real-world problems in the engineering domain, boosting efficiency, automating repetitive tasks, and optimizing workflows.

Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.

Last verified on: 2026-10-06

Quick Facts

Category engineering
Works With Claude
Source sickn33/antigravity-awesome-skills
Stars ⭐ 47.3k
Last Verified 2026-10-06
Risk Level Low
mkdir -p ./skills/gpu-kubernetes-operations && curl -sfL https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/gpu-kubernetes-operations/SKILL.md -o ./skills/gpu-kubernetes-operations/SKILL.md

Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).

Skill Content

# GPU Kubernetes Operations


Run resilient and cost-efficient GPU clusters for production AI workloads.


When to Use This Skill


- Setting up GPU node pools in Kubernetes for AI inference or training

- Configuring NVIDIA device plugin and GPU operator

- Implementing MIG partitioning to share GPUs across workloads

- Building GPU-aware autoscaling policies

- Monitoring GPU health with DCGM and Prometheus

- Troubleshooting GPU scheduling, driver, or OOM issues


Prerequisites


- Kubernetes 1.28+ cluster with GPU-capable nodes

- NVIDIA GPUs (A10, L4, A100, H100, or similar)

- NVIDIA drivers installed on nodes (535+ recommended)

- Helm 3 for operator and plugin installation

- Prometheus stack for metrics collection


NVIDIA GPU Operator Installation


The GPU Operator automates driver, toolkit, device plugin, and DCGM deployment.


bash
# Add NVIDIA Helm repo
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia
helm repo update

# Install GPU Operator
helm install gpu-operator nvidia/gpu-operator \
  --namespace gpu-operator \
  --create-namespace \
  --set driver.enabled=true \
  --set toolkit.enabled=true \
  --set devicePlugin.enabled=true \
  --set dcgmExporter.enabled=true \
  --set migManager.enabled=true \
  --set nodeStatusExporter.enabled=true \
  --version v24.3.0

# Verify installation
kubectl get pods -n gpu-operator
kubectl get nodes -o json | jq '.items[].status.allocatable["nvidia.com/gpu"]'

NVIDIA Device Plugin (Standalone)


If not using the GPU Operator, deploy the device plugin directly.


yaml
# nvidia-device-plugin.yaml
apiVersion: apps/v1
kind: DaemonSet
metadata:
  name: nvidia-device-plugin
  namespace: kube-system
spec:
  selector:
    matchLabels:
      name: nvidia-device-plugin
  template:
    metadata:
      labels:
        name: nvidia-device-plugin
    spec:
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      priorityClassName: system-node-critical
      containers:
        - name: nvidia-device-plugin
          image: nvcr.io/nvidia/k8s-device-plugin:v0.15.0
          securityContext:
            privileged: true
          env:
            - name: FAIL_ON_INIT_ERROR
              value: "false"
            - name: DEVICE_SPLIT_COUNT
              value: "1"
            - name: DEVICE_LIST_STRATEGY
              value: "envvar"
          volumeMounts:
            - name: device-plugin
              mountPath: /var/lib/kubelet/device-plugins
      volumes:
        - name: device-plugin
          hostPath:
            path: /var/lib/kubelet/device-plugins

MIG (Multi-Instance GPU) Partitioning


MIG allows a single A100 or H100 to be split into isolated GPU instances.


yaml
# mig-config.yaml - ConfigMap for MIG Manager
apiVersion: v1
kind: ConfigMap
metadata:
  name: mig-parted-config
  namespace: gpu-operator
data:
  config.yaml: |
    version: v1
    mig-configs:
      # 7 small instances for inference microservices
      all-1g.10gb:
        - devices: all
          mig-enabled: true
          mig-devices:
            "1g.10gb": 7

      # 3 medium instances for mid-size models
      all-2g.20gb:
        - devices: all
          mig-enabled: true
          mig-devices:
            "2g.20gb": 3

      # Mixed: 1 large + 2 small
      mixed-inference:
        - devices: all
          mig-enabled: true
          mig-devices:
            "3g.40gb": 1
            "1g.10gb": 4

      # Full GPU for training (no partitioning)
      all-disabled:
        - devices: all
          mig-enabled: false

bash
# Apply MIG profile to a node
kubectl label nodes gpu-node-01 nvidia.com/mig.config=all-1g.10gb --overwrite

# Verify MIG instances
kubectl exec -it nvidia-device-plugin-xxxxx -n kube-system -- nvidia-smi mig -lgi

# Check available MIG resources
kubectl get nodes gpu-node-01 -o json | jq '.status.allocatable | with_entries(select(.key | startswith("nvidia.com")))'

Requesting MIG Slices in Pods


yaml
# p

🎯 Best For

  • Claude users
  • AI users

💡 Use Cases

  • Using gpu-kubernetes-operations in daily workflow
  • Automating repetitive engineering tasks

📖 How to Use This Skill

  1. 1

    Install the Skill

    Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.

  2. 2

    Load into Your AI Assistant

    Open Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.

  3. 3

    Apply gpu-kubernetes-operations to Your Work

    Provide context for your task — paste source material, describe your audience, or share existing work to guide the AI.

  4. 4

    Review and Refine

    Edit the AI output for accuracy, tone, and completeness. Add human insight where the AI lacks context.

❓ Frequently Asked Questions

How do I install gpu-kubernetes-operations?

Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/gpu-kubernetes-operations/SKILL.md, ready to use.

Can I customize this skill for my team?

Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.

⚠️ Common Mistakes to Avoid

Not reading the full skill

Skills contain important context and edge cases beyond the quick start.

🔗 Related Skills