MR
Mayur Rathi
@sickn33
⭐ 47.3k GitHub stars

multi-tenant-llm-hosting

multi-tenant-llm-hosting is an engineering AI skill with a core value of Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls. It helps developers solve real-world problems in the engineering domain, boosting efficiency, automating repetitive tasks, and optimizing workflows.

Design secure, multi-tenant LLM hosting platforms with tenant isolation, quotas, billing attribution, noisy-neighbor protection, and per-tenant policy controls.

Last verified on: 2026-10-06

Quick Facts

Category engineering
Works With Claude
Source sickn33/antigravity-awesome-skills
Stars ⭐ 47.3k
Last Verified 2026-10-06
Risk Level Low
mkdir -p ./skills/multi-tenant-llm-hosting && curl -sfL https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/multi-tenant-llm-hosting/SKILL.md -o ./skills/multi-tenant-llm-hosting/SKILL.md

Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).

Skill Content

# Multi-Tenant LLM Hosting


Host many teams/customers on shared inference infrastructure without sacrificing security, performance, or cost governance.


Prerequisites


- Kubernetes cluster with GPU node pools

- API gateway or LLM gateway (LiteLLM, Envoy, Kong)

- Prometheus + Grafana for per-tenant observability

- Redis or equivalent for rate limiting state

- Billing system or cost attribution database


Isolation Model


- Strong tenant identity on every request

- Per-tenant API keys and scoped model access

- Namespace or workload isolation for high-risk tenants

- Strict data retention and log partitioning controls


vLLM Multi-Model Serving


yaml
# vllm-deployment.yaml - Multi-model serving with vLLM
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-gpt4o-equivalent
  namespace: llm-serving
  labels:
    app: vllm
    model-tier: premium
spec:
  replicas: 3
  selector:
    matchLabels:
      app: vllm
      model-tier: premium
  template:
    metadata:
      labels:
        app: vllm
        model-tier: premium
      annotations:
        prometheus.io/scrape: "true"
        prometheus.io/port: "8080"
    spec:
      containers:
        - name: vllm
          image: vllm/vllm-openai:v0.4.1
          args:
            - "--model=/models/llama-3.1-70b"
            - "--tensor-parallel-size=2"
            - "--max-model-len=8192"
            - "--gpu-memory-utilization=0.90"
            - "--max-num-seqs=128"
            - "--enable-prefix-caching"
          ports:
            - containerPort: 8000
              name: inference
            - containerPort: 8080
              name: metrics
          resources:
            requests:
              nvidia.com/gpu: 2
              cpu: "8"
              memory: "64Gi"
            limits:
              nvidia.com/gpu: 2
              cpu: "16"
              memory: "128Gi"
          volumeMounts:
            - name: model-weights
              mountPath: /models
              readOnly: true
      volumes:
        - name: model-weights
          persistentVolumeClaim:
            claimName: premium-model-weights
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule
      nodeSelector:
        gpu-type: a100
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-economy
  namespace: llm-serving
  labels:
    app: vllm
    model-tier: economy
spec:
  replicas: 2
  selector:
    matchLabels:
      app: vllm
      model-tier: economy
  template:
    metadata:
      labels:
        app: vllm
        model-tier: economy
    spec:
      containers:
        - name: vllm
          image: vllm/vllm-openai:v0.4.1
          args:
            - "--model=/models/llama-3.1-8b"
            - "--max-model-len=4096"
            - "--gpu-memory-utilization=0.85"
            - "--max-num-seqs=256"
            - "--enable-prefix-caching"
          ports:
            - containerPort: 8000
              name: inference
            - containerPort: 8080
              name: metrics
          resources:
            requests:
              nvidia.com/gpu: 1
              cpu: "4"
              memory: "32Gi"
            limits:
              nvidia.com/gpu: 1
              cpu: "8"
              memory: "64Gi"
          volumeMounts:
            - name: model-weights
              mountPath: /models
              readOnly: true
      volumes:
        - name: model-weights
          persistentVolumeClaim:
            claimName: economy-model-weights
      tolerations:
        - key: nvidia.com/gpu
          operator: Exists
          effect: NoSchedule

Per-Tenant Quota Configuration


yaml
# tenant-quotas-configmap.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: tenant-quotas
  namespace: llm-serving
data:
  quotas.yaml: |
    tenants:
      acme-corp:
        tier: enterprise
        models_allowed:
          - llama-3.1-70b
          - llama-3.1-8b
          - nomic-embed-text
        rate_limits:
          requests_pe

🎯 Best For

  • Claude users
  • AI users

💡 Use Cases

  • Using multi-tenant-llm-hosting in daily workflow
  • Automating repetitive engineering tasks

📖 How to Use This Skill

  1. 1

    Install the Skill

    Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.

  2. 2

    Load into Your AI Assistant

    Open Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.

  3. 3

    Apply multi-tenant-llm-hosting to Your Work

    Provide context for your task — paste source material, describe your audience, or share existing work to guide the AI.

  4. 4

    Review and Refine

    Edit the AI output for accuracy, tone, and completeness. Add human insight where the AI lacks context.

❓ Frequently Asked Questions

How do I install multi-tenant-llm-hosting?

Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/multi-tenant-llm-hosting/SKILL.md, ready to use.

Can I customize this skill for my team?

Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.

⚠️ Common Mistakes to Avoid

Not reading the full skill

Skills contain important context and edge cases beyond the quick start.

🔗 Related Skills