MR
Mayur Rathi
@sickn33
⭐ 47.3k GitHub stars

ai-sre-incident-response

ai-sre-incident-response is an engineering AI skill with a core value of Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions. It helps developers solve real-world problems in the engineering domain, boosting efficiency, automating repetitive tasks, and optimizing workflows.

Build AI-focused SRE incident response practices for LLM outages, degraded quality, runaway cost events, and safety regressions.

Last verified on: 2026-10-06

Quick Facts

Category engineering
Works With Claude
Source sickn33/antigravity-awesome-skills
Stars ⭐ 47.3k
Last Verified 2026-10-06
Risk Level Low
mkdir -p ./skills/ai-sre-incident-response && curl -sfL https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/ai-sre-incident-response/SKILL.md -o ./skills/ai-sre-incident-response/SKILL.md

Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).

Skill Content

# AI SRE Incident Response


Apply SRE rigor to AI systems where incidents include quality regressions, unsafe outputs, and budget explosions.


When to Use This Skill


- An LLM endpoint begins returning degraded or hallucinated answers

- Token spend spikes beyond budget thresholds

- A model provider goes down and traffic must fail over

- Safety guardrails fire at abnormal rates

- A new model deployment causes latency or accuracy regression


Prerequisites


- Prometheus and Alertmanager deployed with scrape targets for AI services

- Grafana dashboards for golden signals (latency, error rate, cost, quality)

- On-call rotation configured in PagerDuty, Opsgenie, or equivalent

- Runbook repository accessible to responders

- Rollback mechanism for model and prompt versions (GitOps or feature flags)


AI Incident Classes


- **Availability incident**: model/provider unavailable, timeout storm.

- **Quality incident**: answer accuracy or tool success drops below SLO.

- **Safety incident**: harmful or policy-violating outputs increase.

- **Cost incident**: unexpected token or provider spend spike.


Severity Framework


| Severity | Criteria | Response Time | Notification |

|----------|----------|---------------|--------------|

| SEV1 | User-facing outage, compliance risk, data leak | 5 min | Page on-call + incident commander |

| SEV2 | Major degradation in key flows | 15 min | Page on-call |

| SEV3 | Limited impact or internal-only issue | 1 hour | Slack alert |

| SEV4 | Cosmetic or low-priority regression | Next business day | Ticket |


Golden Signals for AI Services


- Request success rate

- Latency (queue + generation + tool execution)

- Hallucination/groundedness proxy metrics

- Cost per minute and per tenant

- Guardrail violation rate


Prometheus Alert Rules


yaml
# prometheus-ai-alerts.yaml
groups:
  - name: ai-service-alerts
    rules:
      - alert: ModelEndpointDown
        expr: up{job="llm-inference"} == 0
        for: 2m
        labels:
          severity: sev1
        annotations:
          summary: "LLM inference endpoint {{ $labels.instance }} is down"
          runbook_url: "https://runbooks.internal/ai/model-outage"

      - alert: HighHallucinationRate
        expr: |
          rate(llm_hallucination_detected_total[10m])
          / rate(llm_requests_total[10m]) > 0.15
        for: 5m
        labels:
          severity: sev2
        annotations:
          summary: "Hallucination rate above 15% for {{ $labels.model }}"
          runbook_url: "https://runbooks.internal/ai/quality-regression"

      - alert: TokenCostExplosion
        expr: |
          sum(rate(llm_token_cost_dollars[5m])) by (tenant)
          > 0.50
        for: 3m
        labels:
          severity: sev2
        annotations:
          summary: "Token spend exceeds $0.50/min for tenant {{ $labels.tenant }}"
          runbook_url: "https://runbooks.internal/ai/cost-spike"

      - alert: LatencyP95Exceeded
        expr: |
          histogram_quantile(0.95,
            rate(llm_request_duration_seconds_bucket[5m])
          ) > 5
        for: 5m
        labels:
          severity: sev2
        annotations:
          summary: "LLM p95 latency exceeds 5s for {{ $labels.service }}"

      - alert: GuardrailViolationSpike
        expr: |
          rate(llm_guardrail_violations_total[10m])
          / rate(llm_requests_total[10m]) > 0.05
        for: 5m
        labels:
          severity: sev1
        annotations:
          summary: "Guardrail violations above 5% for {{ $labels.model }}"
          runbook_url: "https://runbooks.internal/ai/safety-incident"

      - alert: ModelQualityDrop
        expr: |
          llm_eval_score{metric="groundedness"} < 0.70
        for: 10m
        labels:
          severity: sev2
        annotations:
          summary: "Groundedness score dropped below 0.70 for {{ $labels.model }}"

      - alert: ProviderErrorRateHigh
        expr: |
          rate(llm_provider_errors_total[5m])
          / rate(llm_provider

🎯 Best For

  • UI designers
  • Product designers
  • Claude users
  • AI users

💡 Use Cases

  • Generating component mockups
  • Creating design system tokens
  • Using ai-sre-incident-response in daily workflow
  • Automating repetitive engineering tasks

📖 How to Use This Skill

  1. 1

    Install the Skill

    Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.

  2. 2

    Load into Your AI Assistant

    Open Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.

  3. 3

    Apply ai-sre-incident-response to Your Work

    Provide context for your task — paste source material, describe your audience, or share existing work to guide the AI.

  4. 4

    Review and Refine

    Edit the AI output for accuracy, tone, and completeness. Add human insight where the AI lacks context.

❓ Frequently Asked Questions

Does this work with Figma?

Some design skills integrate with Figma plugins. Check the Works With section for supported tools.

How do I install ai-sre-incident-response?

Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/ai-sre-incident-response/SKILL.md, ready to use.

Can I customize this skill for my team?

Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.

⚠️ Common Mistakes to Avoid

Skipping usability testing

AI-generated designs should be validated with real users before development.

Not reading the full skill

Skills contain important context and edge cases beyond the quick start.

🔗 Related Skills