ai-inference-service-mesh
ai-inference-service-mesh is an engineering AI skill with a core value of Use service mesh patterns for AI inference traffic management, mTLS, canary releases, policy enforcement, and cross-cluster resilience. It
helps developers solve real-world problems in the engineering domain, boosting
efficiency, automating repetitive tasks, and optimizing workflows.
Use service mesh patterns for AI inference traffic management, mTLS, canary releases, policy enforcement, and cross-cluster resilience.
Quick Facts
mkdir -p ./skills/ai-inference-service-mesh && curl -sfL https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/ai-inference-service-mesh/SKILL.md -o ./skills/ai-inference-service-mesh/SKILL.md Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).
Skill Content
# AI Inference Service Mesh
Apply Istio/Linkerd mesh controls to secure and optimize east-west AI traffic across inference microservices.
Why Mesh for AI
- Enforce mTLS between gateway, retriever, reranker, and model services
- Apply fine-grained traffic policies without app code changes
- Run progressive delivery for model-serving backends
- Observe latency hops for retrieval + generation chains
- Route inference requests by model version, tenant, or priority tier
- Protect expensive GPU-backed services from cascading failures
Prerequisites
# Install Istio with production profile
istioctl install --set profile=default \
--set meshConfig.accessLogFile=/dev/stdout \
--set meshConfig.defaultConfig.holdApplicationUntilProxyStarts=true
# Label inference namespace for sidecar injection
kubectl create namespace ai-inference
kubectl label namespace ai-inference istio-injection=enabled
# Verify installation
istioctl verify-install
istioctl analyze -n ai-inferenceCore Patterns
mTLS Strict Mode Cluster-Wide
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: default
namespace: istio-system
spec:
mtls:
mode: STRICT
---
# Namespace-level override if needed for gradual rollout
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: ai-inference-mtls
namespace: ai-inference
spec:
mtls:
mode: STRICT
portLevelMtls:
# gRPC inference port
8081:
mode: STRICT
# Prometheus metrics port - allow plaintext scraping
9090:
mode: PERMISSIVEAuthorizationPolicy Per Service Account
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
name: model-server-access
namespace: ai-inference
spec:
selector:
matchLabels:
app: model-server
action: ALLOW
rules:
- from:
- source:
principals:
- "cluster.local/ns/ai-inference/sa/api-gateway"
- "cluster.local/ns/ai-inference/sa/orchestrator"
to:
- operation:
methods: ["POST"]
paths: ["/v1/predict", "/v1/embeddings", "/v2/models/*/infer"]
---
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
name: deny-external-to-retriever
namespace: ai-inference
spec:
selector:
matchLabels:
app: vector-retriever
action: DENY
rules:
- from:
- source:
notNamespaces: ["ai-inference"]Egress Policy for Approved Model Endpoints
apiVersion: networking.istio.io/v1alpha3
kind: ServiceEntry
metadata:
name: openai-api
namespace: ai-inference
spec:
hosts:
- api.openai.com
ports:
- number: 443
name: https
protocol: TLS
resolution: DNS
location: MESH_EXTERNAL
---
apiVersion: networking.istio.io/v1alpha3
kind: DestinationRule
metadata:
name: openai-api-tls
namespace: ai-inference
spec:
host: api.openai.com
trafficPolicy:
tls:
mode: SIMPLE
connectionPool:
http:
h2UpgradePolicy: UPGRADE
tcp:
maxConnections: 50
---
apiVersion: security.istio.io/v1beta1
kind: AuthorizationPolicy
metadata:
name: restrict-egress
namespace: ai-inference
spec:
action: ALLOW
rules:
- to:
- operation:
hosts:
- "api.openai.com"
- "models.anthropic.com"
- "*.blob.core.windows.net"Traffic Management
VirtualService for A/B Model Testing
apiVersion: networking.istio.io/v1alpha3
kind: VirtualService
metadata:
name: model-server
namespace: ai-inference
spec:
hosts:
- model-server
http:
# Route by header for explicit model version selection
- match:
- headers:
x-model-version:
exact: "v2-experimental"
route:
- destination:
host: model-server
subset: v2-experimental
timeout: 120s
# Route by header for A/B test cohort
- match:
- headers:
x-ab-cohort:
exact: "treatment"
route:
- destination:
host: model-server
🎯 Best For
- Claude users
- AI users
💡 Use Cases
- Using ai-inference-service-mesh in daily workflow
- Automating repetitive engineering tasks
📖 How to Use This Skill
- 1
Install the Skill
Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.
- 2
Load into Your AI Assistant
Open Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.
- 3
Apply ai-inference-service-mesh to Your Work
Provide context for your task — paste source material, describe your audience, or share existing work to guide the AI.
- 4
Review and Refine
Edit the AI output for accuracy, tone, and completeness. Add human insight where the AI lacks context.
❓ Frequently Asked Questions
How do I install ai-inference-service-mesh?
Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/ai-inference-service-mesh/SKILL.md, ready to use.
Can I customize this skill for my team?
Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.
⚠️ Common Mistakes to Avoid
Not reading the full skill
Skills contain important context and edge cases beyond the quick start.