MR
Mayur Rathi
@sickn33
⭐ 47.3k GitHub stars

ollama-stack

ollama-stack is an engineering AI skill with a core value of Run local LLM workloads with Ollama, Open WebUI, and GPU-aware tuning for private development environments. It helps developers solve real-world problems in the engineering domain, boosting efficiency, automating repetitive tasks, and optimizing workflows.

Run local LLM workloads with Ollama, Open WebUI, and GPU-aware tuning for private development environments.

Last verified on: 2026-10-06

Quick Facts

Category engineering
Works With Claude
Source sickn33/antigravity-awesome-skills
Stars ⭐ 47.3k
Last Verified 2026-10-06
Risk Level Low
mkdir -p ./skills/ollama-stack && curl -sfL https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/ollama-stack/SKILL.md -o ./skills/ollama-stack/SKILL.md

Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).

Skill Content

# Ollama Stack


Deploy a local LLM stack for offline and privacy-first workflows.


When to Use This Skill


Use this skill when:

- Setting up private/local LLM inference for development

- Building air-gapped AI environments

- Running models on personal hardware (Mac, Linux, Windows with GPU)

- Creating team-shared inference endpoints

- Prototyping before committing to cloud LLM APIs


Prerequisites


- 8 GB+ RAM (16 GB+ recommended for 7B+ models)

- For GPU acceleration: NVIDIA GPU with 6 GB+ VRAM, or Apple Silicon Mac

- Docker (for containerized deployment)

- 20 GB+ disk for model storage


Quick Start


bash
# Install Ollama
curl -fsSL https://ollama.com/install.sh -o /tmp/install-ollama.sh && sh /tmp/install-ollama.sh && rm /tmp/install-ollama.sh

# Start the server
ollama serve

# Pull and run a model
ollama pull llama3.1:8b
ollama run llama3.1:8b "Explain Kubernetes pods in one paragraph"

# List available models
ollama list

# Pull specific quantization
ollama pull llama3.1:8b-instruct-q4_K_M

Model Selection Guide


| Model | Size | VRAM | Best For |

|-------|------|------|----------|

| `llama3.1:8b` | 4.7 GB | 6 GB | General chat, coding |

| `llama3.1:70b` | 40 GB | 48 GB | Complex reasoning |

| `codellama:13b` | 7.4 GB | 10 GB | Code generation |

| `mistral:7b` | 4.1 GB | 6 GB | Fast general tasks |

| `mixtral:8x7b` | 26 GB | 32 GB | High-quality MoE |

| `nomic-embed-text` | 274 MB | 1 GB | Embeddings for RAG |

| `llava:13b` | 8 GB | 10 GB | Vision + text |

| `deepseek-coder-v2:16b` | 9 GB | 12 GB | Code generation |

| `qwen2.5:14b` | 9 GB | 12 GB | Multilingual, reasoning |


Docker Compose — Full Stack


yaml
# docker-compose.yml
services:
  ollama:
    image: ollama/ollama:latest
    container_name: ollama
    restart: unless-stopped
    ports:
      - "11434:11434"
    volumes:
      - ollama_data:/root/.ollama
    environment:
      - OLLAMA_HOST=0.0.0.0
      - OLLAMA_NUM_PARALLEL=4
      - OLLAMA_MAX_LOADED_MODELS=2
      - OLLAMA_FLASH_ATTENTION=1
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    healthcheck:
      test: ["CMD", "curl", "-f", "http://localhost:11434/api/tags"]
      interval: 30s
      timeout: 10s
      retries: 3

  open-webui:
    image: ghcr.io/open-webui/open-webui:main
    container_name: open-webui
    restart: unless-stopped
    ports:
      - "3000:8080"
    volumes:
      - webui_data:/app/backend/data
    environment:
      - OLLAMA_BASE_URL=http://ollama:11434
      - WEBUI_AUTH=true
      - WEBUI_SECRET_KEY=${WEBUI_SECRET_KEY:-change-me-in-production}
      - DEFAULT_MODELS=llama3.1:8b
    depends_on:
      ollama:
        condition: service_healthy

  litellm:
    image: ghcr.io/berriai/litellm:main-latest
    container_name: litellm
    restart: unless-stopped
    ports:
      - "4000:4000"
    volumes:
      - ./litellm-config.yaml:/app/config.yaml
    command: ["--config", "/app/config.yaml"]
    depends_on:
      ollama:
        condition: service_healthy

volumes:
  ollama_data:
  webui_data:

LiteLLM Proxy Config


yaml
# litellm-config.yaml
model_list:
  - model_name: llama3
    litellm_params:
      model: ollama/llama3.1:8b
      api_base: http://ollama:11434
  - model_name: codellama
    litellm_params:
      model: ollama/codellama:13b
      api_base: http://ollama:11434
  - model_name: embeddings
    litellm_params:
      model: ollama/nomic-embed-text
      api_base: http://ollama:11434

general_settings:
  master_key: sk-local-dev-key
  max_budget: 0  # unlimited for local

API Usage


Ollama exposes an OpenAI-compatible API:


bash
# Chat completion
curl http://localhost:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1:8b",
    "messages": [{"role": "user", "content": "Hello"}],
    "stream": false
  }'

# Embeddings
curl http://localhost:11434/v1/embeddings \
  -H "Conte

🎯 Best For

  • UI designers
  • Product designers
  • Claude users
  • AI users

💡 Use Cases

  • Generating component mockups
  • Creating design system tokens
  • Using ollama-stack in daily workflow
  • Automating repetitive engineering tasks

📖 How to Use This Skill

  1. 1

    Install the Skill

    Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.

  2. 2

    Load into Your AI Assistant

    Open Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.

  3. 3

    Apply ollama-stack to Your Work

    Provide context for your task — paste source material, describe your audience, or share existing work to guide the AI.

  4. 4

    Review and Refine

    Edit the AI output for accuracy, tone, and completeness. Add human insight where the AI lacks context.

❓ Frequently Asked Questions

Does this work with Figma?

Some design skills integrate with Figma plugins. Check the Works With section for supported tools.

How do I install ollama-stack?

Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/ollama-stack/SKILL.md, ready to use.

Can I customize this skill for my team?

Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.

⚠️ Common Mistakes to Avoid

Skipping usability testing

AI-generated designs should be validated with real users before development.

Not reading the full skill

Skills contain important context and edge cases beyond the quick start.

🔗 Related Skills