MR
Mayur Rathi
@sickn33
⭐ 47.3k GitHub stars

huggingface-community-evals

huggingface-community-evals is an code AI skill with a core value of Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal. It helps developers solve real-world problems in the code domain, boosting efficiency, automating repetitive tasks, and optimizing workflows.

Curated upstream guidance for Huggingface Community Evals; use when the workflow matches the user goal.

Last verified on: 2026-10-06

Quick Facts

Category code
Works With Claude
Source sickn33/antigravity-awesome-skills
Stars ⭐ 47.3k
Last Verified 2026-10-06
Risk Level Low
mkdir -p ./skills/huggingface-community-evals && curl -sfL https://raw.githubusercontent.com/sickn33/antigravity-awesome-skills/main/skills/huggingface-community-evals/SKILL.md -o ./skills/huggingface-community-evals/SKILL.md

Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).

Skill Content

When to Use

- Use when this upstream workflow matches the user's stated goal.

- Use when the task requires the procedures documented in this skill.


# Overview


This skill is for **running evaluations against models on the Hugging Face Hub on local hardware**.


It covers:

- `inspect-ai` with local inference

- `lighteval` with local inference

- choosing between `vllm`, Hugging Face Transformers, and `accelerate`

- smoke tests, task selection, and backend fallback strategy


It does **not** cover:

- Hugging Face Jobs orchestration

- model-card or `model-index` edits

- README table extraction

- Artificial Analysis imports

- `.eval_results` generation or publishing

- PR creation or community-evals automation


If the user wants to **run the same eval remotely on Hugging Face Jobs**, hand off to the `hugging-face-jobs` skill and pass it one of the local scripts in this skill.


If the user wants to **publish results into the community evals workflow**, stop after generating the evaluation run and hand off that publishing step to `~/code/community-evals`.


> All paths below are relative to the directory containing this `SKILL.md`.


# When To Use Which Script


| Use case | Script |

|---|---|

| Local `inspect-ai` eval on a Hub model via inference providers | `scripts/inspect_eval_uv.py` |

| Local GPU eval with `inspect-ai` using `vllm` or Transformers | `scripts/inspect_vllm_uv.py` |

| Local GPU eval with `lighteval` using `vllm` or `accelerate` | `scripts/lighteval_vllm_uv.py` |

| Extra command patterns | `examples/USAGE_EXAMPLES.md` |


# Prerequisites


- Prefer `uv run` for local execution.

- Set `HF_TOKEN` for gated/private models.

- For local GPU runs, verify GPU access before starting:


bash
uv --version
printenv HF_TOKEN >/dev/null
nvidia-smi

If `nvidia-smi` is unavailable, either:

- use `scripts/inspect_eval_uv.py` for lighter provider-backed evaluation, or

- hand off to the `hugging-face-jobs` skill if the user wants remote compute.


# Core Workflow


1. Choose the evaluation framework.

- Use `inspect-ai` when you want explicit task control and inspect-native flows.

- Use `lighteval` when the benchmark is naturally expressed as a lighteval task string, especially leaderboard-style tasks.

2. Choose the inference backend.

- Prefer `vllm` for throughput on supported architectures.

- Use Hugging Face Transformers (`--backend hf`) or `accelerate` as compatibility fallbacks.

3. Start with a smoke test.

- `inspect-ai`: add `--limit 10` or similar.

- `lighteval`: add `--max-samples 10`.

4. Scale up only after the smoke test passes.

5. If the user wants remote execution, hand off to `hugging-face-jobs` with the same script + args.


# Quick Start


Option A: inspect-ai with local inference providers path


Best when the model is already supported by Hugging Face Inference Providers and you want the lowest local setup overhead.


bash
uv run scripts/inspect_eval_uv.py \
  --model meta-llama/Llama-3.2-1B \
  --task mmlu \
  --limit 20

Use this path when:

- you want a quick local smoke test

- you do not need direct GPU control

- the task already exists in `inspect-evals`


Option B: inspect-ai on Local GPU


Best when you need to load the Hub model directly, use `vllm`, or fall back to Transformers for unsupported architectures.


Local GPU:


bash
uv run scripts/inspect_vllm_uv.py \
  --model meta-llama/Llama-3.2-1B \
  --task gsm8k \
  --limit 20

Transformers fallback:


bash
uv run scripts/inspect_vllm_uv.py \
  --model microsoft/phi-2 \
  --task mmlu \
  --backend hf \
  --trust-remote-code \
  --limit 20

Option C: lighteval on Local GPU


Best when the task is naturally expressed as a `lighteval` task string, especially Open LLM Leaderboard style benchmarks.


Local GPU:


bash
uv run scripts/lighteval_vllm_uv.py \
  --model meta-llama/Llama-3.2-3B-Instruct \
  --tasks "leaderboard|mmlu|5,leaderboard|gsm8k|5" \
  --max-samples 20 \
  --use-chat-template

`accelerate` fallback:


text

🎯 Best For

  • UI designers
  • Product designers
  • Claude users
  • Software engineers
  • Development teams

💡 Use Cases

  • Generating component mockups
  • Creating design system tokens
  • Code quality improvement
  • Best practice enforcement

📖 How to Use This Skill

  1. 1

    Install the Skill

    Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.

  2. 2

    Load into Your AI Assistant

    Open Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.

  3. 3

    Apply huggingface-community-evals to Your Work

    Open your project in the AI assistant and ask it to apply the skill. Start with a small module to verify the output quality.

  4. 4

    Review and Refine

    Review AI suggestions before committing. Run tests, check for regressions, and iterate on the skill output.

❓ Frequently Asked Questions

Does this work with Figma?

Some design skills integrate with Figma plugins. Check the Works With section for supported tools.

Is huggingface-community-evals compatible with Cursor and VS Code?

Yes — this skill works with any AI coding assistant including Cursor, VS Code with Copilot, and JetBrains IDEs.

Do I need specific dependencies for huggingface-community-evals?

Check the install command and Works With section. Most code skills only require the AI assistant and your codebase.

How do I install huggingface-community-evals?

Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/huggingface-community-evals/SKILL.md, ready to use.

Can I customize this skill for my team?

Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.

⚠️ Common Mistakes to Avoid

Skipping usability testing

AI-generated designs should be validated with real users before development.

Skipping validation

Always test AI-generated code changes, even for simple refactors.

Missing dependency updates

Check if the skill requires updated dependencies or new packages.

🔗 Related Skills