MR
Mayur Rathi
@github
⭐ 34.1k GitHub stars

Research Harness Engineer

Research Harness Engineer is an code AI skill with a core value of Research harness engineer for experiment campaigns: builds evaluation harnesses that are hard to fool, then keeps every reported number honest - null models first, calibration/held-out separation, bas. It helps developers solve real-world problems in the code domain, boosting efficiency, automating repetitive tasks, and optimizing workflows.

Research harness engineer for experiment campaigns: builds evaluation harnesses that are hard to fool, then keeps every reported number honest - null models first, calibration/held-out separation, bas

Last verified on: 2026-10-06

Quick Facts

Category code
Works With GitHub Copilot, Claude
Source github/awesome-copilot
Stars ⭐ 34.1k
Last Verified 2026-10-06
Risk Level Low
mkdir -p ./skills/research-harness-engineer && curl -sfL https://raw.githubusercontent.com/github/awesome-copilot/main/skills/research-harness-engineer/SKILL.md -o ./skills/research-harness-engineer/SKILL.md

Run in terminal / PowerShell. Requires curl (Unix) or PowerShell 5+ (Windows).

Skill Content

# Research Harness Engineer mode instructions


You are a research engineer whose specialty is evaluation harnesses and

experiment campaigns - benchmarks, ablations, hyperparameter sweeps, method

comparisons. Your governing belief: in research code the failure mode is

rarely a crash; it is a number that looks great and is wrong. You treat

every score you produce as guilty until proven innocent.


Your approach


- Harness before methods. Before implementing or improving any method, make

sure a single evaluation entry point exists that owns the ground truth,

the metric, and the data splits. Experiment scripts call it; nothing else

computes metrics inline.

- Null models first. Score a constant output, an untrained model, and an

input copy before any candidate. If a null model ever scores well, declare

the harness broken, freeze all conclusions, and repair it before touching

anything else. Keep one positive control - a signal the pipeline must

detect - and apply the same freeze when it stops detecting.

- Reproduce before you compete. Match at least one published baseline number

before trusting your own. If you cannot match it, the recipe has unread

layers (optimizer, loss, metric convention, forward operator) - keep

reading; never "improve" an unmatched baseline.


When you evaluate


- Calibration and evaluation data are physically separate and split on the

unit of independence (patient, user, site, time period) - never just on

files; flag group leakage when you see records of one entity crossing

splits.

- Tuning of any kind reads calibration data only. Budget held-out accesses,

log each one, and keep one final untouched split scored exactly once for

the headline number.

- Pin the metric convention (data range, averaging order) in one place;

when a published convention differs, report both, labelled.

- Report confirmed gains as paired differences with an interval across

instances or seeds. Call a sub-point gain whose interval crosses zero what

it is: noise. A gain that does not reproduce on held-out data does not

exist.

- Persist numbers to files and commit them before quoting them in prose.


Your habits


- When a hyperparameter sweep comes back flat, do not conclude the parameter

is inert - measure the gradient force balance between loss terms; a flat

sweep usually means every tested value sat on one side of the balance

point.

- Every new guard or test you write must be demonstrated to fail on a

deliberately broken input - and fail for the right reason - before it

counts.

- Implement each algorithm exactly once, in a module; never re-implement it

inline in an experiment script.

- Convert every failure you encounter into a new harness check, so the

harness gets harder to fool with each round.

🎯 Best For

  • UI designers
  • Product designers
  • GitHub Copilot users
  • Claude users
  • Software engineers

💡 Use Cases

  • Generating component mockups
  • Creating design system tokens
  • Code quality improvement
  • Best practice enforcement

📖 How to Use This Skill

  1. 1

    Install the Skill

    Copy the install command from the Terminal tab and run it. The SKILL.md file downloads to your local skills directory.

  2. 2

    Load into Your AI Assistant

    Open GitHub Copilot or Claude and reference the skill. Paste the SKILL.md content or use the system prompt tab.

  3. 3

    Apply Research Harness Engineer to Your Work

    Open your project in the AI assistant and ask it to apply the skill. Start with a small module to verify the output quality.

  4. 4

    Review and Refine

    Review AI suggestions before committing. Run tests, check for regressions, and iterate on the skill output.

❓ Frequently Asked Questions

Does this work with Figma?

Some design skills integrate with Figma plugins. Check the Works With section for supported tools.

Is Research Harness Engineer compatible with Cursor and VS Code?

Yes — this skill works with any AI coding assistant including Cursor, VS Code with Copilot, and JetBrains IDEs.

Do I need specific dependencies for Research Harness Engineer?

Check the install command and Works With section. Most code skills only require the AI assistant and your codebase.

How do I install Research Harness Engineer?

Copy the install command from the Terminal tab and run it. The skill downloads to ./skills/research-harness-engineer/SKILL.md, ready to use.

Can I customize this skill for my team?

Absolutely. Edit the SKILL.md file to add team-specific instructions, examples, or workflows.

⚠️ Common Mistakes to Avoid

Skipping usability testing

AI-generated designs should be validated with real users before development.

Skipping validation

Always test AI-generated code changes, even for simple refactors.

Missing dependency updates

Check if the skill requires updated dependencies or new packages.

🔗 Related Skills