New to Claude Skills? Learn how to install them →

Personal

~/.claude/skills/

Project

.claude/skills/

One-line install

npx skills add owner/repo/skill --agent claude-code
Docs

Skills for Claude Code

Free agent skills compatible with Claude Code by Anthropic. Search within them or filter by category, level and popularity.

A

Arize Evaluator

github

Streamline LLM evaluation workflows on Arize.

AI & AgentsintermediateShell37.7k repo

Troubleshoot

microsoft

Analyze logs to understand chat agent behavior.

AI & Agentsintermediate188.6k repo
A

Agentic Evaluation

github

Enhance AI outputs through iterative evaluation and refinement.

AI & Agentsintermediate37.7k repo

NeMo Evaluator SDK

davila7

Enterprise-grade LLM benchmarking across 100+ tasks.

AI & AgentsintermediatePython · Shell30.2k repo

LLM Evaluation Harness

davila7

Benchmark LLMs across 60+ academic tasks.

AI & AgentsintermediatePython · Shell30.2k repo

Advanced Evaluation

muratcankoylan

Evaluate LLM outputs with precision and reliability.

AI & AgentsintermediatePython17.7k repo

Retune a Corpus

everyinc

Optimize your skill corpus for improved model performance.

AI & AgentsadvancedNode · Shell24.2k repo

Evals

danielmiessler

An assertion-first framework for AI evaluations.

AI & AgentsintermediateNode · Shell18k repo

Fine-Tuning Method Selection

wshobson

Streamline your fine-tuning decisions effectively.

AI & Agentsintermediate38.7k repo
E

Eval Harness First

wshobson

Essential setup for fine-tuning AI models.

AI & Agentsintermediate38.7k repo

Checkpoint Promotion

wshobson

Ensure your AI model meets quality standards before deployment.

AI & Agentsintermediate38.7k repo
O

Obliteratus

nousresearch

Remove refusal behaviors from LLMs without retraining.

AI & AgentsadvancedPython · Shell228.5k repo

Phoenix Observability

davila7

Open-source observability for LLM applications.

AI & AgentsintermediatePython · Shell30.2k repo

Agent Evaluation

muratcankoylan

Systematically assess agent performance and quality.

AI & AgentsintermediatePython17.7k repo
A

AI Workflow Diagnostics

github

Systematically audit your AI workflows for quality and reliability.

AI & Agentsintermediate37.7k repo

Context Degradation Patterns

sickn33

Understand and mitigate context degradation in language models.

AI & Agentsintermediate44.7k repo

Agent Evaluation

sickn33

Benchmark and assess LLM agents effectively.

AI & Agentsintermediate44.7k repo

Advanced Evaluation

sickn33

Evaluate LLM outputs reliably using advanced techniques.

AI & Agentsintermediate44.7k repo

LLM Evaluation

wshobson

Systematic evaluation strategies for LLM performance.

AI & Agentsintermediate38.7k repo

Creating Online Evaluations

posthog

Automate scoring of AI responses based on real failures.

AI & Agentsintermediate37.6k repo

LLM Evaluation

davila7

Evaluate LLM performance with systematic strategies.

AI & Agentsintermediate30.2k repo

LLM Eval Harness

daymade

Evaluate LLM endpoints for reliability and performance.

AI & AgentsintermediatePython · Shell1.3k repo

GAIA Submission

ruvnet

Streamline your GAIA benchmark submission process.

AI & AgentsintermediateNode · Shell67.6k repo

Harness Learn

ruvnet

Optimize harness genomes with automated learning cycles.

AI & AgentsintermediateNode · Shell67.6k repo