New to Claude Skills? Learn how to install them →

Personal

~/.claude/skills/

Project

.claude/skills/

One-line install

npx skills add owner/repo/skill --agent claude-code
Docs

Skills for Claude Code

Free agent skills compatible with Claude Code by Anthropic. Search within them or filter by category, level and popularity.

A

Arize Evaluator

github

Streamline LLM evaluation workflows on Arize.

AI & AgentsintermediateShell37.7k repo

Troubleshoot

microsoft

Analyze logs to understand chat agent behavior.

AI & Agentsintermediate188.6k repo
A

Agentic Evaluation

github

Enhance AI outputs through iterative evaluation and refinement.

AI & Agentsintermediate37.7k repo

RAG Evaluation

nvidia

Evaluate retrieval-augmented generation benchmarks efficiently.

AI & AgentsintermediatePython · Shell2.8k repo

NV-Reason-CXR

nvidia

Run smoke tests for chest X-ray reasoning models.

AI & AgentsintermediatePython · Shell2.8k repo

Clinical ASR Evaluation

nvidia

Score and evaluate clinical ASR manifests effectively.

AI & AgentsintermediateShell2.8k repo

NeMo Evaluator SDK

davila7

Enterprise-grade LLM benchmarking across 100+ tasks.

AI & AgentsintermediatePython · Shell30.2k repo

LLM Evaluation Harness

davila7

Benchmark LLMs across 60+ academic tasks.

AI & AgentsintermediatePython · Shell30.2k repo

Advanced Evaluation

muratcankoylan

Evaluate LLM outputs with precision and reliability.

AI & AgentsintermediatePython17.7k repo

Evals

danielmiessler

An assertion-first framework for AI evaluations.

AI & AgentsintermediateNode · Shell18k repo

DeepEval

confident-ai

Streamline evaluation workflows for AI applications.

AI & AgentsintermediatePython · Shell17.5k repo

Fine-Tuning Method Selection

wshobson

Streamline your fine-tuning decisions effectively.

AI & Agentsintermediate38.7k repo
E

Eval Harness First

wshobson

Essential setup for fine-tuning AI models.

AI & Agentsintermediate38.7k repo

Checkpoint Promotion

wshobson

Ensure your AI model meets quality standards before deployment.

AI & Agentsintermediate38.7k repo

Hugging Face Community Evals

huggingface

Evaluate Hugging Face models locally with ease.

AI & AgentsintermediatePython · Shell10.9k repo

Agent Platform Eval Flywheel

google

Evaluate and enhance AI models on Google Cloud.

AI & AgentsintermediatePython · Shell17.6k repo

Phoenix Observability

davila7

Open-source observability for LLM applications.

AI & AgentsintermediatePython · Shell30.2k repo

Agent Evaluation

muratcankoylan

Systematically assess agent performance and quality.

AI & AgentsintermediatePython17.7k repo
A

AI Workflow Diagnostics

github

Systematically audit your AI workflows for quality and reliability.

AI & Agentsintermediate37.7k repo

i4h Workflow Validate

nvidia

Validate and evaluate i4h environments efficiently.

AI & AgentsintermediateShell2.8k repo

Context Degradation Patterns

sickn33

Understand and mitigate context degradation in language models.

AI & Agentsintermediate44.7k repo

Agent Evaluation

sickn33

Benchmark and assess LLM agents effectively.

AI & Agentsintermediate44.7k repo

Advanced Evaluation

sickn33

Evaluate LLM outputs reliably using advanced techniques.

AI & Agentsintermediate44.7k repo

Ollama CLI Interface

hkuds

Manage models and generate text from the command line.

AI & AgentsintermediatePython · Shell46.9k repo