New to Claude Skills? Learn how to install them →

Evaluation skills

Free agent skills tagged evaluation, ready to install into any SKILL.md-compatible agent.

E

Eval-Driven Development

github

Automate evaluation for Python LLM applications.

Developer ToolsintermediatePython · Shell37.7k repo

PyTDC

k-dense-ai

Access and evaluate therapeutic datasets with PyTDC.

Data & AnalyticsintermediatePython · Shell33.2k repo

PAIDF AnomalyGen

nvidia

Generate and evaluate synthetic anomaly images efficiently.

Developer ToolsintermediateShell2.8k repo

DINO Object Detection

nvidia

Train and deploy 2D object detection models with DINO.

Developer ToolsintermediateShell2.8k repo

RAG Evaluation

nvidia

Evaluate retrieval-augmented generation benchmarks efficiently.

AI & AgentsintermediatePython · Shell2.8k repo

Code Model Evaluation

davila7

Benchmark code generation models across multiple tasks.

Developer ToolsintermediatePython · Shell30.2k repo

Phoenix Observability

davila7

Open-source observability for LLM applications.

AI & AgentsintermediatePython · Shell30.2k repo
N

Nemotron Customize

nvidia

Streamline your model customization workflows with ease.

Developer Toolsintermediate2.8k repo

i4h Workflow Validate

nvidia

Validate and evaluate i4h environments efficiently.

AI & AgentsintermediateShell2.8k repo
A

AI Engineering Toolkit

sickn33

Transform your AI coding assistant into a senior engineering partner.

AI & AgentsintermediateShell44.7k repo

Create Instance AI Workflow Eval

n8n-io

Efficiently author and manage AI workflow evaluations.

AI & AgentsintermediateNode · Shell200.1k repo

Benchmark Harness

mims-harvard

Enhance ToolUniverse tools with systematic evaluation.

Developer ToolsintermediatePython · Shell1.6k repo

Cosmos Policy Evaluation

orchestra-research

Evaluate NVIDIA Cosmos Policy in simulation environments.

Developer ToolsintermediatePython · Shell11.6k repo

Benchmark Sandbox

vercel

Run Vercel plugin evaluations in isolated environments.

Developer ToolsintermediateNode · Shell249 repo
S

Skill Creator

feiskyer

Build and refine AI agent skills efficiently.

AI & AgentsintermediatePython · Shell1.6k repo

VLM Binary Classification Gap Analysis

Extract false-positive and false-negative gaps from VLM binary-classification-question (BCQ, yes/no) predictions. Use when the user asks to "analyze VLM BCQ gaps", "extract VLM false positives and false negatives", or identify failure cases from a predictions JSON for DEFT root-cause analysis on a binary-classification VLM workflow.

N

NeMo AutoModel Recipe Development

Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.

NSFC Proposal Reviewer

当用户明确要求"评审国自然标书"、"模拟专家评审"、"审阅 NSFC 申请书"时使用。模拟领域专家视角对 NSFC 标书进行多维度评审,输出分级问题与可执行修改建议。⚠️ 不适用:用户只是想写/改标书某个章节(应使用 nsfc-*-writer 系列技能)、只是想了解评审标准(应直接回答)、没有明确"评审/审阅"意图。

Skill Quality Reviewer

This skill should be used when the user asks to "analyze skill quality", "evaluate this skill", "review skill quality", "check my skill", or "generate quality report". Evaluates local skills across description quality, content organization, writing style, and structural integrity.

NeMo Evaluator SDK

Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.

LLM Evaluation Harness

Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.

S

Skill Creator

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.