Evaluation skills
Free agent skills tagged evaluation, ready to install into any SKILL.md-compatible agent.
22 skills
Eval-Driven Development
github
Automate evaluation for Python LLM applications.
PyTDC
k-dense-ai
Access and evaluate therapeutic datasets with PyTDC.
PAIDF AnomalyGen
nvidia
Generate and evaluate synthetic anomaly images efficiently.
DINO Object Detection
nvidia
Train and deploy 2D object detection models with DINO.
RAG Evaluation
nvidia
Evaluate retrieval-augmented generation benchmarks efficiently.
Code Model Evaluation
davila7
Benchmark code generation models across multiple tasks.
Phoenix Observability
davila7
Open-source observability for LLM applications.
Nemotron Customize
nvidia
Streamline your model customization workflows with ease.
i4h Workflow Validate
nvidia
Validate and evaluate i4h environments efficiently.
AI Engineering Toolkit
sickn33
Transform your AI coding assistant into a senior engineering partner.
Create Instance AI Workflow Eval
n8n-io
Efficiently author and manage AI workflow evaluations.
Benchmark Harness
mims-harvard
Enhance ToolUniverse tools with systematic evaluation.
Cosmos Policy Evaluation
orchestra-research
Evaluate NVIDIA Cosmos Policy in simulation environments.
Benchmark Sandbox
vercel
Run Vercel plugin evaluations in isolated environments.
Skill Creator
feiskyer
Build and refine AI agent skills efficiently.
VLM Binary Classification Gap Analysis
Extract false-positive and false-negative gaps from VLM binary-classification-question (BCQ, yes/no) predictions. Use when the user asks to "analyze VLM BCQ gaps", "extract VLM false positives and false negatives", or identify failure cases from a predictions JSON for DEFT root-cause analysis on a binary-classification VLM workflow.
NeMo AutoModel Recipe Development
Create and modify NeMo AutoModel training and evaluation recipes, including YAML structure, builders, and execution flow.
NSFC Proposal Reviewer
当用户明确要求"评审国自然标书"、"模拟专家评审"、"审阅 NSFC 申请书"时使用。模拟领域专家视角对 NSFC 标书进行多维度评审,输出分级问题与可执行修改建议。⚠️ 不适用:用户只是想写/改标书某个章节(应使用 nsfc-*-writer 系列技能)、只是想了解评审标准(应直接回答)、没有明确"评审/审阅"意图。
Skill Quality Reviewer
This skill should be used when the user asks to "analyze skill quality", "evaluate this skill", "review skill quality", "check my skill", or "generate quality report". Evaluates local skills across description quality, content organization, writing style, and structural integrity.
NeMo Evaluator SDK
Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. Use when needing scalable evaluation on local Docker, Slurm HPC, or cloud platforms. NVIDIA's enterprise-grade platform with container-first architecture for reproducible benchmarking.
LLM Evaluation Harness
Evaluates LLMs across 60+ academic benchmarks (MMLU, HumanEval, GSM8K, TruthfulQA, HellaSwag). Use when benchmarking model quality, comparing models, reporting academic results, or tracking training progress. Industry standard used by EleutherAI, HuggingFace, and major labs. Supports HuggingFace, vLLM, APIs.
Skill Creator
Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.
