Personal
~/.claude/skills/Project
.claude/skills/One-line install
npx skills add owner/repo/skill --agent claude-codeSkills for Claude Code
Free agent skills compatible with Claude Code by Anthropic. Search within them or filter by category, level and popularity.
41 skills
Arize Evaluator
github
Streamline LLM evaluation workflows on Arize.
Troubleshoot
microsoft
Analyze logs to understand chat agent behavior.
Agentic Evaluation
github
Enhance AI outputs through iterative evaluation and refinement.
RAG Evaluation
nvidia
Evaluate retrieval-augmented generation benchmarks efficiently.
NV-Reason-CXR
nvidia
Run smoke tests for chest X-ray reasoning models.
Clinical ASR Evaluation
nvidia
Score and evaluate clinical ASR manifests effectively.
NeMo Evaluator SDK
davila7
Enterprise-grade LLM benchmarking across 100+ tasks.
LLM Evaluation Harness
davila7
Benchmark LLMs across 60+ academic tasks.
Advanced Evaluation
muratcankoylan
Evaluate LLM outputs with precision and reliability.
Evals
danielmiessler
An assertion-first framework for AI evaluations.
DeepEval
confident-ai
Streamline evaluation workflows for AI applications.
Fine-Tuning Method Selection
wshobson
Streamline your fine-tuning decisions effectively.
Eval Harness First
wshobson
Essential setup for fine-tuning AI models.
Checkpoint Promotion
wshobson
Ensure your AI model meets quality standards before deployment.
Hugging Face Community Evals
huggingface
Evaluate Hugging Face models locally with ease.
Agent Platform Eval Flywheel
Evaluate and enhance AI models on Google Cloud.
Phoenix Observability
davila7
Open-source observability for LLM applications.
Agent Evaluation
muratcankoylan
Systematically assess agent performance and quality.
AI Workflow Diagnostics
github
Systematically audit your AI workflows for quality and reliability.
i4h Workflow Validate
nvidia
Validate and evaluate i4h environments efficiently.
Context Degradation Patterns
sickn33
Understand and mitigate context degradation in language models.
Agent Evaluation
sickn33
Benchmark and assess LLM agents effectively.
Advanced Evaluation
sickn33
Evaluate LLM outputs reliably using advanced techniques.
Ollama CLI Interface
hkuds
Manage models and generate text from the command line.
