New to Claude Skills? Learn how to install them →

Personal

~/.agents/skills/

Project

.roo/skills/

One-line install

npx skills add owner/repo/skill --agent roo
Docs

Skills for Roo Code

Free agent skills compatible with Roo Code by Roo. Search within them or filter by category, level and popularity.

A

Arize Evaluator

github

Streamline LLM evaluation workflows on Arize.

AI & AgentsintermediateShell37.7k repo

RAG Evaluation

nvidia

Evaluate retrieval-augmented generation benchmarks efficiently.

AI & AgentsintermediatePython · Shell2.8k repo

NV-Reason-CXR

nvidia

Run smoke tests for chest X-ray reasoning models.

AI & AgentsintermediatePython · Shell2.8k repo

Clinical ASR Evaluation

nvidia

Score and evaluate clinical ASR manifests effectively.

AI & AgentsintermediateShell2.8k repo

NeMo Evaluator SDK

davila7

Enterprise-grade LLM benchmarking across 100+ tasks.

AI & AgentsintermediatePython · Shell30.2k repo

LLM Evaluation Harness

davila7

Benchmark LLMs across 60+ academic tasks.

AI & AgentsintermediatePython · Shell30.2k repo

Retune a Corpus

everyinc

Optimize your skill corpus for improved model performance.

AI & AgentsadvancedNode · Shell24.2k repo

Evals

danielmiessler

An assertion-first framework for AI evaluations.

AI & AgentsintermediateNode · Shell18k repo

DeepEval

confident-ai

Streamline evaluation workflows for AI applications.

AI & AgentsintermediatePython · Shell17.5k repo
O

Obliteratus

nousresearch

Remove refusal behaviors from LLMs without retraining.

AI & AgentsadvancedPython · Shell228.5k repo

Hugging Face Community Evals

huggingface

Evaluate Hugging Face models locally with ease.

AI & AgentsintermediatePython · Shell10.9k repo

Agent Platform Eval Flywheel

google

Evaluate and enhance AI models on Google Cloud.

AI & AgentsintermediatePython · Shell17.6k repo

Phoenix Observability

davila7

Open-source observability for LLM applications.

AI & AgentsintermediatePython · Shell30.2k repo

i4h Workflow Validate

nvidia

Validate and evaluate i4h environments efficiently.

AI & AgentsintermediateShell2.8k repo

Ollama CLI Interface

hkuds

Manage models and generate text from the command line.

AI & AgentsintermediatePython · Shell46.9k repo

Agent Evaluation CLI

google

Efficiently evaluate and optimize your AI agents.

AI & AgentsintermediateShell5.5k repo

LLM Eval Harness

daymade

Evaluate LLM endpoints for reliability and performance.

AI & AgentsintermediatePython · Shell1.3k repo

GAIA Submission

ruvnet

Streamline your GAIA benchmark submission process.

AI & AgentsintermediateNode · Shell67.6k repo

Harness Learn

ruvnet

Optimize harness genomes with automated learning cycles.

AI & AgentsintermediateNode · Shell67.6k repo

Cost Counterfactual

ruvnet

Analyze routing costs against baseline models.

AI & AgentsintermediateShell67.6k repo

Behavioral X-Ray

sickn33

Probe AI models for hidden behavioral patterns.

AI & AgentsintermediatePython · Shell44.7k repo

AI Ethics Validator

jeremylongshore

Ensure fairness and compliance in AI models and datasets.

AI & AgentsintermediatePython · Shell2.6k repo

Promptfoo Evaluation

daymade

Efficiently evaluate LLM outputs with Promptfoo.

AI & AgentsintermediatePython · Node · Shell1.3k repo
O

OpenClaw Model Switch

daymade

Easily manage OpenClaw model configurations.

AI & AgentsintermediatePython · Shell1.3k repo