New to Claude Skills? Learn how to install them →

Personal

~/.codex/skills/

Project

.codex/skills/

One-line install

npx skills add owner/repo/skill --agent codex
Docs

Skills for Codex CLI

Free agent skills compatible with Codex CLI by OpenAI. Search within them or filter by category, level and popularity.

RAG Evaluation

nvidia

Evaluate retrieval-augmented generation benchmarks efficiently.

AI & AgentsintermediatePython · Shell2.8k repo

NV-Reason-CXR

nvidia

Run smoke tests for chest X-ray reasoning models.

AI & AgentsintermediatePython · Shell2.8k repo

NeMo Evaluator SDK

davila7

Enterprise-grade LLM benchmarking across 100+ tasks.

AI & AgentsintermediatePython · Shell30.2k repo

LLM Evaluation Harness

davila7

Benchmark LLMs across 60+ academic tasks.

AI & AgentsintermediatePython · Shell30.2k repo

Advanced Evaluation

muratcankoylan

Evaluate LLM outputs with precision and reliability.

AI & AgentsintermediatePython17.7k repo

DeepEval

confident-ai

Streamline evaluation workflows for AI applications.

AI & AgentsintermediatePython · Shell17.5k repo
O

Obliteratus

nousresearch

Remove refusal behaviors from LLMs without retraining.

AI & AgentsadvancedPython · Shell228.5k repo

Hugging Face Community Evals

huggingface

Evaluate Hugging Face models locally with ease.

AI & AgentsintermediatePython · Shell10.9k repo

Agent Platform Eval Flywheel

google

Evaluate and enhance AI models on Google Cloud.

AI & AgentsintermediatePython · Shell17.6k repo

Phoenix Observability

davila7

Open-source observability for LLM applications.

AI & AgentsintermediatePython · Shell30.2k repo

Agent Evaluation

muratcankoylan

Systematically assess agent performance and quality.

AI & AgentsintermediatePython17.7k repo

Ollama CLI Interface

hkuds

Manage models and generate text from the command line.

AI & AgentsintermediatePython · Shell46.9k repo

LLM Eval Harness

daymade

Evaluate LLM endpoints for reliability and performance.

AI & AgentsintermediatePython · Shell1.3k repo

Behavioral X-Ray

sickn33

Probe AI models for hidden behavioral patterns.

AI & AgentsintermediatePython · Shell44.7k repo

AI Ethics Validator

jeremylongshore

Ensure fairness and compliance in AI models and datasets.

AI & AgentsintermediatePython · Shell2.6k repo

Promptfoo Evaluation

daymade

Efficiently evaluate LLM outputs with Promptfoo.

AI & AgentsintermediatePython · Node · Shell1.3k repo
O

OpenClaw Model Switch

daymade

Easily manage OpenClaw model configurations.

AI & AgentsintermediatePython · Shell1.3k repo