Model Evaluation skills
Free agent skills tagged model evaluation, ready to install into any SKILL.md-compatible agent.
17 skills
Arize Experiment
github
Evaluate and compare model performance with ease.
RAG Evaluation
nvidia
Evaluate retrieval-augmented generation benchmarks efficiently.
Scikit-learn
k-dense-ai
Comprehensive machine learning toolkit for Python.
NV-Reason-CXR
nvidia
Run smoke tests for chest X-ray reasoning models.
Senior Data Scientist
davila7
Expert tools for advanced data science and analytics.
Senior Data Scientist
alirezarezvani
Expert tools for statistical modeling and experiment design.
NeMo Evaluator Plugin
nvidia
Evaluate models and datasets efficiently with NeMo.
Retune a Corpus
everyinc
Optimize your skill corpus for improved model performance.
LLM Benchmarking
nousresearch
Evaluate and compare language models with standardized metrics.
Agent Platform Eval Flywheel
Evaluate and enhance AI models on Google Cloud.
Hugging Face Community Evals
huggingface
Evaluate Hugging Face models locally with ease.
Obliteratus
nousresearch
Remove refusal behaviors from LLMs without retraining.
Checkpoint Promotion
wshobson
Ensure your AI model meets quality standards before deployment.
Fine-Tuning on Microsoft Foundry
microsoft
Efficiently fine-tune models with SFT, DPO, or RFT.
Behavioral X-Ray
sickn33
Probe AI models for hidden behavioral patterns.
Scikit-learn
zlanqing
Comprehensive machine learning toolkit for Python.
AI Observability Evaluations
posthog
Streamline AI evaluation management and insights.
