New to Claude Skills? Learn how to install them →

davila7 on GitHub

Senior ML/AI Engineer

Free

Expertly productionize AI and ML models with advanced MLOps.

Get this skill

Free · Opens the source repo

What Senior ML/AI Engineer does

The Senior ML/AI Engineer skill provides a comprehensive toolkit for professionals focused on deploying machine learning models and building scalable ML systems. This skill is tailored for engineers and data scientists who require robust solutions for productionizing AI applications, implementing MLOps practices, and integrating large language models (LLMs) into their workflows. With a focus on advanced production patterns, the skill equips users with the necessary scripts and documentation to streamline their ML operations.

Included in this skill are several core scripts that facilitate essential tasks in ML engineering. The model_deployment_pipeline.py script helps automate the deployment of models, ensuring that they are efficiently integrated into production environments. The rag_system_builder.py allows users to construct retrieval-augmented generation systems, which are critical for enhancing the capabilities of LLMs. Additionally, the ml_monitoring_suite.py provides tools for monitoring deployed models, ensuring they perform optimally over time and adhere to compliance standards.

The skill is supported by extensive reference documentation covering MLOps production patterns, LLM integration, and RAG system architecture. This documentation offers practical guidance on best practices, system design, and performance optimization, making it a valuable resource for both novice and experienced ML engineers. By leveraging this skill, teams can improve their workflows, enhance collaboration, and maintain high standards of quality and performance in their ML projects.

When to use it

Use this skill when you need to implement MLOps practices, deploy ML models, or integrate LLMs into production systems.

When not to use it

This skill may not be suitable for simple ML projects or for users without a background in ML engineering or data science.

What you can build with it

Deploying a New ML Model

Utilize the model deployment pipeline script to automate the deployment of a new machine learning model into a production environment.

Building a RAG System

Leverage the rag system builder script to create a retrieval-augmented generation system that enhances the capabilities of your LLM.

Monitoring Model Performance

Implement the ml monitoring suite to continuously monitor the performance of your deployed models and ensure compliance with performance targets.

How to install Senior ML/AI Engineer

View source

1. Install with the skills CLI

npx skills add davila7/claude-code-templates/senior-ml-engineer --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by davila7

Senior ML/AI Engineer

World-class senior ml/ai engineer skill for production-grade AI/ML/Data systems.

Quick Start

Main Capabilities

# Core Tool 1
python scripts/model_deployment_pipeline.py --input data/ --output results/

# Core Tool 2  
python scripts/rag_system_builder.py --target project/ --analyze

# Core Tool 3
python scripts/ml_monitoring_suite.py --config config.yaml --deploy

Core Expertise

This skill covers world-class capabilities in:

  • Advanced production patterns and architectures
  • Scalable system design and implementation
  • Performance optimization at scale
  • MLOps and DataOps best practices
  • Real-time processing and inference
  • Distributed computing frameworks
  • Model deployment and monitoring
  • Security and compliance
  • Cost optimization
  • Team leadership and mentoring

Tech Stack

Languages: Python, SQL, R, Scala, Go ML Frameworks: PyTorch, TensorFlow, Scikit-learn, XGBoost Data Tools: Spark, Airflow, dbt, Kafka, Databricks LLM Frameworks: LangChain, LlamaIndex, DSPy Deployment: Docker, Kubernetes, AWS/GCP/Azure Monitoring: MLflow, Weights & Biases, Prometheus Databases: PostgreSQL, BigQuery, Snowflake, Pinecone

Reference Documentation

1. Mlops Production Patterns

Comprehensive guide available in references/mlops_production_patterns.md covering:

  • Advanced patterns and best practices
  • Production implementation strategies
  • Performance optimization techniques
  • Scalability considerations
  • Security and compliance
  • Real-world case studies

2. Llm Integration Guide

Complete workflow documentation in references/llm_integration_guide.md including:

  • Step-by-step processes
  • Architecture design patterns
  • Tool integration guides
  • Performance tuning strategies
  • Troubleshooting procedures

3. Rag System Architecture

Technical reference guide in references/rag_system_architecture.md with:

  • System design principles
  • Implementation examples
  • Configuration best practices
  • Deployment strategies
  • Monitoring and observability

Production Patterns

Pattern 1: Scalable Data Processing

Enterprise-scale data processing with distributed computing:

  • Horizontal scaling architecture
  • Fault-tolerant design
  • Real-time and batch processing
  • Data quality validation
  • Performance monitoring

Pattern 2: ML Model Deployment

Production ML system with high availability:

  • Model serving with low latency
  • A/B testing infrastructure
  • Feature store integration
  • Model monitoring and drift detection
  • Automated retraining pipelines

Pattern 3: Real-Time Inference

High-throughput inference system:

  • Batching and caching strategies
  • Load balancing
  • Auto-scaling
  • Latency optimization
  • Cost optimization

Best Practices

Development

  • Test-driven development
  • Code reviews and pair programming
  • Documentation as code
  • Version control everything
  • Continuous integration

Production

  • Monitor everything critical
  • Automate deployments
  • Feature flags for releases
  • Canary deployments
  • Comprehensive logging

Team Leadership

  • Mentor junior engineers
  • Drive technical decisions
  • Establish coding standards
  • Foster learning culture
  • Cross-functional collaboration

Performance Targets

Latency:

  • P50: < 50ms
  • P95: < 100ms
  • P99: < 200ms

Throughput:

  • Requests/second: > 1000
  • Concurrent users: > 10,000

Availability:

  • Uptime: 99.9%
  • Error rate: < 0.1%

Security & Compliance

  • Authentication & authorization
  • Data encryption (at rest & in transit)
  • PII handling and anonymization
  • GDPR/CCPA compliance
  • Regular security audits
  • Vulnerability management

Common Commands

# Development
python -m pytest tests/ -v --cov
python -m black src/
python -m pylint src/

# Training
python scripts/train.py --config prod.yaml
python scripts/evaluate.py --model best.pth

# Deployment
docker build -t service:v1 .
kubectl apply -f k8s/
helm upgrade service ./charts/

# Monitoring
kubectl logs -f deployment/service
python scripts/health_check.py

Resources

  • Advanced Patterns: references/mlops_production_patterns.md
  • Implementation Guide: references/llm_integration_guide.md
  • Technical Reference: references/rag_system_architecture.md
  • Automation Scripts: scripts/ directory

Senior-Level Responsibilities

As a world-class senior professional:

  1. Technical Leadership

    • Drive architectural decisions
    • Mentor team members
    • Establish best practices
    • Ensure code quality
  2. Strategic Thinking

    • Align with business goals
    • Evaluate trade-offs
    • Plan for scale
    • Manage technical debt
  3. Collaboration

    • Work across teams
    • Communicate effectively
    • Build consensus
    • Share knowledge
  4. Innovation

    • Stay current with research
    • Experiment with new approaches
    • Contribute to community
    • Drive continuous improvement
  5. Production Excellence

    • Ensure high availability
    • Monitor proactively
    • Optimize performance
    • Respond to incidents

Frequently asked questions about Senior ML/AI Engineer

Similar skills