New to Claude Skills? Learn how to install them →

k-dense-ai on GitHub

Scanpy

Free

Comprehensive toolkit for single-cell RNA-seq analysis.

Get this skill

Free · Opens the source repo

What Scanpy does

Scanpy is a powerful Python toolkit specifically designed for analyzing single-cell RNA sequencing (scRNA-seq) data. Built on the AnnData structure, it provides a full suite of functionalities for processing, analyzing, and visualizing scRNA-seq datasets. Users can perform essential tasks such as quality control, normalization, dimensionality reduction (including PCA, UMAP, and t-SNE), clustering, and differential expression analysis. The toolkit is particularly useful for researchers and data scientists working in genomics and bioinformatics who need to explore and interpret single-cell data efficiently.

The skill comes with a collection of ready-to-use scripts that streamline common analysis workflows. Each script is designed to handle specific tasks, such as inspecting data, performing quality checks, normalizing datasets, and generating visualizations. This modular approach allows users to either run complete workflows with a single command or execute individual steps as needed, making it flexible for both beginners and advanced users. The scripts automatically manage file formats and logging, ensuring that users can focus on their analysis without worrying about the underlying code.

For those transitioning from R, Scanpy facilitates the conversion of R-native single-cell formats like Seurat and SingleCellExperiment into the h5ad format, which is essential for compatibility with Scanpy. This feature is particularly valuable for researchers who have existing datasets in R and want to leverage Python's ecosystem for further analysis. Overall, Scanpy is an essential tool for anyone involved in single-cell genomics, providing a comprehensive solution for data analysis and visualization.

When to use it

Use this skill when you need to analyze single-cell RNA-seq data, perform quality control, or visualize high-dimensional data.

When not to use it

This skill may not be suitable for users looking for deep learning capabilities beyond the provided workflows or those who require extensive customization not covered by the available scripts.

What you can build with it

Quality Control of scRNA-seq Data

Run the `qc_analysis.py` script to perform quality checks on your raw scRNA-seq data, ensuring that it meets the necessary thresholds before analysis.

Visualizing Clusters in scRNA-seq Data

Use the `plot.py` script to generate UMAP or t-SNE visualizations of your clustered data, helping to interpret the results effectively.

Batch Correction for Integrated Datasets

Apply the `batch_correct.py` script to integrate multiple samples while correcting for batch effects, ensuring accurate downstream analysis.

How to install Scanpy

View source

1. Install with the skills CLI

npx skills add k-dense-ai/scientific-agent-skills/scanpy --agent claude-code

2. Or install it manually

Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.

Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs

Inside SKILL.md

Written by k-dense-ai

Scanpy: Single-Cell Analysis

Overview

Scanpy is a scalable Python toolkit for analyzing single-cell RNA-seq data, built on AnnData. Apply this skill for complete single-cell workflows including quality control, normalization, dimensionality reduction, clustering, marker gene identification, visualization, and trajectory analysis. Current stable release: scanpy 1.12.x (January 2026).

Installation

Requires Python 3.12+ (scanpy 1.12 dropped Python ≤3.11) and anndata ≥0.10.

uv pip install "scanpy[leiden]"

The [leiden] extra installs python-igraph and leidenalg, required for Leiden clustering. For reproducible environments, pin a version: uv pip install "scanpy[leiden]==1.12.1".

For large or out-of-core datasets, many functions support Dask arrays (experimental):

uv pip install "scanpy[leiden]" dask

See the Using dask with Scanpy tutorial. For GPU-accelerated scanpy-like operations, use rapids-singlecell as a separate package.

If the input is an R-native single-cell object (.rds, .RData, Seurat, or SingleCellExperiment), first convert it to .h5ad with R tooling, then load it with Scanpy. Read references/r_interop.md for agent-run installation and conversion instructions across macOS, Linux, and Windows.

For AnnData structure and I/O details, use the anndata skill. For probabilistic models and batch correction, use scvi-tools.

When to Use This Skill

This skill should be used when:

  • Analyzing single-cell RNA-seq data (.h5ad, 10X, CSV formats)
  • Working with R-friendly single-cell datasets (.rds, .RData, Seurat, SingleCellExperiment) that need conversion to .h5ad
  • Performing quality control on scRNA-seq datasets
  • Creating UMAP, t-SNE, or PCA visualizations
  • Identifying cell clusters and finding marker genes
  • Annotating cell types based on gene expression
  • Conducting trajectory inference or pseudotime analysis
  • Generating publication-quality single-cell plots

Script Toolkit (prefer these over writing code from scratch)

This skill bundles ready-to-run CLI scripts in scripts/ for every common step. Run these instead of hand-writing scanpy code — they handle file loading by extension, figure setup, sensible defaults, raw-count preservation, and progress logging. Each reads and writes .h5ad, so they chain together, and each has its own --help. Only drop down to writing scanpy code when a task isn't covered by a script or needs unusual customization.

All scripts use a shared scripts/_common.py helper (loading, saving, figure config) — keep it alongside the others. Run from the skill directory or pass full paths; figures default to ./figures/.

ScriptPurposeTypical call
run_pipeline.pyFull workflow in one command: load → QC → normalize → HVG → PCA → (batch) → UMAP → Leiden → markerspython scripts/run_pipeline.py raw.h5ad -o processed.h5ad
inspect_data.pySummarize an unknown dataset (shape, obs/var, layers, what's already computed, raw vs normalized)python scripts/inspect_data.py data.h5ad
convert.pyLoad any format (10x dir/.h5, csv, loom, mtx) and write .h5adpython scripts/convert.py 10x_dir/ -o data.h5ad
qc_analysis.pyQC metrics, before/after plots, filtering, optional Scrublet doubletspython scripts/qc_analysis.py raw.h5ad -o qc.h5ad --scrublet
preprocess.pyNormalize, log1p, HVG, optional scale/regress (keeps counts layer + raw)python scripts/preprocess.py qc.h5ad -o norm.h5ad
reduce_dimensions.pyPCA + variance plot, neighbors, UMAP, optional t-SNEpython scripts/reduce_dimensions.py norm.h5ad -o red.h5ad
batch_correct.pyIntegration: harmony / bbknn / combatpython scripts/batch_correct.py red.h5ad -o int.h5ad --method harmony --batch-key sample
cluster.pyLeiden (or louvain) at one or many resolutionspython scripts/cluster.py red.h5ad -o clu.h5ad --resolution 0.3 0.6 1.0
find_markers.pyrank_genes_groups + per-group CSVs + marker plotspython scripts/find_markers.py clu.h5ad --groupby leiden -o clu.h5ad
annotate.pyMap clusters → cell types from JSON/CSV; optional marker reference dotplotpython scripts/annotate.py clu.h5ad -o ann.h5ad --mapping map.json
score_genes.pyScore gene signatures (JSON) and/or cell-cycle phasepython scripts/score_genes.py ann.h5ad -o scored.h5ad --gene-sets sigs.json
pseudobulk.pyAggregate counts by sample × cell type → matrix for pydeseq2python scripts/pseudobulk.py ann.h5ad --by sample cell_type --out-prefix pb
subset.pySubset by obs values or gene list (optionally clear stale embeddings)python scripts/subset.py ann.h5ad -o tcells.h5ad --obs cell_type --keep "T cells"
plot.pyGenerate umap/tsne/pca/violin/dotplot/heatmap/etc. from a processed objectpython scripts/plot.py ann.h5ad --kind dotplot --genes CD3D CD14 --groupby cell_type

One-shot end-to-end run

# Counts → clustered, marker-annotated object + figures + marker CSVs
python scripts/run_pipeline.py raw.h5ad -o processed.h5ad \
    --resolution 0.5 --n-top-genes 2000 --scrublet
# With multi-sample integration:
python scripts/run_pipeline.py raw.h5ad -o processed.h5ad --batch-key sample --batch-method harmony
# Reproducible parameters via JSON (keys mirror flag names with underscores):
python scripts/run_pipeline.py raw.h5ad -o processed.h5ad --config params.json

Step-by-step chain (when you need to inspect/iterate between stages)

python scripts/qc_analysis.py        raw.h5ad  -o qc.h5ad   --scrublet
python scripts/preprocess.py         qc.h5ad   -o norm.h5ad --n-top-genes 2000
python scripts/reduce_dimensions.py  norm.h5ad -o red.h5ad  --n-pcs 40
python scripts/cluster.py            red.h5ad  -o clu.h5ad  --resolution 0.3 0.5 0.8
python scripts/find_markers.py       clu.h5ad  -o clu.h5ad  --groupby leiden --use-raw
# inspect results/markers/*.csv, decide labels, write a mapping JSON, then:
python scripts/annotate.py           clu.h5ad  -o ann.h5ad  --mapping celltypes.json

The sections below document the underlying scanpy calls each script performs — read them when customizing beyond the script flags.

Quick Start

Basic Import and Setup

import scanpy as sc
import pandas as pd
import numpy as np

# Configure settings
sc.settings.verbosity = 3
sc.settings.set_figure_params(dpi=80, facecolor='white')
sc.settings.figdir = './figures/'
sc.settings.autosave = True  # Preferred over per-plot save= (deprecated in scanpy 1.12)

Loading Data

# From 10X Genomics
adata = sc.read_10x_mtx('path/to/data/')
adata = sc.read_10x_h5('path/to/data.h5')

# From h5ad (AnnData format)
adata = sc.read_h5ad('path/to/data.h5ad')

# From CSV
adata = sc.read_csv('path/to/data.csv')

For R-native files, do not try to parse Seurat .rds directly in Python. Convert first:

# See references/r_interop.md for installing R and conversion packages.
Rscript convert_rds_to_h5ad.R input.rds output.h5ad
adata = sc.read_h5ad('output.h5ad')

Understanding AnnData Structure

The AnnData object is the core data structure in scanpy:

adata.X          # Expression matrix (cells × genes)
adata.obs        # Cell metadata (DataFrame)
adata.var        # Gene metadata (DataFrame)
adata.uns        # Unstructured annotations (dict)
adata.obsm       # Multi-dimensional cell data (PCA, UMAP)
adata.raw        # Raw data backup

# Access cell and gene names
adata.obs_names  # Cell barcodes
adata.var_names  # Gene names

Standard Analysis Workflow

The seven steps, with code and the parameters that matter at each, are in references/analysis_workflow.md:

  1. Quality control — filter cells and genes; inspect mitochondrial fraction and counts before choosing thresholds rather than copying defaults.
  2. Normalization and preprocessing — normalize, log-transform, select highly variable genes, and keep .raw for later plotting.
  3. Dimensionality reduction — PCA, then the neighbour graph, then UMAP.
  4. Clustering — Leiden at a resolution chosen for the question, not the default.
  5. Marker gene identification — ranked genes per cluster.
  6. Cell type annotation — mapping clusters to types from markers.
  7. Save results — writing the annotated AnnData.

Common follow-on tasks — publication plots, trajectory inference, pseudobulk differential expression between conditions, gene set scoring, and batch correction — are in the same file. See also references/standard_workflow.md and references/plotting_guide.md.

Key Parameters to Adjust

Quality Control

  • min_genes: Minimum genes per cell (typically 200-500)
  • min_cells: Minimum cells per gene (typically 3-10)
  • pct_counts_mt: Mitochondrial threshold (typically 5-20%)

Normalization

  • target_sum: Target counts per cell (default 1e4)

Feature Selection

  • n_top_genes: Number of HVGs (typically 2000-3000)
  • min_mean, max_mean, min_disp: HVG selection parameters

Dimensionality Reduction

  • n_pcs: Number of principal components (check variance ratio plot)
  • n_neighbors: Number of neighbors (typically 10-30)

Clustering

  • resolution: Clustering granularity (0.4-1.2, higher = more clusters)

Common Pitfalls and Best Practices

  1. Always save raw counts: adata.raw = adata before filtering genes
  2. Check QC plots carefully: Adjust thresholds based on dataset quality
  3. Use Leiden clustering: sc.tl.louvain is deprecated in scanpy 1.12
  4. Try multiple clustering resolutions: Find optimal granularity
  5. Validate cell type annotations: Use multiple marker genes
  6. Use use_raw=True for gene expression plots: Shows normalized counts from .raw
  7. Check PCA variance ratio: Determine optimal number of PCs
  8. Save intermediate results: Long workflows can fail partway through
  9. Pseudobulk for DE: Do not treat rank_genes_groups p-values as rigorous DE between conditions
  10. Save plots via settings: Use sc.settings.autosave instead of deprecated save= on plot functions
  11. Convert R objects before Scanpy: Use R packages to convert Seurat or SingleCellExperiment .rds files to .h5ad, preserving counts, metadata, and gene identifiers

Bundled Resources

scripts/ (CLI toolkit)

A composable set of .h5ad-in/.h5ad-out scripts covering the whole workflow plus a one-command end-to-end pipeline. See the Script Toolkit section above for the full table and chaining examples. Each script has --help. Files:

  • _common.py — shared loading/saving/figure helpers imported by the others (not a CLI)
  • run_pipeline.py — full pipeline in one command (flags or --config JSON)
  • inspect_data.py, convert.py — explore and load/convert any input format
  • qc_analysis.py, preprocess.py, reduce_dimensions.py, batch_correct.py, cluster.py — pipeline steps
  • find_markers.py, annotate.py, score_genes.py, pseudobulk.py — markers, annotation, scoring, DE prep
  • subset.py, plot.py — subset by metadata/genes; generate any standard plot

Default to these scripts before writing scanpy code from scratch.

references/standard_workflow.md

Complete step-by-step workflow with detailed explanations and code examples for:

  • Data loading and setup
  • Quality control with visualization
  • Normalization and scaling
  • Feature selection
  • Dimensionality reduction (PCA, UMAP, t-SNE)
  • Clustering (Leiden)
  • Doublet detection (scrublet) and pseudobulk aggregation
  • Marker gene identification
  • Cell type annotation
  • Trajectory inference
  • Differential expression

Read this reference when performing a complete analysis from scratch.

references/api_reference.md

Quick reference guide for scanpy functions organized by module:

  • Reading/writing data (sc.read_*, adata.write_*)
  • Preprocessing (sc.pp.*)
  • Tools (sc.tl.*)
  • Plotting (sc.pl.*)
  • AnnData structure and manipulation
  • Settings and utilities

Use this for quick lookup of function signatures and common parameters.

references/plotting_guide.md

Comprehensive visualization guide including:

  • Quality control plots
  • Dimensionality reduction visualizations
  • Clustering visualizations
  • Marker gene plots (heatmaps, dot plots, violin plots)
  • Trajectory and pseudotime plots
  • Publication-quality customization
  • Multi-panel figures
  • Color palettes and styling

Consult this when creating publication-ready figures.

references/r_interop.md

Agent runbook for installing R on macOS, Linux, and Windows, installing CRAN/Bioconductor conversion packages, inspecting .rds/.RData inputs, converting Seurat or SingleCellExperiment objects to .h5ad, and validating the result in Scanpy.

assets/analysis_template.py

Complete analysis template providing a full workflow from data loading through cell type annotation. Copy and customize this template for new analyses:

cp assets/analysis_template.py my_analysis.py
# Edit parameters and run
python my_analysis.py

The template includes all standard steps with configurable parameters and helpful comments.

assets/ JSON templates

Edit-and-pass templates so you don't author config/mappings from scratch:

  • assets/pipeline_config.json — parameter set for run_pipeline.py --config
  • assets/celltype_mapping.json — cluster → cell-type map for annotate.py --mapping
  • assets/gene_signatures.json — gene-set signatures for score_genes.py --gene-sets

Additional Resources

Tips for Effective Analysis

  1. Start with the template: Use assets/analysis_template.py as a starting point
  2. Run QC script first: Use scripts/qc_analysis.py for initial filtering
  3. Consult references as needed: Load workflow and API references into context
  4. Iterate on clustering: Try multiple resolutions and visualization methods
  5. Validate biologically: Check marker genes match expected cell types
  6. Document parameters: Record QC thresholds and analysis settings
  7. Save checkpoints: Write intermediate results at key steps

Frequently asked questions about Scanpy

Similar skills