
Zarr Python
FreeEfficiently store and manage large N-dimensional arrays.
Free · Opens the source repo
What Zarr Python does
Zarr Python is a library designed for the efficient storage and manipulation of large N-dimensional arrays, making it particularly useful for scientific computing applications. By utilizing chunking and compression, Zarr allows users to handle massive datasets seamlessly, whether they are stored locally or in cloud environments such as S3 or Google Cloud Storage. This skill integrates smoothly with popular Python libraries like NumPy, Dask, and Xarray, enabling users to leverage existing workflows while enhancing their data management capabilities.
The library supports parallel I/O operations, which can significantly improve performance when working with large datasets. Users can create arrays with various shapes and chunk sizes, allowing for optimized access patterns tailored to their specific use cases. Zarr also allows for the organization of data into hierarchical groups, similar to directory structures, which can simplify the management of related datasets. Additionally, users can attach metadata to arrays and groups, providing context and documentation directly within their data structures.
Zarr Python is particularly suited for researchers and data scientists who require a robust solution for handling large-scale data in a cloud-native manner. Its compatibility with other scientific libraries means that it can fit into existing data processing pipelines with minimal friction. Whether you are conducting climate modeling, analyzing large datasets from experiments, or working with time-series data, Zarr Python provides the tools necessary to manage and manipulate your data efficiently.
When to use it
Use this skill when you need to manage large datasets that require efficient storage, retrieval, and manipulation, especially in scientific computing contexts.
When not to use it
This skill may not be suitable for small datasets or applications that do not require advanced data management features like chunking or compression.
What you can build with it
Climate Data Analysis
Researchers can use Zarr Python to store and analyze large climate datasets, leveraging chunking for efficient data access.
Large-Scale Scientific Simulations
Zarr Python enables scientists to manage the output of large simulations, providing a structured way to store and retrieve results.
Time-Series Data Management
Data scientists can utilize Zarr to handle time-series data, allowing for efficient appending and resizing of datasets.
How to install Zarr Python
View source1. Install with the skills CLI
npx skills add k-dense-ai/scientific-agent-skills/zarr-python --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by k-dense-aiZarr Python
Overview
Zarr is a Python library for storing large N-dimensional arrays with chunking and compression. Apply this skill for efficient parallel I/O, cloud-native workflows, and seamless integration with NumPy, Dask, and Xarray.
Current upstream: zarr 3.2.1 (released 2026-05-05). Docs: zarr.readthedocs.io. New arrays default to Zarr format 3; set zarr_format=2 for legacy interop. Zarr 3.2 adds rectilinear chunks and continues to refine the v3 codec pipeline. This skill is a community guide maintained by K-Dense Inc., not an official zarr-developers package.
Quick Start
Installation
uv pip install "zarr==3.2.1"
Requires Python 3.12+ and NumPy 2.0+ for current stable Zarr-Python. For remote stores (S3, GCS, HTTP), pin the optional extras/backends in your project lockfile:
uv pip install "zarr[remote]==3.2.1" "s3fs==2026.4.0" "gcsfs==2026.5.0"
Use a version range such as zarr>=3,<4 only when your project has a committed lockfile and compatibility tests. For Zarr-Python 2 / Python 3.10–3.11 workflows, choose an exact zarr==2.x.y patch version from the support-v2 release notes and commit the resulting lockfile.
Basic Array Creation
import zarr
import numpy as np
# Create a 2D array with chunking and compression
z = zarr.create_array(
store="data/my_array.zarr",
shape=(10000, 10000),
chunks=(1000, 1000),
dtype="f4"
)
# Write data using NumPy-style indexing
z[:, :] = np.random.random((10000, 10000))
# Read data
data = z[0:100, 0:100] # Returns NumPy array
Core Operations
Creating Arrays
Zarr provides multiple convenience functions for array creation:
# Create empty array
z = zarr.zeros(shape=(10000, 10000), chunks=(1000, 1000), dtype='f4',
store='data.zarr')
# Create filled arrays
z = zarr.ones((5000, 5000), chunks=(500, 500))
z = zarr.full((1000, 1000), fill_value=42, chunks=(100, 100))
# Create from existing data
data = np.arange(10000).reshape(100, 100)
z = zarr.array(data, chunks=(10, 10), store='data.zarr')
# Create like another array
z2 = zarr.zeros_like(z) # Matches shape, chunks, dtype of z
Opening Existing Arrays
# Open array (read/write mode by default)
z = zarr.open_array('data.zarr', mode='r+')
# Read-only mode
z = zarr.open_array('data.zarr', mode='r')
# The open() function auto-detects arrays vs groups
z = zarr.open('data.zarr') # Returns Array or Group
Reading and Writing Data
Zarr arrays support NumPy-like indexing:
# Write entire array
z[:] = 42
# Write slices
z[0, :] = np.arange(100)
z[10:20, 50:60] = np.random.random((10, 10))
# Read data (returns NumPy array)
data = z[0:100, 0:100]
row = z[5, :]
# Advanced indexing
z.vindex[[0, 5, 10], [2, 8, 15]] # Coordinate indexing
z.oindex[0:10, [5, 10, 15]] # Orthogonal indexing
z.blocks[0, 0] # Block/chunk indexing
Resizing and Appending
# Resize array (v3: pass shape as a tuple)
z.resize((15000, 15000))
# Append data along an axis
z.append(np.random.random((1000, 10000)), axis=0) # Adds rows
Groups and Hierarchies
Groups organize multiple arrays hierarchically, similar to directories or HDF5 groups.
Creating and Using Groups
# Create root group
root = zarr.group(store='data/hierarchy.zarr')
# Create sub-groups
temperature = root.create_group('temperature')
precipitation = root.create_group('precipitation')
# Create arrays within groups
temp_array = temperature.create_array(
name='t2m',
shape=(365, 720, 1440),
chunks=(1, 720, 1440),
dtype='f4'
)
precip_array = precipitation.create_array(
name='prcp',
shape=(365, 720, 1440),
chunks=(1, 720, 1440),
dtype='f4'
)
# Access using paths
array = root['temperature/t2m']
# Visualize hierarchy
print(root.tree())
# Output:
# /
# ├── temperature
# │ └── t2m (365, 720, 1440) f4
# └── precipitation
# └── prcp (365, 720, 1440) f4
Group API (v3)
Use create_array / require_array (h5py-style create_dataset / require_dataset were removed in v3):
root = zarr.group('data.zarr')
arr = root.create_array('my_data', shape=(1000, 1000), chunks=(100, 100), dtype='f4')
grp = root.require_group('subgroup')
arr2 = grp.require_array('array', shape=(500, 500), chunks=(50, 50), dtype='i4')
Attributes and Metadata
Attach custom metadata to arrays and groups using attributes:
# Add attributes to array
z = zarr.zeros((1000, 1000), chunks=(100, 100))
z.attrs['description'] = 'Temperature data in Kelvin'
z.attrs['units'] = 'K'
z.attrs['created'] = '2024-01-15'
z.attrs['processing_version'] = 2.1
# Attributes are stored as JSON
print(z.attrs['units']) # Output: K
# Add attributes to groups
root = zarr.group('data.zarr')
root.attrs['project'] = 'Climate Analysis'
root.attrs['institution'] = 'Research Institute'
# Attributes persist with the array/group
z2 = zarr.open('data.zarr')
print(z2.attrs['description'])
Important: Attributes must be JSON-serializable (strings, numbers, lists, dicts, booleans, null).
Chunking, Compression, Storage, and Performance
- references/chunking_and_compression.md: sizing chunks to the access pattern (aim for ~1 MB, 5-100 MB on cloud), sharding, and codec choice.
- references/storage_backends.md: local, memory, ZIP, and fsspec remote stores (S3, GCS), with credential guidance — prefer IAM roles or workload identity, and never print credential values.
- references/integration.md: NumPy, Dask, and Xarray integration, thread safety, and consolidated metadata.
- references/performance_and_patterns.md: optimization, appendable time-series and large-matrix patterns, format conversion, and troubleshooting.
- references/api_reference.md and references/v3_migration.md: full API and the v2-to-v3 migration notes.
Additional Resources
Bundled references
| File | Contents |
|---|---|
references/api_reference.md | Function signatures, stores, codecs, indexing |
references/v3_migration.md | Zarr-Python 2→3 breaking changes and WIP features |
Official upstream
- Documentation: https://zarr.readthedocs.io/en/stable/
- 3.0 migration guide: https://zarr.readthedocs.io/en/stable/user-guide/v3_migration/
- Storage backends: https://zarr.readthedocs.io/en/stable/user-guide/storage/
- Zarr specifications: https://zarr-specs.readthedocs.io/
- GitHub: https://github.com/zarr-developers/zarr-python
- Developer chat: https://ossci.zulipchat.com/#narrow/channel/423692-Zarr-Python
Frequently asked questions about Zarr Python
Similar skills
Power BI Semantic Modeling
Optimize your Power BI data models with best practices.
Data Context Extractor
Tailor data analysis skills to your company's needs.
Power BI Performance Troubleshooting
Systematic guidance for optimizing Power BI performance.
Power BI Model Design Review
Optimize your Power BI data models with expert reviews.
Power BI DAX Formula Optimizer
Optimize your DAX formulas for better performance and clarity.
Fabric Lakehouse
Optimize your data solutions with Lakehouse best practices.
