
Python Observability
FreeEnhance your Python apps with structured logging and metrics.
Free · Opens the source repo
What Python Observability does
The Python Observability skill provides essential patterns for instrumenting Python applications with structured logging, metrics collection, and distributed tracing. This skill is particularly useful for developers looking to improve their application's observability without the need to deploy new code when issues arise in production. By implementing these patterns, you can gain insights into your application's performance and behavior, making it easier to diagnose problems and optimize your services.
Structured logging is one of the core concepts emphasized in this skill. By emitting logs in JSON format with consistent fields, you enable machine-readable logs that facilitate powerful querying and alerting. This is crucial for production environments where understanding the context of logged events can significantly reduce the time spent debugging. Additionally, the skill covers the Four Golden Signals—latency, traffic, errors, and saturation—which are vital metrics to monitor for each service boundary.
Another important aspect is the use of correlation IDs, which allow you to trace requests across multiple services. By threading a unique ID through all logs and spans for a single request, you can achieve end-to-end visibility, making it easier to pinpoint where issues occur in complex systems. The skill also discusses the importance of bounded cardinality in metrics collection to avoid excessive storage costs due to unbounded label values.
This skill is ideal for developers and DevOps engineers who are responsible for maintaining and debugging production systems. It provides practical examples and code snippets to help you implement observability best practices in your Python applications effectively.
When to use it
Use this skill when you need to add structured logging, implement metrics collection, or set up tracing in your Python applications.
When not to use it
This skill may not be suitable for applications that do not require detailed observability or for environments where minimal logging is preferred.
What you can build with it
Adding Structured Logging
Implement structured logging in your Python application using the provided patterns to enhance log readability and query capabilities.
Setting Up Distributed Tracing
Use the correlation ID patterns to set up distributed tracing across microservices, enabling better tracking of requests.
Debugging Production Issues
Utilize the observability patterns to quickly diagnose and resolve production issues without needing to deploy new code.
How to install Python Observability
View source1. Install with the skills CLI
npx skills add wshobson/agents/python-observability --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by wshobsonPython Observability
Instrument Python applications with structured logs, metrics, and traces. When something breaks in production, you need to answer "what, where, and why" without deploying new code.
When to Use This Skill
- Adding structured logging to applications
- Implementing metrics collection with Prometheus
- Setting up distributed tracing across services
- Propagating correlation IDs through request chains
- Debugging production issues
- Building observability dashboards
Core Concepts
1. Structured Logging
Emit logs as JSON with consistent fields for production environments. Machine-readable logs enable powerful queries and alerts. For local development, consider human-readable formats.
2. The Four Golden Signals
Track latency, traffic, errors, and saturation for every service boundary.
3. Correlation IDs
Thread a unique ID through all logs and spans for a single request, enabling end-to-end tracing.
4. Bounded Cardinality
Keep metric label values bounded. Unbounded labels (like user IDs) explode storage costs.
Quick Start
import structlog
structlog.configure(
processors=[
structlog.processors.TimeStamper(fmt="iso"),
structlog.processors.JSONRenderer(),
],
)
logger = structlog.get_logger()
logger.info("Request processed", user_id="123", duration_ms=45)
Fundamental Patterns
Pattern 1: Structured Logging with Structlog
Configure structlog for JSON output with consistent fields.
import logging
import structlog
def configure_logging(log_level: str = "INFO") -> None:
"""Configure structured logging for the application."""
structlog.configure(
processors=[
structlog.contextvars.merge_contextvars,
structlog.processors.add_log_level,
structlog.processors.TimeStamper(fmt="iso"),
structlog.processors.StackInfoRenderer(),
structlog.processors.format_exc_info,
structlog.processors.JSONRenderer(),
],
wrapper_class=structlog.make_filtering_bound_logger(
getattr(logging, log_level.upper())
),
context_class=dict,
logger_factory=structlog.PrintLoggerFactory(),
cache_logger_on_first_use=True,
)
# Initialize at application startup
configure_logging("INFO")
logger = structlog.get_logger()
Pattern 2: Consistent Log Fields
Every log entry should include standard fields for filtering and correlation.
import structlog
from contextvars import ContextVar
# Store correlation ID in context
correlation_id: ContextVar[str] = ContextVar("correlation_id", default="")
logger = structlog.get_logger()
def process_request(request: Request) -> Response:
"""Process request with structured logging."""
logger.info(
"Request received",
correlation_id=correlation_id.get(),
method=request.method,
path=request.path,
user_id=request.user_id,
)
try:
result = handle_request(request)
logger.info(
"Request completed",
correlation_id=correlation_id.get(),
status_code=200,
duration_ms=elapsed,
)
return result
except Exception as e:
logger.error(
"Request failed",
correlation_id=correlation_id.get(),
error_type=type(e).__name__,
error_message=str(e),
)
raise
Pattern 3: Semantic Log Levels
Use log levels consistently across the application.
| Level | Purpose | Examples |
|---|---|---|
DEBUG | Development diagnostics | Variable values, internal state |
INFO | Request lifecycle, operations | Request start/end, job completion |
WARNING | Recoverable anomalies | Retry attempts, fallback used |
ERROR | Failures needing attention | Exceptions, service unavailable |
# DEBUG: Detailed internal information
logger.debug("Cache lookup", key=cache_key, hit=cache_hit)
# INFO: Normal operational events
logger.info("Order created", order_id=order.id, total=order.total)
# WARNING: Abnormal but handled situations
logger.warning(
"Rate limit approaching",
current_rate=950,
limit=1000,
reset_seconds=30,
)
# ERROR: Failures requiring investigation
logger.error(
"Payment processing failed",
order_id=order.id,
error=str(e),
payment_provider="stripe",
)
Never log expected behavior at ERROR. A user entering a wrong password is INFO, not ERROR.
Pattern 4: Correlation ID Propagation
Generate a unique ID at ingress and thread it through all operations.
from contextvars import ContextVar
import uuid
import structlog
correlation_id: ContextVar[str] = ContextVar("correlation_id", default="")
def set_correlation_id(cid: str | None = None) -> str:
"""Set correlation ID for current context."""
cid = cid or str(uuid.uuid4())
correlation_id.set(cid)
structlog.contextvars.bind_contextvars(correlation_id=cid)
return cid
# FastAPI middleware example
from fastapi import Request
async def correlation_middleware(request: Request, call_next):
"""Middleware to set and propagate correlation ID."""
# Use incoming header or generate new
cid = request.headers.get("X-Correlation-ID") or str(uuid.uuid4())
set_correlation_id(cid)
response = await call_next(request)
response.headers["X-Correlation-ID"] = cid
return response
Propagate to outbound requests:
import httpx
async def call_downstream_service(endpoint: str, data: dict) -> dict:
"""Call downstream service with correlation ID."""
async with httpx.AsyncClient() as client:
response = await client.post(
endpoint,
json=data,
headers={"X-Correlation-ID": correlation_id.get()},
)
return response.json()
Detailed worked examples and patterns
Detailed sections (starting with ## Advanced Patterns) live in references/details.md. Read that file when the navigation summary above is insufficient.
Best Practices Summary
- Use structured logging - JSON logs with consistent fields
- Propagate correlation IDs - Thread through all requests and logs
- Track the four golden signals - Latency, traffic, errors, saturation
- Bound label cardinality - Never use unbounded values as metric labels
- Log at appropriate levels - Don't cry wolf with ERROR
- Include context - User ID, request ID, operation name in logs
- Use context managers - Consistent timing and error handling
- Separate concerns - Observability code shouldn't pollute business logic
- Test your observability - Verify logs and metrics in integration tests
- Set up alerts - Metrics are useless without alerting
Frequently asked questions about Python Observability
Similar skills
Turborepo
Optimized build system for JavaScript/TypeScript monorepos.
Azure Pipelines Validation
Streamline your Azure DevOps pipeline changes locally.
Azure Developer CLI
Streamline your Azure project workflows with best practices.
Azure Container Registry CLI
Manage Azure Container Registry resources with ease.
Aspire
Build and orchestrate polyglot distributed applications seamlessly.
Vercel CLI
Manage and deploy Vercel projects from the command line.
