
Azure AI Gateway
OfficialFreeManage AI models and tools with Azure API Management.
Free · Opens the source repo
What Azure AI Gateway does
The Azure AI Gateway skill enables developers to configure Azure API Management (APIM) as a gateway for AI models, tools, and agents. This skill is particularly valuable for those looking to implement governance and management features for AI applications. By using this skill, you can effectively manage semantic caching, token limits, and load balancing, ensuring that your AI models operate efficiently and cost-effectively. The skill also facilitates the integration of Azure OpenAI and AI Foundry models, allowing for a seamless backend setup.
With this skill, you can enforce policies that govern AI model usage, such as content safety measures and rate limiting for tools. This is crucial for maintaining compliance and ensuring that your AI applications do not produce harmful content. The skill provides a structured approach to configuring these policies, making it easier for developers to implement best practices in AI governance. Additionally, it offers troubleshooting guidance to address common issues that may arise during configuration.
The Azure AI Gateway skill is ideal for developers and organizations that are already leveraging Azure services and want to enhance their AI capabilities. Whether you are looking to optimize costs through token metrics or improve the safety of your AI outputs, this skill provides the tools necessary to achieve those goals. It also supports testing AI endpoints, which is essential for validating your configurations before deployment.
Overall, the Azure AI Gateway skill is a comprehensive solution for managing AI models and tools within the Azure ecosystem, making it a valuable addition for developers focused on AI governance and management.
When to use it
Use this skill when you need to manage AI model governance, configure backend integrations, or enforce policies for AI applications.
When not to use it
This skill may not be suitable if you are not using Azure services or if you require features outside the scope of API management and governance.
What you can build with it
Integrating Azure OpenAI
Easily add Azure OpenAI as a backend to your API Management service, enabling advanced AI capabilities.
Implementing Content Safety Policies
Configure content safety measures to filter harmful outputs from your AI models, ensuring compliance.
Optimizing API Costs
Use semantic caching and token metrics to reduce costs associated with API calls to AI models.
How to install Azure AI Gateway
View source1. Install with the skills CLI
npx skills add microsoft/azure-skills/azure-aigateway --agent claude-code2. Or install it manually
Download the skill folder and drop it into ~/.claude/skills/ for all projects, or .claude/skills/ to scope it to one repo. Restart Claude Code so it picks up the new skill.
Anthropic's agentic coding CLI, and the reference implementation of Agent Skills. Drop a skill folder into ~/.claude/skills and Claude Code loads it automatically whenever a task matches the skill's description. Claude Code docs
Inside SKILL.md
Written by microsoftAzure AI Gateway
Configure Azure API Management (APIM) as an AI Gateway for governing AI models, MCP tools, and agents.
To deploy APIM, use the azure-prepare skill. See APIM deployment guide.
When to Use This Skill
| Category | Triggers |
|---|---|
| Model Governance | "semantic caching", "token limits", "load balance AI", "track token usage" |
| Tool Governance | "rate limit MCP", "protect my tools", "configure my tool", "convert API to MCP" |
| Agent Governance | "content safety", "jailbreak detection", "filter harmful content" |
| Configuration | "add Azure OpenAI backend", "configure my model", "add AI Foundry model" |
| Testing | "test AI gateway", "call OpenAI through gateway" |
Quick Reference
| Policy | Purpose | Details |
|---|---|---|
azure-openai-token-limit | Cost control | Model Policies |
azure-openai-semantic-cache-lookup/store | 60-80% cost savings | Model Policies |
azure-openai-emit-token-metric | Observability | Model Policies |
llm-content-safety | Safety & compliance | Agent Policies |
rate-limit-by-key | MCP/tool protection | Tool Policies |
Get Gateway Details
# Get gateway URL
az apim show --name <apim-name> --resource-group <rg> --query "gatewayUrl" -o tsv
# List backends (AI models)
az apim backend list --service-name <apim-name> --resource-group <rg> \
--query "[].{id:name, url:url}" -o table
# Get subscription key
az apim subscription keys list \
--service-name <apim-name> --resource-group <rg> --subscription-id <sub-id>
Test AI Endpoint
GATEWAY_URL=$(az apim show --name <apim-name> --resource-group <rg> --query "gatewayUrl" -o tsv)
curl -X POST "${GATEWAY_URL}/openai/deployments/<deployment>/chat/completions?api-version=2024-02-01" \
-H "Content-Type: application/json" \
-H "Ocp-Apim-Subscription-Key: <key>" \
-d '{"messages": [{"role": "user", "content": "Hello"}], "max_tokens": 100}'
Common Tasks
Add AI Backend
See references/patterns.md for full steps.
# Discover AI resources
az cognitiveservices account list --query "[?kind=='OpenAI']" -o table
# Create backend
az apim backend create --service-name <apim> --resource-group <rg> \
--backend-id openai-backend --protocol http --url "https://<aoai>.openai.azure.com/openai"
# Grant access (managed identity)
az role assignment create --assignee <apim-principal-id> \
--role "Cognitive Services User" --scope <aoai-resource-id>
Apply AI Governance Policy
Recommended policy order in <inbound>:
- Authentication - Managed identity to backend
- Semantic Cache Lookup - Check cache before calling AI
- Token Limits - Cost control
- Content Safety - Filter harmful content
- Backend Selection - Load balancing
- Metrics - Token usage tracking
See references/policies.md for complete example.
Troubleshooting
| Issue | Solution |
|---|---|
| Token limit 429 | Increase tokens-per-minute or add load balancing |
| No cache hits | Lower score-threshold to 0.7 |
| Content false positives | Increase category thresholds (5-6) |
| Backend auth 401 | Grant APIM "Cognitive Services User" role |
See references/troubleshooting.md for details.
References
- Detailed Policies - Full policy examples
- Configuration Patterns - Step-by-step patterns
- Troubleshooting - Common issues
- AI-Gateway Samples
- GenAI Gateway Docs
SDK Quick References
- Content Safety: Python | TypeScript
- API Management: Python | .NET
Frequently asked questions about Azure AI Gateway
Similar skills
Build MCP Server
Streamline your MCP server development with guided discovery.
MCP Server Development Guide
Build high-quality MCP servers for LLM integration.
NemoClaw Documentation Access
Streamline your access to NemoClaw documentation with AI agents.
Agent Connect
Seamlessly connect agents to external APIs and services.
OpenAPI to MCP
Transform OpenAPI specs into MCP servers effortlessly.
Harness Evolve
Evolve your harness configurations without retraining.
