APIM Policy Patterns for AI Governance: Part 1 – Rate Limits, Token Quotas & Observability
How I would use Azure API Management to control AI request rates and token consumption, attribute usage, emit useful telemetry and handle backend throttling.
Practical insights on Platform Engineering, Azure, GitHub, Terraform, and AI-powered delivery
How I would use Azure API Management to control AI request rates and token consumption, attribute usage, emit useful telemetry and handle backend throttling.
An Azure AI Landing Zone should make the governed route the easiest route for delivery teams. This post covers how Azure API Management, Azure Policy, identity, networking, quotas, telemetry and clear ownership boundaries work together to control AI consumption without turning the platform team into an approval bottleneck.
Measuring AI-assisted engineering by activity alone misses the point. Active users, token usage, generated lines of code, and agent sessions are useful signals, but they do not tell you whether the work was reviewable, trusted, safe, or worth the cost.
This post looks at what platform teams should measure instead: workflow success, review effort, guardrail failures, context reuse, cost per useful outcome, and developer confidence.
How I built the azure-pricing skill for GitHub Copilot, using Azure MCP and the Azure Retail Prices API to bring live Azure pricing into architecture and engineering workflows.