AI-SPECIFIC API GATEWAY ARCHITECTURE
Standard API gateways (Kong, Envoy, AWS API Gateway) handle HTTP-level concerns but lack AI-specific features needed for model endpoint security. An AI-specific gateway layer must understand: token-based pricing and rate limiting (not just request count, but token volume and generation length), streaming connections (maintaining SSE or WebSocket connections with per-user limits on connection count and duration), model-specific routing (directing requests to the appropriate model version or LoRA adapter), and GPU-aware load balancing (routing to nodes with available KV cache capacity or GPU memory).
The gateway architecture for AI inference uses a two-tier approach. Tier 1 is the standard HTTP gateway handling: TLS termination, authentication (API keys, OAuth2, JWT), basic rate limiting (requests per second per user), and request validation. Tier 2 is the AI gateway handling: token counting and budget enforcement, streaming connection management, model selection routing, cost tracking and billing integration, and GPU utilization-aware scheduling. The AI gateway sits between the HTTP gateway and the inference engine, maintaining per-user token buckets. Implementations build on custom Envoy filters, NGINX Lua modules, or purpose-built proxies like LiteLLM, BentoML Gateway, or custom Go/Rust proxies. The AI gateway must handle 10,000+ concurrent streaming connections with sub-millisecond routing decisions.
| Gateway Feature | Tier 1 (HTTP) | Tier 2 (AI-Specific) |
|---|---|---|
| Authentication | API keys, OAuth2, JWT | API key + usage plan mapping |
| Rate limiting | Requests/sec per user | Tokens/sec + requests/sec |
| Streaming support | Connection limits | SSE/WebSocket per-user tracking |
| Routing | Path-based | Model version + adapter routing |
| Cost tracking | Per-request | Per-token billing with tiered pricing |
| GPU awareness | None | Load-aware: KV cache, memory |
| Caching | Response-level | Semantic cache integration |
TOKEN-AWARE RATE LIMITING STRATEGIES
AI inference rate limiting must account for both request volume and token consumption. A user sending 10 requests with single-token responses uses vastly less GPU compute than a user sending 10 requests with 4,000-token responses. Token-aware rate limiting tracks both dimensions: request rate (RPS) and token consumption rate (tokens per second). The rate limiter maintains a dual-bucket algorithm: a request bucket (permits N requests per minute) and a token bucket (permits M tokens output per minute). A request consumes from both buckets, preventing users from exhausting GPU capacity through long generations.
Implementation uses sliding window counters in Redis (token bucket state) and in-memory counters for burst tolerance. For a production system with 10K+ users, the rate limiter must handle: 100K+ token bucket lookups per second (each request checks both buckets), sub-5ms lookup latency, atomic increment and decrement operations for bucket state, and distributed consistency across gateway replicas. The token bucket variable cost is estimated per request before generation begins using the maximum possible tokens (4K for chat, 8K for document analysis) or a user-provided max_tokens parameter. The estimate ensures users cannot burn through their budget by requesting unlimited generation lengths. For streaming responses, the rate limiter updates token consumption incrementally as tokens are generated, allowing the gateway to enforce mid-stream budget limits by truncating or terminating the generation.
AUTHENTICATION AND MULTI-TENANT AUTHORIZATION
Multi-tenant model serving requires fine-grained authorization beyond API key validation. Each request carries: tenant identity (organization, team, or user), model access permissions (which models the tenant is authorized to use), usage tier (pricing plan determining rate limits and available features), and data handling classification (PII, PHI, or general data). The authorization layer checks all four dimensions before routing the request. For regulated deployments (HIPAA, GDPR), the authorization check must also verify data residency requirements (ensure PII data is processed in approved geographic regions).
The authorization infrastructure uses a policy-based access control (PBAC) system. Policies are defined in a structured language (OPA Rego, AWS Cedar) and evaluated at gateway level. Example policy: "Users from tenant ACME can access Llama 3.1 70B and GPT-4o at tier-2 rate limits, but only from EU-located GPU nodes. PII-scanned requests require approved SaaS API mode." The policy engine must evaluate in under 5ms per request, caching compiled policies from a centralized store. Policy evaluation logs feed into the audit trail for compliance verification. For fine-grained access, model endpoints can be further restricted by: input length, output length, specific features (function calling, vision, structured output), and deployment environment (development, staging, production).
| Authorization Dimension | Check Performed | Latency Impact |
|---|---|---|
| API key validity | Signature verification | 1-3ms |
| Model access permission | Policy evaluation | 2-5ms |
| Usage tier check | Rate limit bucket lookup | 1-3ms |
| Data residency check | Region attribute mapping | 0.5-2ms |
| Feature restriction | Policy evaluation | 1-3ms |
| Total authorization | All dimensions | 5-15ms |
GPU-AWARE LOAD BALANCING AND SCHEDULING
AI inference load balancing must account for GPU state, not just request count. Two requests to a 70B model on an 8x H100 node have very different GPU impact depending on: current batch size and composition, KV cache utilization, active LoRA adapter residency, and remaining GPU memory for the new request's context. GPU-unaware round-robin load balancing leads to cascading failures where heavily loaded nodes receive new requests while idle nodes sit unused. An AI-specific load balancer uses GPU telemetry to make routing decisions.
The load balancer maintains per-node metrics via a shared telemetry bus (Prometheus + gRPC streaming): current request count, active KV cache memory, GPU utilization percentage, inference latency P50/P99 over the last minute, and available LoRA adapter slots. The routing algorithm scores each node using a weighted combination: node_score = w1 * (1 - GPU_util) + w2 * (1 - KV_cache_usage) + w3 * (available_adapters / max_adapters) + w4 * (latency_normalized). The node with the highest score receives the request. Nodes are also flagged for cooldown if inference latency exceeds thresholds (e.g., P99 > 5s). The load balancer must update node scores with sub-100ms freshness to avoid routing to stale state. For multi-region deployments, regional affinity and data residency policies add routing constraints that are evaluated before node scoring.
USAGE BUDGET ENFORCEMENT AND COST TRACKING
For model endpoints offered as APIs (either internally or to external customers), usage budget enforcement prevents cost overruns and ensures fair resource allocation. Each tenant has a budget plan defining: monthly token allowance (input + output tokens), concurrent request limit, max tokens per request, and available models. The budget enforcement layer checks these limits at request time and pre-checks: current month-to-date token consumption, remaining budget, and whether the current request would exceed budget. Hard enforcement stops requests that exceed budget; soft enforcement sends warning headers.
The budget tracking infrastructure requires: a token counting service that tracks every input and output token per tenant, a cost aggregation pipeline that converts token counts to monetary cost using per-model pricing tiers, and a budget alert system that notifies tenants at configurable thresholds (50%, 80%, 90%, 100% of budget). The token counter must handle streaming responses by charging incrementally as tokens are delivered, not just at completion. For a multi-tenant platform serving 100M tokens/day across 100 tenants, the tracking infrastructure processes 5,000-20,000 token events per second, with sub-5 second freshness for budget dashboards. On ClusterBid, teams can deploy the budget tracking infrastructure on the same GPU nodes as the inference engine, using GPU-generated metrics for cost accounting that is verifiable and auditable.
| Budget Component | Implementation | Scale Target |
|---|---|---|
| Monthly token allowance | Redis counter per tenant | 10M tokens default |
| Concurrent request limit | In-memory counter per tenant | 50 concurrent default |
| Max tokens per request | Pre-check at gateway | 4,096 / 8,192 / 32,768 tiers |
| Model access list | Policy evaluation | Per-tenant model allowlist |
| Cost aggregation | Hourly batch to data warehouse | 99.9% accuracy |
| Budget alerts | Webhook + email at thresholds | 50/80/90/100% |
SECURITY CONSIDERATIONS AND AUDIT
Model endpoints present a unique attack surface compared to traditional APIs. Key security considerations: prompt injection attacks (adversarial inputs that manipulate model behavior), model extraction attacks (repeated API calls to distill a model through black-box queries), denial-of-wallet attacks (consuming excessive expensive tokens to drive up costs), and data exfiltration (embedding PII in generated outputs). The access control infrastructure must defend against all four vectors simultaneously. Prompt injection detection adds a classifier layer between the gateway and the inference engine, scanning inputs for jailbreak patterns, prompt leakage attempts, and system prompt override attempts.
Model extraction defenses include rate limiting on the number of requests per session (not just per user), input diversity detection (flagging users who send requests spanning the same semantic space, indicating model distillation), and output perturbation (adding controlled noise to logprobs if model extraction is detected). Denial-of-wallet prevention uses budget thresholds with automatic overage blocks and anomaly detection on usage patterns. All security events are logged to the audit trail with severity classification and automated incident response triggers. The security infrastructure adds 10-50ms of pre-inference latency for prompt injection scanning (using small classifier models on GPU or CPU) and 5-15ms for post-inference output scanning. For high-security deployments, the total access control stack adds approximately 30-80ms end-to-end latency, compared to 15-30ms for basic authentication and rate limiting.
