Skip to main content

Core Capabilities

AI Gateway capabilities are organized into separate task guides. This page provides the capability map. Start with the Quickstart to make a model request.

1. Unified AI API Access​

Connect public inference, dedicated endpoints, and commercial APIs for text, embeddings, images, audio, and other supported tasks. See Connect and Call Models for endpoints, SDK usage, and model differences.

2. Multi-Model and Multi-Provider Unified Scheduling​

Configure several upstreams for a model and distribute requests using round robin, weights, and session hashing. See Routing and Reliability and Performance and Cost Optimization for configuration and cache-affinity boundaries.

3. Authentication and Quota Management​

Use personal or organization API keys to manage caller identity, expiry, and spending limits. Key limits, token usage, TPM, and account balances are distinct concepts. See Identity and Access and Limits, Metering, and Costs.

4. Content Safety Inspection​

Apply enabled checks to input and output in streaming and non-streaming scenarios. Configuration, trusted-request exceptions, and data-handling boundaries are covered in Content Safety and Data Handling.

5. High Availability and Automatic Failover​

Health checks, circuit breaking, and request-level failover help reduce upstream failure impact. Outcomes depend on the request stage, available upstreams, and configuration. See Routing and Reliability.

6. Request Logging and Data Retention​

Use request records to investigate inputs, outputs, and tool calls, and gateway metrics to review traffic and health. See Monitoring, Logs, and Troubleshooting. Request-body collection and reuse follow deployment data policies; see Content Safety and Data Handling.

7. Usage Statistics and Billing​

Measure tokens, calls, duration, or other configured units for supported tasks. See Limits, Metering, and Costs for usage management and Operating Paid Model Services for pricing and financial operations.

Team onboarding and operations​

Coordinate personal, organization, and administrator responsibilities for onboarding, credentials, and maintenance. See Team Management and Operations.

Performance and cost optimization​

Use routing, Prompt Cache, inference configuration, and model evaluation to compare quality, latency, and charges. See Performance and Cost Optimization.

MCP and Agent scenarios​

Distinguish remote MCP connections, platform-hosted services, Agents, and sandbox proxies. See MCP and Agent Access.