Public inference (Serverless API)
Public inference is deployed and maintained by the current platform's administrator. Use that platform's model, address and API Key; a private deployment does not automatically call the public community.
Use the service
- Select a model with public inference enabled and submit a short playground request.
- Create an API Key in the intended personal or organization settings. Do not use a Git Access Token.
- Copy the API address, model identifier and example from View code on the model page.
- Verify a non-streaming response first, then use streaming or additional parameters if supported. Check usage records and the request result.
Usage and fees
Check trial allowances and prices on the selected service page; not every model or request is free.
In commercial operating environments, Token charges belong to the Key's personal or organization account and cannot be offset with compute vouchers. Scaling, concurrency and caching depend on the deployed service configuration.
For dedicated resources and custom runtime settings, see create an endpoint.