Models
Live public models, capabilities, limits, pricing, readiness, and available capacity.
The catalog below loads from GET /v1/models once per page visit. Runtime fields are never maintained by hand in guide prose. If live discovery fails, the page keeps working from the build snapshot and labels the data as a snapshot.
Live model catalog
Fetched once on load; build snapshot available as fallback.
Choose by capability¶
Use supported_endpoints, input_modalities, and supported_features instead of assuming every chat model supports tools, vision, reasoning, or structured output. Context and output limits can differ by modality.
Read the metadata directly¶
curl https://api.koscompute.com/v1/models
Pricing is returned in the public model metadata. Chat token rates use per-token fields such as prompt, completion, and input_cache_read; display clients can multiply those values to their preferred unit. Audio and speech models expose media-specific pricing units where applicable.
Qwen3.8 availability¶
qwen/qwen3.8-27b uses idle GPUs that are offered for rental. It has no permanently reserved GPU capacity. An internal, disposable worker starts after a card becomes free and joins the API only after loading and readiness checks pass. Rentals and host GPU checks take priority: the worker is removed from API routing and stopped before the card is used for either purpose. In-flight requests can be interrupted, including while no rental is active.
When no backend is ready, Qwen requests return 503; the service does not substitute another model. Check GET /v1/models for current readiness and handle 503 with backoff. The live state can change between discovery and request admission.