Overview
Quickstart
Make a successful authenticated Chat Completions request in a few minutes.
1. Set your API key¶
Store the key in an environment variable. Do not commit it or place it in browser-side code.
export KOSCOMPUTE_API_KEY="your-api-key"
2. Send a request¶
curl https://api.koscompute.com/v1/chat/completions \
-H "Authorization: Bearer $KOSCOMPUTE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen/qwen3.8-27b",
"messages": [
{"role": "system", "content": "Answer clearly and briefly."},
{"role": "user", "content": "What is speculative decoding?"}
],
"max_completion_tokens": 180
}'
A successful response contains choices[0].message.content and a usage object. max_completion_tokens is the total completion budget, including reasoning tokens when reasoning is enabled.
3. Use an SDK¶
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["KOSCOMPUTE_API_KEY"],
base_url="https://api.koscompute.com/v1",
)
response = client.chat.completions.create(
model="qwen/qwen3.8-27b",
messages=[{"role": "user", "content": "Return one practical use for embeddings."}],
max_completion_tokens=120,
)
print(response.choices[0].message.content)
4. Handle temporary capacity¶
Retry 429, 503, and 504 responses according to Retry-After when present. Use exponential backoff with jitter and do not retry invalid 400 requests unchanged. The default chat model uses rental spare capacity, so it may be unavailable while cards are rented; check its live readiness before depending on it.
Where to go next¶
- Query the live model catalog.
- Add streaming.
- Learn tool calling or structured output.
- Review errors and retries before production rollout.