KosComputeAPI
Overview

Quickstart

Make a successful authenticated Chat Completions request in a few minutes.

View Markdown

1. Set your API key

Store the key in an environment variable. Do not commit it or place it in browser-side code.

export KOSCOMPUTE_API_KEY="your-api-key"

2. Send a request

curl https://api.koscompute.com/v1/chat/completions \
  -H "Authorization: Bearer $KOSCOMPUTE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.8-27b",
    "messages": [
      {"role": "system", "content": "Answer clearly and briefly."},
      {"role": "user", "content": "What is speculative decoding?"}
    ],
    "max_completion_tokens": 180
  }'

A successful response contains choices[0].message.content and a usage object. max_completion_tokens is the total completion budget, including reasoning tokens when reasoning is enabled.

3. Use an SDK

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["KOSCOMPUTE_API_KEY"],
    base_url="https://api.koscompute.com/v1",
)

response = client.chat.completions.create(
    model="qwen/qwen3.8-27b",
    messages=[{"role": "user", "content": "Return one practical use for embeddings."}],
    max_completion_tokens=120,
)
print(response.choices[0].message.content)

4. Handle temporary capacity

Retry 429, 503, and 504 responses according to Retry-After when present. Use exponential backoff with jitter and do not retry invalid 400 requests unchanged. The default chat model uses rental spare capacity, so it may be unavailable while cards are rented; check its live readiness before depending on it.

Where to go next

Type to search guides, models, and API reference.