Updates
Changelog
User-visible API, model, capability, limit, pricing, and reliability changes.
2026-10-01¶
- Removed the MiMo model family from the public API.
- Qwen3.8 27B now runs only on available rental spare capacity. No GPU is reserved for the API by default. Rental jobs take priority and can interrupt active Qwen requests. When no Qwen backend is ready, requests return
503rather than falling back to another model. - Host GPU checks also take priority over spare-capacity Qwen workers. Qwen readiness can briefly change even when no new rental appears.
2026-09-29¶
- Added
dealignai/mimo-v2.6-flash-uncensoredwith themimo-v2.6-flash-uncensoredalias, text chat, reasoning, tools, structured output, 131,072-token context, and published token pricing. - Kept
qwen/qwen3.8-27bavailable on opportunistic GPU capacity. Compatible text requests may use the configured MiMo fallback when no Qwen backend is ready; the requested model ID and Qwen pricing remain in effect.
2026-08-17¶
- Rebuilt the public documentation as an English-only, Markdown-first developer portal.
- Added generated OpenAPI 3.1 JSON/YAML, integrated API reference, raw Markdown pages, and
/llms.txt. - Made the public model catalog live from
GET /v1/modelswith a timestamped fallback snapshot. - Clarified total completion-token budgets, stateless Responses behavior, Anthropic-specific
529mapping, combined vision upload limits, and Zero Data Retention metadata boundaries. - Added contract, example, link, rendering, drift, and public-artifact leak validation.
Future entries include only changes that affect API users. Internal deployment and infrastructure details remain in the private operations knowledge base.