Running models often means cloud infrastructure.
Ember serves small models on hardware Aitium owns.
Ember is an inference engine for serving small models on 24GB GPUs.
Why it exists
Ember serves small models on hardware Aitium owns.
Ember targets 24GB consumer GPUs.
Architecture
Every request takes the same path, and the whole path runs on Aitium-operated hardware.
A client sends an OpenAI-compatible request.
Validates the request and selects a registered model.
Holds each model's runtime configuration.
Owns model lifecycle and GPU assignment.
Starts the backend runtime for the selected model.
Executes inference on one or two local 24GB GPUs.
Specification
| Model | Qwen3.8-27B |
|---|---|
| Parameters | 27B |
| GPU | 1–2 × 24GB GPUs |
| API | OpenAI-compatible |
Ember does not silently degrade precision to make a model fit. Backends are selected per model archetype. Quantized KV cache is a standard lever. Spec-decode and prefix caching are on the roadmap for evaluation, not promised performance features.
Ember is not a cloud API and not a multi-tenant serving platform. It is private/on-device inference for models on hardware Aitium controls.
What it does
Ember is built around small models that fit on 24GB GPUs.
The registry-driven gateway and GPU worker are written and operated by Aitium.
Applications call it like any OpenAI endpoint. No client changes, no data leaving your hardware.
Ember never quietly drops precision to make a model fit. Backends are chosen per model archetype.
Process
A model is added to the registry with the runtime configuration Ember needs to serve it. The gateway uses that registry entry to decide how requests are routed.
The GPU worker loads a small model that fits the 24GB target.
Applications call the gateway endpoint, and Ember routes the request to the registered model runtime. The serving path stays on Aitium-operated hardware.
Access
Ember is live on Aitium hardware. Talk to us about serving your models with it.
Talk about EmberInference engine · live
Questions / answers
Ember is Aitium's inference engine. It serves small models on 24GB GPUs through a registry-driven gateway and a GPU worker.
Through an OpenAI-compatible endpoint, so existing clients work unchanged.
Aitium's own models — CyberGuard and Cogenics — and any small model that fits on 24GB cards.
The platform that serves our models and governs agents.