Skip to contentAitium

Platform / Model serving

Live

Small-model serving on 24GB GPUs.

Ember is an inference engine for serving small models on 24GB GPUs.

Why it exists

Problems we are solving

Running models often means cloud infrastructure.

Ember serves small models on hardware Aitium owns.

Many serving stacks target high-end GPUs.

Ember targets 24GB consumer GPUs.

Architecture

How a request is served

Every request takes the same path, and the whole path runs on Aitium-operated hardware.

  1. Request

    A client sends an OpenAI-compatible request.

  2. Gateway

    Validates the request and selects a registered model.

  3. Registry

    Holds each model's runtime configuration.

  4. Worker

    Owns model lifecycle and GPU assignment.

  5. Runner

    Starts the backend runtime for the selected model.

  6. GPU

    Executes inference on one or two local 24GB GPUs.

Specification

Live serving specification

ModelQwen3.8-27B
Parameters27B
GPU1–2 × 24GB GPUs
APIOpenAI-compatible

Design principles

Ember does not silently degrade precision to make a model fit. Backends are selected per model archetype. Quantized KV cache is a standard lever. Spec-decode and prefix caching are on the roadmap for evaluation, not promised performance features.

What it is not

Ember is not a cloud API and not a multi-tenant serving platform. It is private/on-device inference for models on hardware Aitium controls.

What it does

The serving layer for Aitium models.

24GB GPU target

Ember is built around small models that fit on 24GB GPUs.

Aitium code

The registry-driven gateway and GPU worker are written and operated by Aitium.

OpenAI-compatible

Applications call it like any OpenAI endpoint. No client changes, no data leaving your hardware.

No silent degradation

Ember never quietly drops precision to make a model fit. Backends are chosen per model archetype.

Process

How it works

  1. Register the model

    A model is added to the registry with the runtime configuration Ember needs to serve it. The gateway uses that registry entry to decide how requests are routed.

  2. Load on GPU

    The GPU worker loads a small model that fits the 24GB target.

  3. Serve

    Applications call the gateway endpoint, and Ember routes the request to the registered model runtime. The serving path stays on Aitium-operated hardware.

Access

Talk about Ember

Ember is live on Aitium hardware. Talk to us about serving your models with it.

Talk about Ember
  • Ember

    Inference engine · live

    Live

Requirements

  • A 24GB GPU
  • A small model to serve

Questions / answers

Ember FAQ

What does Ember do?

Ember is Aitium's inference engine. It serves small models on 24GB GPUs through a registry-driven gateway and a GPU worker.

How do applications call it?

Through an OpenAI-compatible endpoint, so existing clients work unchanged.

Which models use it?

Aitium's own models — CyberGuard and Cogenics — and any small model that fits on 24GB cards.

The platform that serves our models and governs agents.