Research
Small, checkable AI on hardware you control.
Useful AI that runs on hardware you control, with limits you can enforce and results you can check. Aitium builds small specialized models, the engine that serves them on one or two 24GB GPUs, and the runtime that keeps AI agents inside the boundaries their operators set. We publish what we measure, including what fails.
Research areas
The questions we are working on
Small specialized models
In evaluationCan a small model, sized for one local GPU, do a narrow expert task well enough to trust?
CyberGuard targets security triage and threat analysis, measured on precision, recall and false-positive rate before any release.
See the work ↗Pharmacogenomics
ResearchCan matching DNA variants to medicine compatibility become a model query without losing expert-level accuracy?
Cogenics validates variant-to-drug associations before any model is trained, using CPIC prescribing guidelines as the evaluation axis and ClinVar as the reference dataset.
See the work ↗Local serving
LiveHow capable a model can run privately on one or two 24GB consumer GPUs without silently degrading precision?
Ember is a registry-driven gateway and GPU worker, live on Aitium hardware and serving an OpenAI-compatible API.
See the work ↗Governed agents
In developmentHow do you give an AI agent real autonomy while keeping its reach, spend and actions enforceably bounded?
Foxtail Harness puts a tool registry, delegation scopes, atomic budgets, review gates and an audit trail between agents and production systems.
See the work ↗Roadmap
What comes next
- Live nowEmber inference engine serving on Aitium hardware.
- Q4 2026CyberGuard evaluation results published.
- Q4 2026Foxtail Harness early access.
- Q1 2027Cogenics validation complete; training decision.
Funding focus
What support makes possible
Outside funding goes to three things, each tied to a public result on our roadmap.
01Compute for evaluation and training
GPU time to train CyberGuard and run it against Cybench, CyberSecEval and the AWS Deception Benchmark.
Result Full evaluation results published, misses included.
Q4 2026
02Expert validation for Cogenics
Blinded review by pharmacogenomics specialists, checking variant-to-drug associations against CPIC guidelines before any model is trained.
Result Validation published, with a go/no-go training decision.
Q1 2027
03Taking Foxtail Harness to early access
Engineering time to harden the governed agent runtime — registry, budgets, review gates and audit trail — for outside teams.
Result Early access for external teams.
Q4 2026
Evidence
Check the work
We publish what we measure, including the misses.