Big models, half the memory.
Matched on benchmarks.
Uluka builds compressed models and the CUDA runtime that serves them. Our 26B mixture-of-experts fits in 8.73 GB, and on four public benchmarks it finished level with the standard 4-bit builds of its class while taking about half their memory — so the same work fits on a smaller GPU, in an account you own.
An example exchange, replayed in your browser. No model runs on this page.
One GPU, one process
The pack loads into video memory once and answers from there. No inference service in the middle, no second machine.
Sparse by design
A router picks 8 of 128 experts in every layer for each token, so a 26B model reads only about 4B parameters per step — and the whole model is 8.73 GB on disk.
The API you already call
An OpenAI-compatible endpoint with streaming and tool calls, at an address you choose.
The model, and the runtime
built to serve it.
Uluka ships as a weight pack and a CUDA runtime written for that pack: its own kernels, its own memory plan, one GPU. You get an endpoint, not a subscription.
- Model
- UlukaLabs 26B
- Design
- Mixture of experts: 30 layers of 128 experts, 8 picked per token
- Parameters
- 26B in total, about 4B read for each token
- On disk
- 8.73 GB, the whole pack — the same 26B at full precision is 51.6 GB
- Context
- 8K tokens
- Input
- Text and images
- API
- OpenAI-compatible: chat completions, streaming, tool calls
- Runs on
- One NVIDIA GPU with 12 GB or more — a desktop RTX card or a T4, L4, L40S, A100, H100 in a rack
- Where
- A workstation, your own server, or an instance in your cloud account
Our own kernels
The runtime is written for this pack rather than adapted to it, which is what gets 86 tok/s out of one mainstream desktop card.
It fits the GPU it finds
At start-up it reads the hardware and plans how the model sits in memory, instead of assuming one machine.
Vision in the same pack
Images go in alongside text — screenshots, scans, photographs — with no second model to host.
Tool calls that hold up
95 out of 100 on our tool-calling gate, so it can drive the systems around it.
Streams as it answers
Tokens arrive as they are produced, so your own product can show them the moment they exist.
One file to move
8.73 GB of weights and a runtime. That is the whole deployment.
Measured,
not promised.
Every number on this page comes from a run with the items, the settings and the seed written down first, and the intervals printed beside the result. Where we have not measured something, it says so.
Quality four public benchmarks, same items, same settings, 95% intervals
Mean of MMLU, HellaSwag, GSM8K and IFEval. Build A is 15.63 GB installed, build B 17.22 GB, Uluka 8.73 GB. Uluka Labs benchmarks, September 2026.
A tie, not a win. We match the 4-bit builds at about half their size, and we publish the intervals that back it: on MMLU and HellaSwag we are behind, on GSM8K ahead. Full report to customers on request.
- MMLU
- 59.0 — build A 61.4, build B 58.8
- HellaSwag
- 45.1 — build A 47.3, build B 47.2
- GSM8K
- 87.8 — build A 84.5, build B 87.2
- IFEval
- 87.8 — build A 88.5, build B 87.5
- Paired
- On the mean: −0.5 [−1.7, +0.6] against A and −0.3 [−1.4, +0.8] against B — both intervals cross zero
- Tool calls
- 95 out of 100 on the tool-calling gate
What it runs on and how big a model each one can hold
Click a card to open it · drag to turn it
Drag to turn the card · scroll to zoom
—
—
- Memory
- —
- For models
- —
- Uluka 26B
- —
- Side by side
- —
- Speed
- —
How big a model fits
—
—
Half the memory
is half the GPU you rent.
Size is not a spec sheet line. A 15 to 17 GB model needs a 24 GB accelerator; 8.73 GB fits a 16 GB one with room for context. Same answers, one instance class down, every hour it runs — and on the GPUs you already own, the same card holds a model nearly twice the size.
Every row is a complete installed model, weights and all. Uluka Labs measurements, September 2026.
- Cheaper box
- One GPU class down for every endpoint, for as long as it runs
- Same box
- On a card you already own, room for a model about 1.8× larger
- More at once
- 8 copies of the pack fit in an 80 GB accelerator instead of 4
- Longer context
- Memory the weights do not take is memory the context can
- Quality
- Level with both 4-bit builds on the four-benchmark mean
The saving is compute you no longer rent, not a licence you no longer buy. It applies the same way to a card under a desk: the hardware you already have does more.
Compute you stop renting your numbers, not ours
AWS EC2 on-demand list prices, Linux, US East (N. Virginia), instance only, checked 24 September 2026: $0.805 an hour for a 24 GB GPU instance, $0.526 for a 16 GB one. Which instance you need is decided by what has to fit in video memory; what a request costs depends on your workload.
$9,769a year, not rented
- 24 GB box
- $2,350 / month
- 16 GB box
- $1,536 / month
- Difference
- $814 / month
- No per-token bill and no per-seat bill — you rent the GPU, not the answers
- No rate limit and no queue you share with strangers
- 8.73 GB to move, sized for the GPUs your fleet already has
Your data never leaves
the account you run it in.
Uluka answers from a GPU you control. There is no inference service in the middle, so prompts, documents and answers stay inside the machine or the account you deployed it to.
Runs where you put it
A workstation, a server in your rack, or an instance in your own cloud account. The endpoint's address is yours.
Nothing is sent to us
No prompts, no documents, no answers, no usage reports. The runtime serves requests and nothing else.
Signed builds
Builds are code-signed, so the operating system verifies them before anything runs.
Checks its own files
Every installed file is recorded with a checksum and verified, so a damaged or partial install is reported rather than failing strangely later.
What crosses the wire
- Requests
- From your systems to the address you gave the runtime, and no further
- Weights
- Delivered once, then resident on your hardware
- Answers
- Returned on the same connection
- To us
- Nothing
Where the endpoint sits, who may reach it and how it is logged are yours to decide — the runtime does not take that out of your hands.
Put it on your own GPU.
What ships is the weight pack and the runtime that serves it, with its own private dependencies, so it does not disturb what is already on the machine.
On a workstation
The build isn't published yet. This is where it will be, with its SHA-256 checksum.
- GPU
- NVIDIA RTX with 12 GB of video memory or more (RTX 3060 12 GB, 4070 SUPER, 4080, 4090, 5000-series)
- Memory
- 32 GB system RAM recommended
- Disk
- 25 GB free, SSD recommended
- Driver
- A recent NVIDIA driver
In your own account
One endpoint per GPU, inside your network or your cloud account, speaking the OpenAI-compatible API. Serving one model across several GPUs is still on the roadmap.
- GPU
- 16 GB or more — T4, L4, L40S, A100, H100
- Endpoint
- Chat completions, streaming and tool calls
- Apple
- Apple silicon is in development; NVIDIA is what runs today