Uluka.Labs
AWS portalBenchmarks
Model + runtime · one GPU

Big models, half the memory.
Matched on benchmarks.

Uluka builds compressed models and the CUDA runtime that serves them. Our 26B mixture-of-experts fits in 8.73 GB, and on four public benchmarks it finished level with the standard 4-bit builds of its class while taking about half their memory — so the same work fits on a smaller GPU, in an account you own.

uluka-runtime — 127.0.0.1:8321UlukaLabs 26B · cuda:0
0.0 tok/s 0 tokens VRAM 8.4 / 12 GB ● nothing left the host
Try:

An example exchange, replayed in your browser. No model runs on this page.

One GPU, one process

The pack loads into video memory once and answers from there. No inference service in the middle, no second machine.

Sparse by design

A router picks 8 of 128 experts in every layer for each token, so a 26B model reads only about 4B parameters per step — and the whole model is 8.73 GB on disk.

The API you already call

An OpenAI-compatible endpoint with streaming and tool calls, at an address you choose.