Control your AI — when it runs, where it runs.A private AI that only exists while it thinks.

Uno GPU runs your model, your weights, your fine-tunes on a dedicated slice that wakes in under a second and sleeps for free. The privacy of dedicated hardware, billed like serverless.

Talk to us from $199/mo · early access
slice/qwen3-27b-yourco — request trace right now: asleep
00.00 POST /v1/chat/completions hits your endpoint. Slice asleep at $0/h
00.02 Snapshot streams into VRAM: 27 GB
00.60 First token. Meter starts
04.10 Response complete. Meter stops: 3.5 s billed
05.30 Idle. Slice snapshots itself back to sleep: $0/h again
OpenAI-compatible: same SDK, new base URL. 0.6 s to first token, measured on Qwen3 27B. Our hardware, not a projection.

The middle ground that didn't exist.

Control or economics: until now you picked one. GPU pods bill around the clock. Serverless GPUs cold-start in 5–30 seconds. Token APIs read your prompts. We took the empty corner.

the corner nobody built you pay only while it thinks privacy & control Dedicated server Hetzner, OVH bare metal · rented by the month GPU pods & VMs Nebius, RunPod pods · your stack, billed 24/7 Serverless GPU Modal, RunPod serverless · cold start 5–30 s Token APIs OpenAI, DeepSeek · provider reads your prompts Uno GPU your VM · awake in 0.6 s

Just need cheap tokens with no privacy constraints? Use a token API, honestly. Uno GPU is for the work that can't go there.

The frontier of open weights, awake in 0.6 s.

The model

Qwen3 27B: open weights at near-frontier intelligence. Agents, coding, tool use — the model you'd actually build on, not a demo-sized 7B. Or bring your own: any weights up to 35 GB, including fine-tunes of it.

Why 0.6 seconds

The model's whole state lives as a snapshot next to the GPU. Waking streams it straight into VRAM: a memory copy, not a boot. No image pull, no weight load, no warm-up — generation resumes mid-state.

stream 27 GB → VRAM · 0.58 s resume · 0.02 s → first token at 0.60 s

Not an endpoint. A machine.

Uno GPU is a full Linux VM with root and a GPU slice attached. Run vLLM, your own stack, whatever you want.

A real VM, yours

SSH in as root. Install anything, run anything: your inference server, your tools, your daemons. It's a server, not a sandbox.

A home for an agent, memory included

Spawn an agent with its own private or custom LLM. It works, falls asleep, and wakes exactly where it was: processes, files, context intact. One API call.

Proven on Uno boxes

Snapshot sleep and wake runs in production on Uno Cloud today. The GPU slice rides the same mechanism.

root@slice-01 — your machine
$ ssh root@slice-01.yourco.uno root@slice-01:~# nvidia-smi H200 slice · 35 GB VRAM — yours root@slice-01:~# ./agent --model ./weights/qwen3-27b-ft ✓ agent up · thinking # goes idle → slice sleeps itself · $0/h $ curl -X POST api.uno4.dev/slices/01/wake ✓ resumed mid-thought · 0.6 s

What it's for.

Experiments with custom weights

Park a shelf of models: this Qwen, that Qwen, your fine-tune. Swap the active one in a second, all under your control.

Teams that need a private contour

Your perimeter, your network, your VM. Inference that never leaves the contour you control.

Personal work: rare but private

A model you call twice a week costs parking, not rent. It waits asleep at $0/hour.

Compliance-sensitive workloads

Single-tenant VM, no-log, EU jurisdiction, DPA. The paperwork matches the wiring.

Built for work that can't leak.

Wiring, not a policy page.

network
Endpoint inside your WireGuard mesh. Prompts never touch the public internet.
logs
No-log by design. Prompts and outputs aren't stored. Nothing to leak, subpoena, or train on.
tenancy
Single-tenant. Your model, your VM, your slice. No strangers in the queue.
jurisdiction
Netherlands hardware, DPA on request. You know which country your data thinks in.

Run any model that fits.

Park a shelf of models and swap the active one in a second. And it's a shelf of your models: your fine-tune, your quant, any repo on Hugging Face — up to 35 GB a slice.

hf.co/

Paste a link, we do the rest: weights, quantization, chat template, endpoint. White-glove while we're in early access.

or start from a template
ModelGood forWake
Qwen3 27B FP8frontier-class agents, coding, chat0.6 smeasured
Qwen3 8B FP8routing, extraction, classification0.3 smeasured
Qwen3 VL 30BOCR, vision, document parsing~0.7 s
Whisper large-v3transcription, calls, subtitles~0.3 s
XTTS / CosyVoicevoice bots, narration~0.3 s
FLUX.1 dev FP8product images, avatars, creatives~0.5 s
BGE-M3 + rerankerembeddings, RAG search~0.1 s

A full AI backend (LLM, speech, images, embeddings) parks together and wakes piece by piece. Your VM, your call: custom weights, fine-tunes, or models the big platforms won't host. We don't inspect what you load.

From $199a month, all in.

A fixed monthly pool of thinking minutes, sized to your workload at onboarding. Plus parking for the models you keep.

While the model generatesmetered, from the pool
Waking, waiting, idle, parked$0
Parking your weightsflat fee, sized at onboarding
When the pool runs outit stops, never a surprise invoice
Talk to us

Tell us what you'd run.

Early access is hand-picked. We onboard a few teams at a time and set up your models with you.

Two lines is enough. What model, what for, roughly how many people.
Straight to the founders. This lands in our Telegram, not a CRM queue.
Fast answer. Usually within the hour, always within a day.

No mailing lists, no follow-up sequences. One human reply.

delivered

Got it — it's in our Telegram.

We'll get back to you shortly. Impatient? Ping @get_uno_support.