Uno GPU runs your model, your weights, your fine-tunes on a dedicated slice that wakes in under a second and sleeps for free. The privacy of dedicated hardware, billed like serverless.
POST /v1/chat/completions hits your endpoint. Slice asleep at $0/h
Control or economics: until now you picked one. GPU pods bill around the clock. Serverless GPUs cold-start in 5–30 seconds. Token APIs read your prompts. We took the empty corner.
Just need cheap tokens with no privacy constraints? Use a token API, honestly. Uno GPU is for the work that can't go there.
Qwen3 27B: open weights at near-frontier intelligence. Agents, coding, tool use — the model you'd actually build on, not a demo-sized 7B. Or bring your own: any weights up to 35 GB, including fine-tunes of it.
The model's whole state lives as a snapshot next to the GPU. Waking streams it straight into VRAM: a memory copy, not a boot. No image pull, no weight load, no warm-up — generation resumes mid-state.
Uno GPU is a full Linux VM with root and a GPU slice attached. Run vLLM, your own stack, whatever you want.
SSH in as root. Install anything, run anything: your inference server, your tools, your daemons. It's a server, not a sandbox.
Spawn an agent with its own private or custom LLM. It works, falls asleep, and wakes exactly where it was: processes, files, context intact. One API call.
Snapshot sleep and wake runs in production on Uno Cloud today. The GPU slice rides the same mechanism.
Park a shelf of models: this Qwen, that Qwen, your fine-tune. Swap the active one in a second, all under your control.
Your perimeter, your network, your VM. Inference that never leaves the contour you control.
A model you call twice a week costs parking, not rent. It waits asleep at $0/hour.
Single-tenant VM, no-log, EU jurisdiction, DPA. The paperwork matches the wiring.
Wiring, not a policy page.
Park a shelf of models and swap the active one in a second. And it's a shelf of your models: your fine-tune, your quant, any repo on Hugging Face — up to 35 GB a slice.
Paste a link, we do the rest: weights, quantization, chat template, endpoint. White-glove while we're in early access.
| Model | Good for | Wake |
|---|---|---|
| Qwen3 27B FP8 | frontier-class agents, coding, chat | 0.6 smeasured |
| Qwen3 8B FP8 | routing, extraction, classification | 0.3 smeasured |
| Qwen3 VL 30B | OCR, vision, document parsing | ~0.7 s |
| Whisper large-v3 | transcription, calls, subtitles | ~0.3 s |
| XTTS / CosyVoice | voice bots, narration | ~0.3 s |
| FLUX.1 dev FP8 | product images, avatars, creatives | ~0.5 s |
| BGE-M3 + reranker | embeddings, RAG search | ~0.1 s |
A full AI backend (LLM, speech, images, embeddings) parks together and wakes piece by piece. Your VM, your call: custom weights, fine-tunes, or models the big platforms won't host. We don't inspect what you load.
A fixed monthly pool of thinking minutes, sized to your workload at onboarding. Plus parking for the models you keep.
Early access is hand-picked. We onboard a few teams at a time and set up your models with you.