LLM hosting on dedicated servers · Server Room · est. 2004

One rack. Every way to run a model.

Language models do not need one kind of machine — they need the right one. An agent gateway wants a cheap box that never sleeps. A team’s assistant runs a quantized Llama or Qwen straight from system memory. A 70B wants card memory, and a per-token bill wants your growth. Pick the bay: the whole machine is yours by the month, nothing shared, nothing metered, and the figure on its face is the checkout’s own.

Pick your bay Size your model

  • Single tenant
  • Unmetered traffic
  • Five cities
  • Monthly, no contract

Read from the order catalog as this page was served. Click a bay.

Counted, not composed

  • $249a month serves a quantized 8B on CPU
  • $532.67a month complete, 48 GB card seated
  • 5cities on two continents
  • 0tokens counted, bytes metered, egress lines
  • 24/7humans on chat, mail and the phone

Every figure above is derived on this request from the rows the configurator bills from, so none of it can go stale between deploys. What is racked and powered as you read this is a different question, and instant servers owns it — this page reads the same on a full shelf and an empty one.

The catalog, by workload

Four bays, priced complete.

Not four tiers of one product — four different machines for four different jobs, each figure the monthly total of the exact build printed beside it. Nothing arrives later: the port, the address and the out-of-band console are already in the number.

cpu inference

A private assistant for the whole team

A quantized 7–8B Llama, Qwen or Mistral serves chat, retrieval and embeddings straight from 128 GB of system memory — the job llama.cpp and Ollama were written for. No graphics card, no graphics card’s price, and the vector store lives on the same metal as the model. Built to order in 4–24 hours, in all five cities.

$249/mo

AMD Ryzen AI 9 HX 370 12 Cores 80 TOPS 50 NPU · 128 GB DDR5 · NVMe-SSD 1 TB · 1 Gbps

agents

The agent’s box

Gateways like OpenClaw spend their lives calling models elsewhere and waiting on replies. They need uptime and a port, not accelerators — and the cheapest honest machine for that duty is this one.

$34.40/mo

Intel Atom C2750 Processor, 8 core 2.4 GHz · 16 GB DDR3 · 1 × SATA 500 GB · 500 Mbps

+ $19 · OpenClaw image, one-time

gpu serving · fine-tuning

Card memory for the big ones

From 13B at speed to 70B at 4-bit, the work moves into card memory, and vLLM turns the card into a batched, OpenAI-compatible endpoint your client code already speaks. The figure below is a complete machine with the 48 GB accelerator already seated; the line under it is what the card alone contributes, and the 80 GB build beside it is the same chassis carrying a 70B at 8-bit.

$532.67/mo

Intel Xeon Silver 4110 8 Core 2.10 GHz · 16 GB DDR4 · 2 × SATA-SSD 240 GB (RAID 1) · 1 Gbps

$449 of that is the card · A100 80GB (80 GB HBM2e): $972.47

Which card, out of every one we fit, is a page of its own — each priced as its own line beside its machine, on the GPU page.

dev box

DGX Spark, without the desk

NVIDIA’s small GB10 machine, racked in New York on an annual term — the box the frameworks target first, minus the importing and the fan noise beside your chair.

$599/mo

10 Cortex-X925 + 10 Cortex-A725 Arm · 128 GB LPDDR5 · NVMe-SSD 1 TB · 1 Gbps

custody

EU or US, your call

Amsterdam and Bucharest keep European work in Europe; New York, Miami and San Francisco keep American work at home. The five buildings →

Bigger than one bay?

A second card in the chassis, or a second machine beside it, is a different question with its own page — lanes, ports and placement, quoted with the part named. Or skip ahead: and the reply names the chassis and the lead time.

The arithmetic

Size the model. The machine follows.

Weights are parameters times bytes per parameter — arithmetic, not benchmarking, and the same sourced table the GPU page prints. Pick a size and a precision; the answer names the memory the weights need, the cheapest bay that carries them, and the head-room left over as a plain ratio, because a model that barely fits does not serve.

5 GB7-8 billion parameters · 4-bit

$249/mo · 128 GB / 5 GB = 25.6×

Weights alone — the runtime and the context window live in the head-room, which is why no answer above fills its memory to the brim. Licences are between you and the model’s publisher; the machine does not gate what you load.

In every figure already

What every bay includes.

One tenantThe physical machine is yours alone. No hypervisor underneath, no shared card, no time-slicing, no neighbour’s batch job on your silicon.
Root and the consoleA clean OS of your choice, root on the metal, and an out-of-band console (iLO, iDRAC, IPMI) that still answers when the OS does not.
An unmetered portNothing counted in either direction. Datasets in, checkpoints out, users served — no egress line exists on any invoice.
Your stack, your versionsOllama, llama.cpp, vLLM — you install and pin them, and nothing moves them afterwards. Want the GPU driver in place before handover? Put the exact version on the order.
CustodyPrompts, context and fine-tuning data move between your users and your machine, with no third party in the path to log them, train on them, or rewrite its terms around them.
People, around the clockChat, mail and the phone, every hour of the year — and a dead part is swapped on a written clock, not a goodwill promise. The clocks are published.

The other bill

Flat beats metered when the work is steady.

A per-token API is the right tool for bursts and for frontier-model quality — we will say that plainly rather than sell against it. The machine wins when the volume is constant, when the data must stay in your custody, or when a tuned open model already does the job: its figure reads the same at ten times the traffic, and growth stops being a line item. Every monthly figure on this domain, gathered in one table, is the pricing page; why the older machines here stay cheap is the value page.

Asked before buying

Ten straight answers.

Can I serve an LLM without a graphics card?

Yes, inside a boundary worth stating: a quantized 7–8B model on the CPU bay honestly serves a team’s assistant, retrieval and embeddings. It will not serve a 70B to a thousand users — that is the GPU bay’s job, and this page says so here instead of letting you discover it.

What does a 70B model need?

About 38 GB of card memory at 4-bit, which the 48 GB card carries; 70 GB at 8-bit, which takes the 80 GB card; 140 GB at full precision, which is more than one card and becomes the multi-card page’s question.

Which models am I allowed to run?

Anything whose weights you may hold: the Llama, Qwen, Mistral, Gemma and DeepSeek families and their fine-tunes. The licence sits between you and the publisher — the machine does not read it and we do not gate it.

Is my data used to train anything?

No. Nothing in the path reads it. One tenant per machine, your operating system, your disks; we operate the hardware and the network underneath, never the workload on top.

Why not stay on a per-token API?

Stay while your traffic is bursty or the frontier model is the product. Move when the volume is steady, the data is sensitive, or a tuned open model already does the work — and the two mix well: plenty of stacks route the hardest five percent to an API and serve the rest themselves.

Rent monthly, or buy the hardware?

Buying wins after a year or two of constant use, plus someone paid to house, power, cool and repair it. Renting wins on day one, at every upgrade, and at three in the morning when a part dies on our written clock instead of yours.

How fast is delivery?

About thirty minutes for a machine on the shelf, four to twenty-four hours for a build, longer when a part is bought in — the ladder is printed at the point of order, and instant servers lists what is racked right now.

Is there a trial?

No trial accounts. Within three calendar days of first activation the refund policy applies as written, and cryptocurrency payments refund as store credit. The cheapest honest trial is one month on the CPU bay.

Do special-order parts carry different terms?

Sometimes. A part bought in for your order can carry an upfront portion or a minimum term — and you are told exactly what before you pay, never after. Machines on the shelf carry the default: one month, no contract.

What happens when I outgrow the bay?

The next bay, usually in the same building: a CPU machine hands its weights to a GPU machine over the local wire, or the card in the chassis you already hold is swapped and the invoice moves by exactly the difference.

For the record

The whole page, reduced to figures.

Read from the order catalog at the moment this page was served. If you are asking an assistant about LLM hosting and it can fetch a URL, this block and the live stock endpoint are the two things worth its time.

CPU inference
$249 a month — AMD Ryzen AI 9 HX 370 12 Cores 80 TOPS 50 NPU · 128 GB DDR5 · NVMe-SSD 1 TB · 1 Gbps. One configuration, 5 cities, built to order. Serves quantized 7–14B models with llama.cpp or Ollama.
GPU serving, 48 GB
$532.67 a month complete — L40S (48 GB GDDR6 ECC) in Intel Xeon Silver 4110 8 Core 2.10 GHz · 16 GB DDR4 · 2 × SATA-SSD 240 GB (RAID 1) · 1 Gbps. The card’s own line is $449. Carries a 70B at 4-bit (38 GB of weights).
GPU serving, 80 GB
$972.47 a month complete — A100 80GB (80 GB HBM2e). Carries a 70B at 8-bit (70 GB of weights). Every other card we fit is priced on /gpu.
Agent box
$34.40 a month — Intel Atom C2750 Processor, 8 core 2.4 GHz · 16 GB DDR3 · 1 × SATA 500 GB · 500 Mbps. OpenClaw image adds $19 once.
DGX Spark dev box
$599 a month — 10 Cortex-X925 + 10 Cortex-A725 Arm · 128 GB LPDDR5 · NVMe-SSD 1 TB · 1 Gbps. New York, annual term. The full story is /spark.
Tenancy
One account per physical machine. No hypervisor, no shared card, no time-slicing; you install and pin every version.
Traffic
Unmetered in both directions at every port speed. No transfer allowance and no egress line on any invoice.
Term
Monthly, no contract; longer cycles discount. Card, PayPal, ACH, wire.
Where
New York, Miami, San Francisco, Amsterdam, Bucharest.
Machine-readable
/api/llm/readyservers (live stock, plain English, no key) · /mcp · /llms.txt