More than one card, more than one machine · Server Room · est. 2004

Scaling up rides PCIe. Scaling out rides the port.

A single card in a single machine is settled arithmetic, priced complete on the GPU page. Past one is where the engineering starts: the next card contends for lanes and watts inside the chassis, the next machine talks through a port that is bought per node, and the platform that maxes one of those axes bottoms out on the other. Our catalog states that trade in numbers, and this page prints them.

See the lane-versus-port trade Price the fabric

Single tenant on every chassis, PCIe passthrough on every card, and no byte counter on any port — traffic between your nodes is nobody’s line item.

Counted, not composed

  • 4Platforms that seat a card
  • 100 GbpsTop port speed, per node
  • 2Socket ceiling, one chassis
  • 1 TBRAM ceiling, one chassis
  • 2Cities stocking all of them

Derived from the order catalog on this request — the same rows the checkout bills from, so none of it can be stale. What is physically racked and powered as you read this is a separate question, and instant servers keeps that list.

First, name the bottleneck

A second card and a second machine cure different shortages.

“We need more compute” is not yet a specification. Whether the shortage sits inside the chassis or across the rack decides what gets ordered, how the money runs, and which ceiling you meet next — so the page splits here, the same way the order will.

Inside the chassis

Add a card to the machine you have

The symptom is a starved accelerator or an evicted working set: batches queue on the bus, or the model has outgrown the card’s memory. The cure is a second card — and how many the chassis seats is set by slot clearance, the per-socket lane budget and the power supply’s headroom, three numbers that belong to a specific chassis holding a specific card rather than to any catalog row. We quote it with the part named instead of selling a checkbox.

What decides it: the chassis. Name the card, and the quote names the chassis that seats that many of it, and the date.

What two cards are not: one big card. Their memory pools stay separate unless your runtime splits the work across them — tensor or pipeline parallel for one large job, which pays a toll on the link between the cards. Work that is already many independent jobs simply runs on twice the silicon, unchanged.

Across the rack

Stand a second machine beside it

The symptom is a full machine: sockets busy, DIMM slots populated, jobs queueing behind jobs. The best any card-taking platform here offers is 2 sockets and 1 TB, and past that wall no bigger single box exists in this catalog — the next unit of capacity is another server, not a larger one.

What decides it: the building and the port. Node-to-node traffic is only useful inside one facility, and each node buys its own port — the fabric bill is the port price times the node count, every month.

What a cluster is here: n separate servers. Ordered separately, invoiced separately, upgraded and replaced separately. No cluster SKU exists, which also means no cluster premium and no minimum node count exist either.

The trade, measured

The widest bus and the fastest port are on different machines.

Every platform we will seat a card in, on the axes that size a cluster rather than a server. The bus generations are the same sourced table the GPU page prints — one file, two pages, no way to disagree — and everything else is the order catalog speaking for itself.

Live from the order catalog on this request. Follow any row across: the bus column and the port column pull against each other.
PlatformPCIe to the cardPort ceilingSocketsRAM ceilingCities
Intel Xeon Silver / Gold48 lanes per socketPCIe 3.0100 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v3/v440 lanes per socketPCIe 3.040 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v1/v240 lanes per socketPCIe 3.020 Gbps1 Gbps included2256 GB3
AMD EPYC128 lanes per socketPCIe 4.020 Gbps1 Gbps included21 TB2

How an engineer reads that table

  • PCIe lanes bound host-to-card transfer — batch streaming, scene loads, anything spilling out of card memory. A model that is resident and staying resident stops caring about them.
  • The port ceiling bounds node-to-node transfer. At one machine it is irrelevant; at two it is the whole problem.
  • Cities is the column to check first for a cluster, not last: a platform that is perfect and absent from your building is not a candidate.
  • Therefore: bus-fed single-node work belongs on the widest lanes and can ignore the port; partitioned work belongs where the port tops out and can tolerate an older bus. A build demanding both maxima at once does not exist in this catalog, and we print that here rather than let you find it out at invoice time.

Which silicon goes in the slot, and each card’s own monthly line, is GPU dedicated servers — the one page on this domain permitted to price a card. Nothing below repeats a figure from it.

The fabric bill

Multiply the port by n before you call it a budget.

A port is priced per machine, so the interconnect under a cluster is the same line item once per node. The table carries that multiplication already done, because budgets composed from single-server prices are how cluster quotes come back as surprises.

Monthly, per machine, on top of the machine itself. A cluster buys its speed once per node — the multiplied columns are the fabric bill.
SpeedPer node× 2 nodes× 4 nodesReaches it
1 Gbpsincludedincludedincludedall 4 platforms
2 Gbps$189$378$756all 4 platforms
3 Gbps$299$598$1,196all 4 platforms
5 Gbps$459$918$1,836all 4 platforms
10 Gbps$529$1,058$2,116all 4 platforms
20 Gbps$1,629$3,258$6,516all 4 platforms
40 Gbps$2,799$5,598$11,1962 of 4 platforms
100 Gbps$3,999$7,998$15,9961 of 4 platforms

Worth noticing in that ladder

  • The rungs are not priced linearly. At several points, n machines one rung down cost less than n−1 machines one rung up — and buy you a whole extra server. Whether that trade is available to you depends only on whether the job partitions.
  • The right-hand column thins toward the top. The highest speeds exist on fewer of the platforms, and not on the one with the widest bus — the platform table’s argument, read from the network side.

What is not in the catalog, stated rather than implied

Between two of our machines the listed link is an Ethernet port. There is no InfiniBand line item and no card-to-card bridge on any order form, and a page about clusters that writes around that fact is misleading you with layout. If a build needs either, ask before ordering: the answer is a quote, and sometimes the answer is no — both are cheaper to hear first.

Every rung is unmetered in both directions, with no allowance and no egress line anywhere on an invoice — replication, shuffles and checkpoints between your nodes cost the port and nothing further. For one port on one machine, judged on its own merits, unmetered ports is the page.

Placement

A cluster is placed into a building before it is configured.

Latency makes the facility a hard constraint: nodes that share a job share a room. Our five rooms do not stock the same platforms, so an engineer who picks the platform first can pick one their city does not carry — check the room, then configure. The grid below is derived, so a platform reaching a new city appears here with nobody editing anything.

  • New York, US4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Bucharest, EU4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Miami, US3 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2
  • San Francisco, US2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4
  • Amsterdam, EU2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4

The buildings themselves — power, carriers, and who answers the phone in each — are on data centers. If the latency between two of our cities matters to your pair of nodes, ask support to measure it between those rooms; a number beats a map.

From order to racked

The date arrives with the quote, not after the payment.

A multi-node build is several machines and, often, a purchase order’s worth of parts — so what gets published is not a delivery slogan but the lead-time register build to order and the GPU page already use, with the long pole identified before you commit to anything.

  1. ~30 minutes

    Racked already, taken as-is

    No fitting and no rack visit — the machines exist and get handed over. Instant servers is that list, and for the plain nodes of a cluster it is the first place worth reading.

  2. 4–24 hours

    Fitted from our own stock

    The standard case: cards seated, arrays built, operating systems installed — each node worked in parallel rather than queued behind its neighbour.

  3. 5–10 working days

    Bought in for the build

    A part on the specification is not one we shelve, so it is ordered first. For a multi-card chassis this is the usual rung, and it is usually the chassis rather than the silicon that sets it.

  4. 10–15 working days

    On allocation

    Current-generation accelerators arrive when the distributor says they arrive. We name that dependency up front, and hardware bought in against a single order takes three months in advance — a week-two cancellation would leave a machine built for nobody.

Every rung’s clock starts when payment and fraud review clear — the one step that carries no promised duration. On a large first order, a wire tends to clear review faster than a card does; worth choosing deliberately when a rack visit is waiting on it.

A cluster survives chat better than a form. — how many nodes, whether the job partitions, and what it has to talk to — and the person answering has quoted one before. If one machine covers it, they will tell you that instead of selling you four.

For the record

The whole page, reduced to figures.

Read from the order catalog at the moment this page was served. Card prices are missing on purpose — GPU dedicated servers prints every one of them beside the machine each figure belongs to.

Cards in one chassis
Above one, the count is quoted: slot clearance, the per-socket lane budget and the power envelope belong to a specific chassis holding a specific card, so the honest number needs the part named. Name it, and the quote comes back with the chassis, the count and the date.
Nodes in one cluster
Any number from one up. Each node is a whole physical server with its own invoice and its own term, ordered, upgraded and replaced independently — no cluster SKU exists here, so neither does a cluster premium or a minimum count.
Card-taking platforms
4, and they specialise: AMD EPYC gives a card the widest bus (PCIe 4.0, 128 lanes per socket) while topping out at 20 Gbps of network in 2 of 5 cities; the 100 Gbps ceiling belongs to Intel Xeon Silver / Gold, which feeds a card an older bus. Scaling up and scaling out are therefore different platform choices.
The link between nodes
An Ethernet port on each node — 1 Gbps included on each of them, priced rungs to 100 Gbps — bought per machine per month, so a cluster pays the rung once per node. No InfiniBand line and no card-to-card bridge is in the catalog; anything past the port is quoted, and the answer can be no.
Where one machine stops
2 sockets and 1 TB of RAM, on the best of these platforms. Past either figure the unit of capacity is the next machine, not a bigger one.
Cities
5: New York, Miami, San Francisco, Amsterdam, Bucharest. Nodes that share a job belong in one of them together, and New York and Bucharest stock the complete platform set.
Traffic
Unmetered at every speed, both directions, with no allowance and no egress line — replication, shuffles and checkpoints between your nodes cost the port and nothing further.
Tenancy
One account per physical machine, on every node of the cluster, each card on PCIe passthrough into that account’s own operating system — no hypervisor, no MIG partition, no vGPU profile, no time-slicing.
Delivery
About half an hour for machines racked as-is, four to twenty-four hours fitted from our stock, five to ten working days bought in, ten to fifteen for an allocation part — and a multi-node quote names its long pole before you pay.
Term
Monthly per machine, no contract; three, six and twelve-month cycles discount up to 15%. Hardware bought in against a single order takes three months up front.
Card money
Absent from this page by rule. Every card we seat and its own monthly line are on GPU dedicated servers, each figure printed beside the exact machine it belongs to.