不止一张卡,不止一台机器 · Server Room · 始于 2004 年

纵向扩展走 PCIe。横向扩展走端口。

单台机器里的单张卡是已经算清的算术,GPU 页面上标着完整价格。超过一张,工程才真正开始:下一张卡要在机箱里争抢通道和功耗,下一台机器要通过按节点购买的端口通信,而把其中一条轴拉到极限的平台,在另一条轴上就会触底。我们的产品目录用数字说明这笔权衡,本页则把这些数字印出来。

看看通道与端口的权衡 给互联网络估价

每台机箱独占租户,每张卡都走 PCIe 直通,任何端口都没有流量计数器 — 节点之间的流量不会成为任何人账单上的条目。

数出来的,不是编出来的

  • 4Platforms that seat a card
  • 100 GbpsTop port speed, per node
  • 2Socket ceiling, one chassis
  • 1 TBRAM ceiling, one chassis
  • 2Cities stocking all of them

源自本次请求时的订单目录 — 与结账计费所用的同一批数据行,所以不可能过期。你读到这句话时实际已上架并通电的机器是另一回事,即时服务器维护着那份清单。

首先,说出瓶颈在哪

第二张卡和第二台机器解决的是不同的短缺。

“我们需要更多算力”还不是一份规格说明。短缺在机箱内部还是横跨机架,决定了要订什么、钱怎么花,以及你下一个撞上的天花板是什么 — 所以本页在这里分岔,和订单将来的分法一样。

机箱内部

给现有机器加一张卡

症状是加速器吃不饱,或者工作集被逐出:批次在总线上排队,或者模型已经超出了卡上显存。解法是第二张卡 — 而机箱能装多少张,取决于插槽间隙、每插座的通道预算和电源的余量,这三个数字属于装着某张具体卡的某台具体机箱,而不是目录里的任何一行。我们报价时会写明部件名称,而不是卖一个勾选框。

决定因素:机箱。说出卡的名字,报价单就会写出能装下那么多张卡的机箱,以及交付日期。

两张卡不是什么:一张大卡。它们的显存池保持独立,除非你的运行时把工作拆分到它们上面 — 单个大任务用张量并行或流水线并行,这要在卡间链路上付出代价。本来就是许多独立任务的工作,则直接在双倍硅片上运行,无需改动。

横跨机架

在旁边再立一台机器

症状是机器已满:插槽繁忙,DIMM 插满,任务在任务后面排队。这里任何能插卡的平台最多提供 2 sockets and 1 TB,超过那道墙,本目录里不存在更大的单机 — 下一个容量单位是另一台服务器,而不是一台更大的。

决定因素:机房和端口。节点间流量只有在一个设施内部才有用,而且每个节点都要买自己的端口 — 互联网络账单就是端口价格乘以节点数,每个月都是。

这里的集群是什么:n 台独立的服务器。分别下单,分别开票,分别升级和更换。不存在集群 SKU,这同时意味着不存在集群溢价,也不存在最低节点数。

权衡,用数字衡量

最宽的总线和最快的端口不在同一台机器上。

我们会为之插卡的每一个平台,都按决定集群规模而非服务器规模的轴来排列。总线代际与 GPU 页面印出的来源表相同 — 一个文件,两个页面,不可能不一致 — 其余一切都由订单目录自己说话。

Live from the order catalog on this request. Follow any row across: the bus column and the port column pull against each other.
PlatformPCIe to the cardPort ceilingSocketsRAM ceilingCities
Intel Xeon Silver / Gold48 lanes per socketPCIe 3.0100 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v3/v440 lanes per socketPCIe 3.040 Gbps1 Gbps included21 TB5
Intel Xeon E5-2600 v1/v240 lanes per socketPCIe 3.020 Gbps1 Gbps included2256 GB3
AMD EPYC128 lanes per socketPCIe 4.020 Gbps1 Gbps included21 TB2

工程师怎么读那张表

  • PCIe 通道约束主机到卡的传输 — 批次流式传输、场景加载、任何溢出卡上显存的内容。一个常驻且保持常驻的模型就不再关心它们。
  • 端口上限约束节点到节点的传输。一台机器时它无关紧要;两台机器时它就是全部问题。
  • 城市是集群要最先查看的列,而不是最后:一个完美却不在你所在机房的平台,不是候选。
  • 因此:总线驱动的单节点工作属于最宽通道,可以忽略端口;分区工作属于端口到顶的地方,可以容忍较老的总线。同时要求两个最大值的构建在本目录中不存在,我们在这里印出这一点,而不是让你在开票时才发现。

插槽里放什么芯片,以及每张卡自己的月度费用,见 GPU 独立服务器 — 本域名上唯一允许给卡标价的页面。下面没有任何内容重复那里的数字。

互联网络账单

在把它叫作预算之前,先把端口乘以 n

端口按机器计价,所以集群底下的互联就是每个节点一次的同一条目。表格已经替你算好了乘法,因为用单服务器价格拼出来的预算,正是集群报价变成惊吓的原因。

Monthly, per machine, on top of the machine itself. A cluster buys its speed once per node — the multiplied columns are the fabric bill.
SpeedPer node× 2 nodes× 4 nodesReaches it
1 Gbpsincludedincludedincludedall 4 platforms
2 Gbps$189$378$756all 4 platforms
3 Gbps$299$598$1,196all 4 platforms
5 Gbps$459$918$1,836all 4 platforms
10 Gbps$529$1,058$2,116all 4 platforms
20 Gbps$1,629$3,258$6,516all 4 platforms
40 Gbps$2,799$5,598$11,1962 of 4 platforms
100 Gbps$3,999$7,998$15,9961 of 4 platforms

那架梯子上值得注意的

  • 梯级不是线性定价的。在好几个点上,低一档的 n 台机器比高一档的 n−1 台更便宜 — 还多给你一整台服务器。这笔权衡是否对你可用,只取决于任务能否分区。
  • 右列越往上越稀。最高速度存在于更少的平台上,而且不在总线最宽的那个平台上 — 平台表的论点,从网络这一侧读出来。

目录里没有的,明说而不是暗示

我们两台机器之间的所列链路是以太网端口。任何订单表单上都没有 InfiniBand 条目,也没有卡到卡的桥接,而一个关于集群的页面如果绕着这个事实写,就是在用排版误导你。如果构建需要其中任何一样,下单前先问:答案是报价单,有时答案是不行 — 两样都先听到更便宜。

每一级都是双向不限流量,账单上任何地方都没有额度也没有出站费用 — 节点之间的复制、混洗和检查点只花端口钱,再无其他。单台机器上的单个端口,按它自身的优劣来评判,见 不限流量端口

放置位置

集群先放进一栋楼,然后才配置。

延迟让设施成为硬约束:共享一个任务的节点共享一个机房。我们的五个机房并不存放相同的平台,所以先挑平台的工程师可能挑到一个他们城市没有的 — 先查机房,再配置。下面的网格是派生出来的,所以平台到达新城市时,这里会自动出现,无需任何人编辑。

  • New York, US4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Bucharest, EU4 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2 · AMD EPYC
  • Miami, US3 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4 · Intel Xeon E5-2600 v1/v2
  • San Francisco, US2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4
  • Amsterdam, EU2 of 4 platformsIntel Xeon Silver / Gold · Intel Xeon E5-2600 v3/v4

这些建筑本身 — 电力、运营商、每个机房谁接电话 — 见 数据中心。如果我们两个城市之间的延迟对你的节点对很重要,请让支持团队在两个机房之间实测;一个数字胜过一张地图。

从下单到上架

日期随报价单一起到,而不是付款之后。

多节点构建是好几台机器,而且往往是一张采购单量的部件 — 所以公布的并不是一句交付口号,而是 按单定制 和 GPU 页面已经在用的交付周期登记表,在你承诺任何事之前就标出最长的那根杆。

  1. 约 30 分钟

    已经上架,按现状交付

    无需装配,无需进机房 — 机器已经存在,直接移交。即时服务器就是那份清单,对于集群里的普通节点,它是第一个值得读的地方。

  2. 4–24 小时

    用我们自己的库存装配

    标准情况:插卡、建阵列、装操作系统 — 每个节点并行处理,而不是排在邻居后面。

  3. 5–10 个工作日

    为这次构建采购

    规格上的某个部件不是我们常备的,所以先下单采购。对于多卡机箱,这是常见的梯级,而且通常决定它的是机箱而不是芯片。

  4. 10–15 个工作日

    按配额分配

    当前一代加速器什么时候到,取决于分销商说什么时候到。我们一开始就点明这个依赖,而且针对单个订单采购的硬件要提前三个月付款 — 第二周取消就会留下一台为谁都不存在的机器。

每一级的时钟从付款和欺诈审核通过时开始 — 这是唯一没有承诺时长的步骤。大额首单时,电汇通常比信用卡更快通过审核;当一次进机房在等它时,值得刻意选择。

集群在聊天里比在表单里活得更久。 — 多少节点、任务能否分区、它要和什么通信 — 接话的人以前报过价。如果一台机器就够,他们会直接告诉你,而不是卖给你四台。

For the record

The whole page, reduced to figures.

Read from the order catalog at the moment this page was served. Card prices are missing on purpose — GPU dedicated servers prints every one of them beside the machine each figure belongs to.

Cards in one chassis
Above one, the count is quoted: slot clearance, the per-socket lane budget and the power envelope belong to a specific chassis holding a specific card, so the honest number needs the part named. Name it, and the quote comes back with the chassis, the count and the date.
Nodes in one cluster
Any number from one up. Each node is a whole physical server with its own invoice and its own term, ordered, upgraded and replaced independently — no cluster SKU exists here, so neither does a cluster premium or a minimum count.
Card-taking platforms
4, and they specialise: AMD EPYC gives a card the widest bus (PCIe 4.0, 128 lanes per socket) while topping out at 20 Gbps of network in 2 of 5 cities; the 100 Gbps ceiling belongs to Intel Xeon Silver / Gold, which feeds a card an older bus. Scaling up and scaling out are therefore different platform choices.
The link between nodes
An Ethernet port on each node — 1 Gbps included on each of them, priced rungs to 100 Gbps — bought per machine per month, so a cluster pays the rung once per node. No InfiniBand line and no card-to-card bridge is in the catalog; anything past the port is quoted, and the answer can be no.
Where one machine stops
2 sockets and 1 TB of RAM, on the best of these platforms. Past either figure the unit of capacity is the next machine, not a bigger one.
Cities
5: New York, Miami, San Francisco, Amsterdam, Bucharest. Nodes that share a job belong in one of them together, and New York and Bucharest stock the complete platform set.
Traffic
Unmetered at every speed, both directions, with no allowance and no egress line — replication, shuffles and checkpoints between your nodes cost the port and nothing further.
Tenancy
One account per physical machine, on every node of the cluster, each card on PCIe passthrough into that account’s own operating system — no hypervisor, no MIG partition, no vGPU profile, no time-slicing.
Delivery
About half an hour for machines racked as-is, four to twenty-four hours fitted from our stock, five to ten working days bought in, ten to fifteen for an allocation part — and a multi-node quote names its long pole before you pay.
Term
Monthly per machine, no contract; three, six and twelve-month cycles discount up to 15%. Hardware bought in against a single order takes three months up front.
Card money
Absent from this page by rule. Every card we seat and its own monthly line are on GPU dedicated servers, each figure printed beside the exact machine it belongs to.