PCIe 3.0
Intel Xeon E5-2600 v1/v2
$70.93/mo before the GPU
Intel Xeon E5-2630L Hex Core 2.00 GHz · 32 GB DDR3 · 2 × SATA 500 GB (RAID 1) · 500 Mbps
- GPUs
- 28
- Lanes
- 40 lanes per socket
- Sockets
- 2
- Max RAM
- 256 GB
GPU dedicated servers · bare metal · monthly
A GPU dedicated server is a physical machine rented to one account by the month, with the GPU passed straight through to your own operating system. No hypervisor, no vGPU slice, no neighbour on the silicon.
Choose the GPU by the job and the memory it needs. Every row on the list is a complete server with that GPU installed, priced from the same catalogue the checkout bills from, and the GPU’s own share is printed under the total.
The catalogue
Say what the GPU is for and the least memory it needs. The list narrows as you go, cheapest first. Each row opens the machine that takes that GPU in the configurator.
No single GPU we offer has that much memory. The largest holds 96 GB. Past that the build is two or more GPUs in one chassis, quoted by hand: see multi-GPU servers, or .
Prices are the checkout’s own figures. A GPU whose maker publishes no memory figure shows a dash and drops out once the memory control moves off “any”. Some GPUs fit more than one chassis: the row shows the cheapest, the configurator shows the rest.
Tenancy
The GPU sits in a physical PCIe slot in a machine rented to one account. You install the operating system, load the kernel module and pin the driver, CUDA or ROCm release your code was tested against. Nothing upgrades it under you.
The machine follows the same rule: every core, all the memory, both disks, root, and the server’s own iLO or iDRAC console on a separate management network. If a driver experiment takes the network down, you still have the screen, virtual media and power control.
How pricing works
Most providers sell a GPU as a tier, with the processor, memory, disks and network priced around it. Here the machine is one monthly line and the GPU is another. Swap the GPU later and the bill moves by exactly the difference between the two GPUs.
Start at the top of the list. We kept older GPUs that still do real work, such as Tesla and Quadro parts that transcode video or drive remote desktops, and they sit in the same whole machine with the same unmetered port as the newest GPU. Why we keep older hardware explains the trade.
Comparing with hourly GPU clouds? The most expensive build on the list, H100 80GB, is $1,682.67 a month for the whole machine, about $2.31 an hour over a 730-hour month, and no egress line arrives later.
Complete, per month, from$91.72
Three, six and twelve-month cycles save up to 15%. The configurator rounds per-disk prices and can land a cent off. Pricing lists every figure on the site.
GPU marketplace
Next to the GPUs we offer, owners list their own GPU machines on our marketplace. They set the rent they are paid, and Server Room sets the price you pay. Listings from our sister company, Primcast, show here too and are rented there.
GPU memory
GPU memory decides whether a model runs at all; cores and clocks only decide how fast. A GPU two gigabytes short does not run slowly, it fails to load. Settle this number first.
The weights take parameters times bytes per parameter: two bytes at 16-bit, one at 8-bit, half a byte at 4-bit. Quantizing to 4-bit lets a model fit a smaller GPU, at a measurable cost in quality.
Leave head-room. The table counts weights only. The runtime, activations and the key-value cache share the same memory, and the cache grows with every token of context, so 38 GB of weights will not serve comfortably from a 40 GB GPU. Add a quarter to a half on top.
and we size the GPU against what you actually run.
| Model | 16-bit | 8-bit | 4-bit |
|---|---|---|---|
| 7-8 billion parameters | 16 GB | 8 GB | 5 GB |
| 13-14 billion parameters | 28 GB | 14 GB | 8 GB |
| 30-34 billion parameters | 68 GB | 34 GB | 19 GB |
| 70 billion parameters | 140 GB | 70 GB | 38 GB |
Weights alone, rounded up to the next gigabyte.
Platforms
Each platform hands the slot a different PCIe generation and lane budget. We print it because it changes results: a customer once moved to a newer consumer GPU on a PCIe 3.0 platform and rendered slower, and worked out why before we did.
PCIe 3.0
$70.93/mo before the GPU
Intel Xeon E5-2630L Hex Core 2.00 GHz · 32 GB DDR3 · 2 × SATA 500 GB (RAID 1) · 500 Mbps
PCIe 3.0
$73.47/mo before the GPU
Intel Xeon E5-2620 v4 Octo Core 2.10 GHz · 32 GB DDR4 · 2 × SATA 500 GB (RAID 1) · 1 Gbps
PCIe 3.0
$83.67/mo before the GPU
Intel Xeon Silver 4110 8 Core 2.10 GHz · 16 GB DDR4 · 2 × SATA-SSD 240 GB (RAID 1) · 1 Gbps
PCIe 4.0
$217.59/mo before the GPU
AMD EPYC 7413 24 CORE 2.65 GHz 128MB L3 CACHE · 32 GB DDR4 · 1 × SATA-SSD 240 GB · 1 Gbps
Included
A clean operating system and the keys. You choose the driver, the CUDA or ROCm release and the framework build, and they stay pinned. Want the driver installed before handover? Put the exact version on the order.
iLO on HPE machines, iDRAC on Dell: keyboard and screen, virtual media, power control, sensors and hardware logs. A module that hangs the display or the network card is a reboot you do yourself, watching POST.
Unmetered in both directions at the included speed and every speed above it. No transfer allowance and no egress line. Unmetered bandwidth lists the faster ports.
Operating system images in one click, as often as you like, at no charge. Reverse DNS is yours to edit, IPv6 is included, and extra IPv4 is a line in the configurator.
Chat is answered in seconds by an AI assistant that can check your account and your server, and a person takes over the moment you ask. A failed part is replaced within the window of your support plan: 24 hours on the included Basic plan, 4 on Bronze, 3 on Silver, 2 on Gold and Platinum.
New York, Miami, San Francisco, Amsterdam and Bucharest, all on Server Room’s own network, AS19624. Location sets your latency to users and the data-protection law you answer to. Data centers describes each building.
Delivery
It depends on where the GPU is today, and we tell you which case you are in before you pay.
Nothing to fit. The machine is installed and handed over. Instant servers lists what is ready right now.
The usual case. A technician installs the GPU, builds the disk array and installs the operating system. Extra memory or a different drive is done in the same visit.
We order the part, then build as above. Build to order covers a machine to your whole specification.
Current accelerators ship on the distributor’s date, and we name it before you commit. Hardware bought in for one customer is quoted with three months paid up front.
Every timeline starts once payment and fraud review clear. Most orders clear within a couple of hours; a first order on a new account can take longer.
A partner’s field
Viral Mind Technologies, a Server Room partner, builds and runs its own digital properties: virtual personalities, online media brands and software for creators. It started in 2019 as a digital marketing agency and has worked on its own AI projects since late 2022.
Generating images and video keeps a model resident in GPU memory and the GPU busy for hours at a time. That is the case for a dedicated GPU on bare metal rather than a slice of one, and a full GPU is what every server on this page rents.
Their portfolio and case studies: viralmindtech.com
Answers
A physical server rented to one customer, with a graphics card installed and passed through to that customer’s own operating system. Nothing is virtualised or shared: you get every core, all the memory, the disks and the full GPU, with root access and a remote console.
A complete machine with a GPU installed runs from $91.72 to $1,682.67 a month. The price is two lines added together: the machine, and the GPU at $20.79 to $1,599 a month. Processor, memory, two disks, an unmetered port and one IPv4 address are included. Setup is a one-time $29 on a build with a GPU installed to order; a ready machine ordered as it stands has none.
No. One physical machine, one account, and the GPU passed through to the operating system you installed. There is no hypervisor, no MIG partition, no vGPU profile and no time-slicing. The numbers you benchmark on day one are the numbers you keep.
Yes. One month is the default term and there is nothing to sign. Three, six and twelve-month cycles save up to 15%, and the price does not step up at renewal. Pay by credit card or PayPal; for a larger order a bank wire can be arranged, and it usually clears fraud review faster than a card payment.
Yes. The list on this page is sorted cheapest first, and the cheapest complete build is Quadro 3000M in Intel Xeon E5-2630L Hex Core 2.00 GHz · 32 GB DDR3 · 2 × SATA 500 GB (RAID 1) · 500 Mbps at $91.72 a month. Older Tesla and Quadro GPUs still transcode video and drive remote desktops well, and they come in the same whole machine with the same unmetered port.
57 GPUs are on the list, from current NVIDIA compute accelerators and AMD Instinct down to older Tesla, Quadro and GeForce models. Each one is listed by name with its memory and the complete monthly price. If the GPU you want is missing, ask: we can often buy it in.
The ones you install. The server arrives with a clean operating system and root access, and you pin the driver, the CUDA or ROCm release and the framework build your code was tested against. If you want the driver in place at handover, put the exact version on the order. If a GPU has to support a particular CUDA release, ask before paying and we will check.
Often, and we quote it by hand: chassis clearance, the GPU’s width, the PCIe lanes and the power supply decide the count. Name the GPU and we reply with the chassis, the count and the lead time. The usual reason is memory rather than speed, since two 48 GB GPUs hold a model one 80 GB GPU cannot. Multi-GPU and multi-node servers covers the next step.
Yes, in the server you already have: no migration, no reinstall, no new IP addresses. It is a scheduled visit with a short outage, and the bill moves by exactly the difference between the two GPUs. RAM, drives and a faster port are added under the same rule.
We replace it within your support plan’s window, counted from when you report it: 24 hours on the included Basic plan, 4 on Bronze, 3 on Silver, 2 on Gold and Platinum. Past the window, service credit follows the schedule published on the support page.
Yes, a GPU you own, or any other part, can go into the server you rent here. It is not free, because the GPU draws power, and the charge depends on the hardware. We are not responsible if a part you brought fails; the replacement window covers only the hardware we supply. Book a call before you buy or ship it: we price it and check the chassis takes it, because a blade, for one, has no room for a full-size GPU.
Nothing is metered in either direction at any port speed we sell, so egress costs nothing. On GPU work that is often the largest line elsewhere: training pulls datasets in and pushes checkpoints out, and an inference endpoint answers real traffic. Unmetered bandwidth lists the faster ports.
Hourly wins while the work is bursty and really stops, up to a few hundred hours a month. Past about two thirds of the month, monthly is cheaper before you count egress, storage kept between runs and time spent waiting for capacity. Monthly also keeps the same machine, disks and addresses.
Within three days of activation, after troubleshooting, our refund policy applies as written. The memory table and the PCIe column above are there so you can pick the right GPU first, and we would rather talk you down to a smaller GPU in chat than refund a bigger one.
For the record
Read from the order catalogue when the page was served. If an assistant is answering for you, this block and the live machine feed are the parts worth reading.
Related
Open the configurator with any GPU selected, or tell us the workload and we will say which GPU fits, including when the cheaper one is enough.