What the models weigh
The files are the specification. ComfyUI loads a checkpoint, or a diffusion model plus its text encoders and VAE, and the card has to hold the part that runs every step. These are the files as published.
| Model | Files | On disk |
|---|---|---|
| SDXL base 1.0 | sd_xl_base_1.0.safetensors | 6.94 GB |
| FLUX.1-dev, fp8 checkpoint | flux1-dev-fp8.safetensors (model, text encoders, VAE) | 17.2 GB |
| FLUX.1-dev, 16-bit | flux1-dev 23.8 GB, t5xxl_fp16 9.79 GB, clip_l 246 MB, ae 335 MB | about 34 GB |
| FLUX.1-dev, 16-bit model, fp8 T5 | the same with t5xxl_fp8_e4m3fn 4.89 GB | about 29 GB |
By default ComfyUI unloads a model to system memory once it has been used, and its README says asynchronous weight streaming runs even the biggest open models on as little as 4 GB of VRAM and 8 GB of RAM. Streaming moves weights across PCIe on every pass and needs that much more system memory. A card that holds the diffusion model whole skips that traffic. --highvram keeps models in GPU memory instead of unloading them, which only works when they fit.
FLUX.1-dev is published under a non-commercial licence: its official repository is gated behind it, and repackaged files such as the fp8 checkpoint carry it too; SDXL base is under OpenRAIL++. Licences are between you and the model's publisher. The machine does not check what you load.
Card memory, model by model
- 8 to 12 GB (RTX 2070 to RTX 4070 class): SDXL at its native resolution. Flux by streaming weights from system memory.
- 16 GB (RTX A4000, RTX 5080, Tesla T4): SDXL with headroom for LoRAs and ControlNets.
- 24 GB (RTX 3090, RTX 4090D, RTX A5000): the 17.2 GB fp8 Flux checkpoint whole. The 23.8 GB 16-bit Flux model alone fills the card, so run fp8 weights there.
- 48 GB (L40S, RTX A6000, RTX 6000 Ada): the full 16-bit Flux set, about 34 GB, resident at once with
--highvram. - 80 to 96 GB (A100 80GB, H100 80GB, RTX PRO 6000 96GB): several large models resident side by side, or long video workflows.
Every card in that list is on the GPU servers page, each priced as its own line beside the machine it seats in, so moving up a card changes the invoice by the difference between the two cards. The largest single card is 96 GB; past that the build is two or more cards in one chassis, a multi-GPU build quoted rather than ticked.
The rest of the machine
The card is not the whole bill of materials. Most GPU builds on the shelf carry 16 or 32 GB of system memory and two mirrored disks of 240 or 500 GB. Two of those numbers bite on Flux.
- System memory. ComfyUI's own Flux guide recommends the 16-bit T5 text encoder only with more than 32 GB of RAM, and the fp8 encoder otherwise. Unloaded models also wait in RAM between steps.
- Disk. One full Flux set is about 34 GB before LoRAs, upscalers and outputs. A model library outgrows a 240 GB mirror quickly.
- Both are build options. Built to order, memory and drives are your choice, each priced part by part on the build page.
Why a whole machine for ComfyUI
The card sits in a PCIe slot in a machine rented to one account and is passed straight to your operating system. No hypervisor, no MIG partition, no vGPU profile, no time-slicing: a queue of renders that takes an hour at noon takes an hour at four in the morning.
You load the kernel module and pin the driver, and CUDA stays at the release your workflow was tested against. The machine's own iLO or iDRAC sits on a separate management network for console, virtual media and power, so a driver experiment that drops the network still leaves you the screen. Nothing on the port counts bytes: models in, images and video out.
Installing ComfyUI on Linux
- Install the NVIDIA driver and confirm the card with
nvidia-smi. - As an ordinary user (here
comfy), clone the repository, create a virtual environment and install PyTorch, then ComfyUI's requirements. The PyTorch line is the one in ComfyUI's README, built for CUDA 13.0. - Start it. With no flags it listens on 127.0.0.1, port 8188.
- Put checkpoints in
models/checkpoints, Flux diffusion models inmodels/diffusion_models, text encoders inmodels/text_encodersand VAEs inmodels/vae.extra_model_paths.yamlpoints ComfyUI at a library elsewhere.
sudo apt install -y git python3-venv
git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv venv
. venv/bin/activate
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
python main.py
CUDA 13.0 needs NVIDIA driver 580 or later and a card of compute capability 7.5 (Turing) or newer. Anything older than Turing (Tesla P40, P100, V100, TITAN V, Quadro P6000) needs a PyTorch build for CUDA 12.6: use --index-url https://download.pytorch.org/whl/cu126 on the PyTorch line instead.
Reaching it without exposing port 8188
ComfyUI has no user accounts and no password. --listen with no address binds every IPv4 and IPv6 address, and then anyone who reaches port 8188 can queue jobs and run whatever custom nodes are installed. Keep the default bind and come in over SSH: the tunnel puts the interface on your laptop at http://localhost:8188.
# on your laptop
ssh -N -L 8188:127.0.0.1:8188 you@your-server
--tls-keyfile and --tls-certfile make ComfyUI serve HTTPS, which encrypts the traffic but still asks nobody for a password. For a team, put a reverse proxy that checks credentials and serves TLS in front, and leave ComfyUI on 127.0.0.1 behind it.
Running it day to day
A systemd unit keeps ComfyUI up across reboots, still bound to localhost:
# /etc/systemd/system/comfyui.service
[Unit]
Description=ComfyUI
After=network-online.target
[Service]
User=comfy
WorkingDirectory=/home/comfy/ComfyUI
ExecStart=/home/comfy/ComfyUI/venv/bin/python main.py --listen 127.0.0.1 --port 8188
Restart=on-failure
[Install]
WantedBy=multi-user.target
- Start it:
sudo systemctl daemon-reload, thensudo systemctl enable --now comfyui. - Custom nodes are code. Each one is Python that runs with the
comfyuser's rights. Install the ones you trust. ComfyUI-Manager'snormal-security level blocks its high-risk features whenever--listenis set to anything but a 127.x address. - Outputs and inputs can live on the big disk:
--output-directoryand--input-directorymove them. - Back up the work, not the weights. Models download again; your workflows, custom node list and outputs do not.
When this is the wrong place
A few images a week are cheaper on a service billed per image or per minute. A machine by the month earns its place when the queue runs daily, when the models and LoRAs are your own, or when the inputs should stay on hardware nobody else touches.
Video models and stacked high-resolution workflows can outgrow one card. Above 96 GB on one card the answer is a multi-card build, and how a workflow uses a second card depends on the workflow and its custom nodes; test that before you order the pair.
Questions
How much VRAM does ComfyUI need for Flux?
The FLUX.1-dev fp8 checkpoint is 17.2 GB, so a 24 GB card holds it whole. The 16-bit set is about 34 GB with its T5 encoder, which wants a 48 GB card to stay resident. Smaller cards still run Flux by streaming weights from system memory, which leans on host RAM and the PCIe bus.
Can a Tesla P40 run ComfyUI?
Yes, with the right PyTorch. Its 24 GB holds the fp8 Flux checkpoint, but it is a Pascal card, and the CUDA 13.0 build in ComfyUI's README does not support Pascal. Install the CUDA 12.6 build of PyTorch instead, and pin it: newer CUDA builds will not drive the card.
Is it safe to run ComfyUI with --listen on a server?
Not on a public address. ComfyUI has no login, so anyone who reaches port 8188 can run jobs and any installed custom node. Keep it on 127.0.0.1 and come in over SSH, or front it with a proxy that checks credentials over TLS.
Where do the model files go?
Checkpoints in models/checkpoints, Flux diffusion models in models/diffusion_models, text encoders in models/text_encoders and VAEs in models/vae, under the ComfyUI folder. To keep a library on another disk, list its paths in extra_model_paths.yaml.
Do I share the GPU with anyone?
No. The card is passed through to your operating system on a machine rented to one account: no MIG partition, no vGPU profile, no time-slicing. You install and pin the driver yourself.
Which models may I run?
Any whose licence lets you use them. FLUX.1-dev, for one, is under a non-commercial licence, whichever copy of the file you download. The licence is between you and the publisher; the machine does not check it.