Add items to comparison
Add items to wish list
0
Add products to cart

AI Workstations

Sort:
View:

An AI workstation is a computer where the key spec is video memory: it decides which neural networks you can run locally and how fast they will work. PowerUp's AI workstations come with one, two or three graphics cards and up to 192 GB of total video memory: NVIDIA RTX PRO Blackwell, GeForce RTX 50, NVIDIA RTX A, Quadro RTX and Intel Arc Pro. Processors range from Core Ultra and Ryzen 9 to Threadripper PRO 9000, Xeon W and AMD EPYC, with 32 to 256 GB of RAM. Every station passes a 4-hour stress test and comes with a 24-month warranty.

How much video memory your task needs

When choosing an AI PC, start from the model: it has to fit entirely in video memory, or speed drops several times over. For large language models (LLMs) the math is simple: about 0.6 GB per billion parameters with 4-bit quantization, about 1.1 GB at 8 bits and 2 GB in FP16, plus 10–20% for context.

Total video memoryWhat runs locallyConfigurations in the catalog
16 GB7–14B-parameter LLMs at 4 bits, Stable Diffusion XL, Flux in FP8GeForce RTX 5080, Tesla V100 16 GB
24–32 GBLLMs up to 32B at 4 bits, 14B at 8 bits, image generation with FluxGeForce RTX 5090 32 GB, Quadro RTX 6000 24 GB, Arc Pro B60 24 GB, 2× RTX 5060 Ti 16 GB, 2× RTX 5070 Ti, 2× RTX A4000, 2× RTX PRO 2000 Blackwell
40–48 GBLLMs around 70B at 4 bits, 32B at 8 bitsRTX PRO 5000 Blackwell 48 GB, 2× RTX PRO 4000 Blackwell, 2× Arc Pro B60, 2× RTX A5000, 2× Quadro RTX 6000, 2× RTX A4500, 3× RTX 5060 Ti 16 GB, 3× RTX A4000
72–96 GB70B LLMs at 8 bits (needs 96 GB), 100–120B MoE models at 4 bits, LoRA fine-tuning of models up to 30B2× RTX PRO 5000 Blackwell (96 GB), 3× RTX PRO 4000 Blackwell (72 GB), 3× Arc Pro B60 (72 GB)
192 GB70B LLMs in FP16, MoE models of 200B+ at 4 bits, full fine-tuning of 7–8B models2× RTX PRO 6000 Blackwell 96 GB

Graphics cards in our AI workstations

Graphics cardMemory per cardArchitectureBest for
NVIDIA RTX PRO 6000 Blackwell96 GB GDDR7 with ECCBlackwellThe largest models on a single card, fine-tuning, 24/7 operation
NVIDIA RTX PRO 5000 Blackwell48 GB GDDR7 with ECCBlackwell70B-parameter LLMs on a single card
NVIDIA RTX PRO 4000 / 2000 Blackwell24 / 16 GB GDDR7 with ECCBlackwellTwo- and three-card builds: the 4000 is single-slot, the 2000 draws 70 W
GeForce RTX 509032 GB GDDR7BlackwellThe fastest inference and image generation for the money
GeForce RTX 5060 Ti 16 GB / 5070 Ti / 508016 GB GDDR7BlackwellAn affordable entry into AI, two- and three-card builds
NVIDIA RTX A5000 / A4500 / A400024 / 20 / 16 GB GDDR6 with ECCAmpereInference and fine-tuning with proven drivers
Quadro RTX 6000 / RTX 500024 / 16 GB GDDR6TuringLLM inference with plenty of memory for less money
Intel Arc Pro B60 / B5024 / 16 GB GDDR6BattlemageInference via OpenVINO, PyTorch and llama.cpp; the cheapest video memory
Tesla V10016 GB HBM2VoltaDouble-precision (FP64) scientific computing, training classic models

NVIDIA cards support CUDA, so virtually all AI software runs on them. Blackwell (RTX PRO and RTX 50) accelerates FP4 and FP8 math. Turing and Volta do not support BF16 or FlashAttention 2: they are fine for inference, but Ampere and Blackwell are better for training modern models. Intel Arc Pro gives you more memory for the same money, but software written only for CUDA will not run on it. For the differences between professional card generations, see our overview of the NVIDIA professional graphics lineup; we also explain why GeForce RTX 50 cards have so little memory.

One powerful card or several

  • Language models — Ollama, llama.cpp and vLLM split a model across cards, so 2× 24 GB can run a model that needs 48 GB. A single card with the same amount of memory usually generates text faster, because data does not travel between cards over PCIe.
  • Image generation — Stable Diffusion and ComfyUI render each image on one card. Several cards let you run several jobs in parallel, but they do not pool their memory.
  • Training — several cards noticeably speed up training when the model fits in each card's memory. Without NVLink, the gain is smaller than the card count.
  • Card-to-card links — RTX PRO Blackwell and GeForce RTX 50 cards have no NVLink, so data moves over PCIe, which is why the number of CPU PCIe lanes matters for two or three cards.

Processor and platform

PlatformCPU PCIe lanesMemory channelsCards in our buildsWhen to choose it
Core Ultra 5/7, 14th gen Core i5/i7, Ryzen 9 7900X/9900X/9950X20–242 × DDR51–2Inference on one or two cards, image generation
Threadripper 3970X64 PCIe 4.04 × DDR42–3Three cards at an affordable price
AMD EPYC 7551, 7F52, 7F72, 7C13128 PCIe 3.0/4.08 × DDR4 ECC2–3Plenty of lanes and memory for less money
Threadripper PRO 9955WX–9995WX128 PCIe 5.08 × DDR5 ECC2–3Flagship stations with RTX PRO 5000/6000 Blackwell
Xeon w5-3423, w9-3475X112 PCIe 5.08 × DDR5 ECC2–3Three cards on an Intel platform, up to 256 GB of memory

On desktop Core and Ryzen systems, two cards usually run at x8/x8 or x16/x4. On Threadripper PRO, Xeon W and EPYC, each card gets its own 16 lanes. Eight memory channels matter when part of a large model is offloaded to system RAM — see our article on why memory channels matter on a workstation. How server EPYC chips such as the 7C13 differ from retail models is covered in our review of OEM EPYC versions.

RAM, storage and software

  • RAM — at least as much as your total video memory, because the model is loaded into system RAM first. Our builds have 32 to 256 GB, and you can add more when ordering.
  • NVMe SSD of 1 TB or more — a 70B model at 4 bits takes about 40 GB, and datasets and training checkpoints take hundreds of gigabytes. Our builds use 1–4 TB SSDs.
  • Software — Ollama, LM Studio, llama.cpp and vLLM for language models, ComfyUI and Stable Diffusion for images, PyTorch and TensorFlow for training. vLLM officially runs only on Linux; everything else also runs on Windows.

For a headless 24/7 service such as an internal company chatbot, see our server PCs. Single-card GeForce RTX 50 builds for work are collected under computers with RTX 50 series graphics, and other lines are in the workstation catalog.

Testing, warranty, delivery and payment

  • Testing — every station runs under load in AIDA64 and FurMark for 4 hours. An unactivated copy of Windows 10/11 is installed for testing; a license key can be purchased separately.
  • Your configuration — we can change the memory size, the number and model of graphics cards, the processor within the platform and the drives. For a station built around your model, use the custom configuration form.
  • 24-month warranty — repairs take 1–2 days on average plus shipping, and a replacement unit is available in the meantime. Warranty terms.
  • Delivery and payment — Nova Poshta across Ukraine in 1–3 days, or pickup from our Kyiv office, where you can test the computer before paying. Pay by bank transfer with no fee, by invoice with VAT, or in installments from PrivatBank and PUMB for up to 6 months. Details are on the payment and delivery page.

AI workstation FAQ

How much video memory do I need to run a language model locally?

Multiply the parameter count in billions by 0.6 GB for 4-bit quantization, by 1.1 GB for 8 bits or by 2 GB for FP16, then add 10–20% for context. A 14B model at 4 bits fits in 16 GB, a 32B model needs 24 GB, a 70B model needs 48 GB, and a 70B model at 8 bits needs 96 GB.

Which is better for AI: the GeForce RTX 5090 or RTX PRO Blackwell?

The RTX 5090 with 32 GB delivers the most speed for the money as long as the model fits in that memory. The RTX PRO 5000 and 6000 Blackwell offer 48 and 96 GB per card, ECC memory and drivers built for 24/7 operation — they are the choice for models of 70B parameters and up and for multi-card stations.

Can I combine the video memory of several cards?

For language models, yes: Ollama, llama.cpp and vLLM spread the model across cards, so two 24 GB cards can run a 48 GB model. For Stable Diffusion and ComfyUI, no: each image is rendered on one card, while several cards let you run jobs in parallel.

Is the Intel Arc Pro B60 good for neural networks?

Yes, for inference: its 24 GB of video memory costs less than on NVIDIA cards, and models run through OpenVINO, PyTorch and llama.cpp. The one limitation is that software built only for CUDA will not run on Intel Arc, so NVIDIA is the safer choice for training and less common tools.

What is the difference between the RTX 6000 cards?

Four different cards have carried this name: the Quadro RTX 6000 (Turing, 24 GB GDDR6), the RTX A6000 (Ampere, 48 GB), the RTX 6000 Ada (48 GB) and the RTX PRO 6000 Blackwell (96 GB GDDR7). Our AI workstations use the Quadro RTX 6000 as an affordable 24 GB option and the RTX PRO 6000 Blackwell for the largest models.

Do I need a powerful processor for AI work?

When the model sits entirely in video memory, the processor has little effect on generation speed, and a Core Ultra or Ryzen 9 is enough for one or two cards. Threadripper PRO, Xeon W and EPYC are needed for three cards (PCIe lanes), for offloading part of a model to system RAM (eight channels) and for preparing large datasets.

Which operating system should I use: Windows or Linux?

Ollama, LM Studio and ComfyUI run on Windows, while vLLM and most training tools are easier to run on Linux, such as Ubuntu. An unactivated copy of Windows 10/11 is installed for testing; you can activate it with your own key or replace it with Linux.

Can I change the configuration and buy an AI workstation as a company?

Yes. We can change the memory size, the number and model of graphics cards, the processor within the platform and the drives — add your requests in the order comment or agree on them with a manager. Payment options include bank transfer with no fee, invoice with VAT and installments from PrivatBank and PUMB for up to 6 months; the warranty is 24 months.