GenAIHub
Back to Technical
Hardware

GPU Infrastructure

Choose the right GPU for LLM inference and training. Compare consumer, datacenter, and cloud GPU options. Understand VRAM requirements and costs.

NVIDIA GPU Comparison

GPU VRAM FP16 TFLOPS Type Cost (New)
RTX 4090 24GB ~83 Consumer $1,600
RTX 4080 16GB ~49 Consumer $1,000
A10G 24GB ~31 Cloud (AWS) ~$1/hr
L4 24GB ~121 Cloud (GCP) ~$0.80/hr
A100 40GB 40GB ~312 Datacenter ~$2/hr
A100 80GB 80GB ~312 Datacenter ~$3/hr
H100 80GB 80GB ~990 Datacenter ~$4/hr
H200 141GB 141GB ~990 Datacenter ~$6/hr

VRAM Requirements by Model Size

Model FP16 Q8 (8-bit) Q4 (4-bit) Recommended GPU
7B 14GB 7GB 4GB RTX 4060 Ti 16GB
13B 26GB 13GB 7GB RTX 4090 24GB
30-34B 68GB 34GB 17GB 2x RTX 4090 or A100
70B 140GB 70GB 35GB 2x A100 80GB or H100
405B 810GB 405GB 202GB 8x H100 (DGX)

πŸ’‘ Formula: VRAM (FP16) β‰ˆ Parameters Γ— 2 bytes. Add ~20% overhead for KV cache during inference.

Cloud GPU Pricing (On-demand)

Provider GPU Instance $/hr
AWS 1x A10G g5.xlarge $1.00
AWS 8x A100 p4d.24xlarge $32.77
GCP 1x L4 g2-standard-4 $0.70
GCP 8x H100 a3-highgpu-8g $32.00
Azure 1x A100 NC24ads_A100_v4 $3.67
Lambda Labs 1x A100 gpu_1x_a100 $1.10
RunPod 1x A100 On-demand $1.64

Quantization for Less VRAM

GGUF (llama.cpp)

CPU + GPU hybrid. Q4_K_M is the sweet spot for quality vs size.

AWQ

4-bit, GPU-optimized. Best for vLLM production inference.

GPTQ

4-bit, widely supported. Good for consumer GPUs.

FP8

8-bit, minimal quality loss. Requires H100/Ada GPUs.

Build vs Rent Decision

Factor Build (Buy GPUs) Rent (Cloud)
Upfront Cost High ($10K-$500K+) None
Break-even ~6-12 months 24/7 usage Never (pay per use)
Best For Continuous training, privacy Burst workloads, experimentation
Maintenance You handle everything Provider handles

Related Topics