NVIDIAVendor documented

A30

Ampere · Ampere · PCIe · 2021

Power-efficient inference GPU for mainstream enterprise servers, with MIG for consolidating smaller workloads.

LLM inferenceVisionSpeechRAG
Precision fingerprint
6432t3216BF168i8
Memory

24 GB

HBM2

Bandwidth

0.933 TB/s

peak

TDP

165 W

air

Max model

~7B

FP16, planning est.

Compute throughput

FP6410.3 TFLOPS
FP3210.3 TFLOPS
TF3282 TFLOPS
FP16165 TFLOPS
BF16165 TFLOPS
FP8
INT8330 TFLOPS

Platform & software

InterconnectNVLink — 200 GB/s
PCIePCIe 4.0 x16
Coolingair
MIGSupported
PartitioningUp to 4× MIG
VirtualizationvGPU, MIG
FrameworksCUDA, TensorRT, Triton
Availabilitybare-metal
Known limitations
  • ·Mid-range throughput
  • ·24 GB caps larger models