NVIDIAVendor documented
A30
Ampere · Ampere · PCIe · 2021
Power-efficient inference GPU for mainstream enterprise servers, with MIG for consolidating smaller workloads.
LLM inferenceVisionSpeechRAG
Precision fingerprint
6432t3216BF168i8
Memory
24 GB
HBM2
Bandwidth
0.933 TB/s
peak
TDP
165 W
air
Max model
~7B
FP16, planning est.
Compute throughput
| FP64 | 10.3 TFLOPS |
| FP32 | 10.3 TFLOPS |
| TF32 | 82 TFLOPS |
| FP16 | 165 TFLOPS |
| BF16 | 165 TFLOPS |
| FP8 | — |
| INT8 | 330 TFLOPS |
Platform & software
InterconnectNVLink — 200 GB/s
PCIePCIe 4.0 x16
Coolingair
MIGSupported
PartitioningUp to 4× MIG
VirtualizationvGPU, MIG
FrameworksCUDA, TensorRT, Triton
Availabilitybare-metal
Known limitations
- ·Mid-range throughput
- ·24 GB caps larger models