NVIDIAVendor documented
A40
Ampere · Ampere · PCIe · 2020
Professional-visualization and inference GPU with 48 GB GDDR6. Common for graphics, Omniverse, and mixed inference.
GraphicsVisionLLM inferenceDev sandbox
Precision fingerprint
6432t3216BF168i8
Memory
48 GB
GDDR6
Bandwidth
0.696 TB/s
peak
TDP
300 W
air
Max model
~13B
FP16, planning est.
Compute throughput
| FP64 | — |
| FP32 | 37.4 TFLOPS |
| TF32 | 74.8 TFLOPS |
| FP16 | 149.7 TFLOPS |
| BF16 | 149.7 TFLOPS |
| FP8 | — |
| INT8 | 299 TFLOPS |
Platform & software
InterconnectNVLink — 112 GB/s
PCIePCIe 4.0 x16
Coolingair
MIGNot supported
VirtualizationvGPU
FrameworksCUDA, TensorRT, Triton, Omniverse
AvailabilityAzure, bare-metal
Known limitations
- ·GDDR6 bandwidth well below HBM parts
- ·No MIG