NVIDIAVendor documented

A100 80GB SXM

Ampere · Ampere · SXM4 · 2020

The prior-generation standard. Still highly capable and widely available; lacks FP8 so newer inference kernels are less efficient.

LLM trainingLLM inferenceFine-tuningHPCVision
Precision fingerprint
6432t3216BF168i8
Memory

80 GB

HBM2e

Bandwidth

2.039 TB/s

peak

TDP

400 W

air or liquid

Max model

~30B

FP16, planning est.

Compute throughput

FP6419.5 TFLOPS
FP3219.5 TFLOPS
TF32156 TFLOPS
FP16312 TFLOPS
BF16312 TFLOPS
FP8
INT8624 TFLOPS

Platform & software

InterconnectNVLink 3 — 600 GB/s
PCIePCIe 4.0 x16
Coolingair or liquid
MIGSupported
PartitioningUp to 7× MIG
VirtualizationvGPU, MIG
FrameworksCUDA, TensorRT, Triton, vLLM
AvailabilityAWS, Azure, GCP, OCI, bare-metal
Known limitations
  • ·No FP8 / Transformer Engine — slower on modern LLM stacks
  • ·Prior generation, but broadly available and cost-effective