NVIDIAVendor documented
L40S
Ada Lovelace · Ada · PCIe · 2023
The inference/fine-tuning value pick: strong FP8 throughput and 48 GB in an air-cooled 350 W card that fits standard servers.
LLM inferenceFine-tuningRAGMultimodalVision
Precision fingerprint
6432t3216BF168i8
Memory
48 GB
GDDR6
Bandwidth
0.864 TB/s
peak
TDP
350 W
air
Max model
~13B
FP16, planning est.
Compute throughput
| FP64 | — |
| FP32 | 91.6 TFLOPS |
| TF32 | 183 TFLOPS |
| FP16 | 366 TFLOPS |
| BF16 | 366 TFLOPS |
| FP8 | 733 TFLOPS |
| INT8 | 733 TFLOPS |
Platform & software
InterconnectPCIe only
PCIePCIe 4.0 x16
Coolingair
MIGNot supported
VirtualizationvGPU
FrameworksCUDA, TensorRT-LLM, Triton, vLLM
AvailabilityAWS, Azure, GCP, OCI, bare-metal
Known limitations
- ·No NVLink — scale-out over PCIe/network only
- ·GDDR6 bandwidth is the ceiling for big models