NVIDIAVendor documented

L40S

Ada Lovelace · Ada · PCIe · 2023

The inference/fine-tuning value pick: strong FP8 throughput and 48 GB in an air-cooled 350 W card that fits standard servers.

LLM inferenceFine-tuningRAGMultimodalVision
Precision fingerprint
6432t3216BF168i8
Memory

48 GB

GDDR6

Bandwidth

0.864 TB/s

peak

TDP

350 W

air

Max model

~13B

FP16, planning est.

Compute throughput

FP64
FP3291.6 TFLOPS
TF32183 TFLOPS
FP16366 TFLOPS
BF16366 TFLOPS
FP8733 TFLOPS
INT8733 TFLOPS

Platform & software

InterconnectPCIe only
PCIePCIe 4.0 x16
Coolingair
MIGNot supported
VirtualizationvGPU
FrameworksCUDA, TensorRT-LLM, Triton, vLLM
AvailabilityAWS, Azure, GCP, OCI, bare-metal
Known limitations
  • ·No NVLink — scale-out over PCIe/network only
  • ·GDDR6 bandwidth is the ceiling for big models