NVIDIAVendor documented

L4

Ada Lovelace · Ada · PCIe (low-profile) · 2023

72 W single-slot GPU for dense, power-efficient inference and video. Fits nearly any server without new power planning.

LLM inferenceVisionSpeechRAGDev sandbox
Precision fingerprint
6432t3216BF168i8
Memory

24 GB

GDDR6

Bandwidth

0.3 TB/s

peak

TDP

72 W

air

Max model

~7B

FP16, planning est.

Compute throughput

FP64
FP3230 TFLOPS
TF3260 TFLOPS
FP16121 TFLOPS
BF16121 TFLOPS
FP8242 TFLOPS
INT8242 TFLOPS

Platform & software

InterconnectPCIe only
PCIePCIe 4.0 x16
Coolingair
MIGNot supported
VirtualizationvGPU
FrameworksCUDA, TensorRT-LLM, Triton
AvailabilityAWS, GCP, bare-metal
Known limitations
  • ·Low memory bandwidth
  • ·Best for small models and video/vision, not training