NVIDIAVendor documented
L4
Ada Lovelace · Ada · PCIe (low-profile) · 2023
72 W single-slot GPU for dense, power-efficient inference and video. Fits nearly any server without new power planning.
LLM inferenceVisionSpeechRAGDev sandbox
Precision fingerprint
6432t3216BF168i8
Memory
24 GB
GDDR6
Bandwidth
0.3 TB/s
peak
TDP
72 W
air
Max model
~7B
FP16, planning est.
Compute throughput
| FP64 | — |
| FP32 | 30 TFLOPS |
| TF32 | 60 TFLOPS |
| FP16 | 121 TFLOPS |
| BF16 | 121 TFLOPS |
| FP8 | 242 TFLOPS |
| INT8 | 242 TFLOPS |
Platform & software
InterconnectPCIe only
PCIePCIe 4.0 x16
Coolingair
MIGNot supported
VirtualizationvGPU
FrameworksCUDA, TensorRT-LLM, Triton
AvailabilityAWS, GCP, bare-metal
Known limitations
- ·Low memory bandwidth
- ·Best for small models and video/vision, not training