Operate

Live Telemetry

Current fleet state from the configured telemetry provider — per-cluster utilization, thermals, power, and ECC, with alerts assessed against the reference bands and deep-linked to their troubleshooting workflows.

Provider sample — a synthetic snapshot derived from the sample fleet + GPU catalog (no cluster required). Set TELEMETRY_PROVIDER=prometheus to scrape a real DCGM/ROCm Prometheus.

sample · synthetic
GPUs

1,048

9 clusters

Avg util

65%

fleet-wide

Draw

377 kW

board power

Critical

1

1 clusters

Warnings

98

3 clusters

ECC / throttle

0 / 1

uncorrectable / capped

Firing alerts
Cluster health
ClusterVendorUtilizationPeak tempPowerFaultsState

prod-train-us

H100 SXM5 · 256 GPU

NVIDIA
82%
92°C
154.3 kWthrottle 1critical

prod-train-eu

H100 SXM5 · 128 GPU

NVIDIA
77%
76°C
73.8 kWwatch

prod-train-us2

Instinct MI300X · 64 GPU

AMD
68%
78°C
36.8 kWwatch

prod-inf-us

L40S · 192 GPU

NVIDIA
64%
71°C
49.3 kWhealthy

prod-inf-eu

L40S · 96 GPU

NVIDIA
57%
67°C
22.9 kWhealthy

rag-us

A100 80GB SXM · 64 GPU

NVIDIA
71%
74°C
20.1 kWhealthy

vision-us

L4 · 120 GPU

NVIDIA
55%
65°C
5.8 kWwatch

dev-us

A30 · 48 GPU

NVIDIA
24%
52°C
3.5 kWhealthy

legacy-us

A100 40GB PCIe · 80 GPU

NVIDIA
36%
59°C
10.8 kWhealthy

Snapshot 2026-07-26 20:59:24 UTC · bands from the observability reference.