Megatron-LM: Arbitrary code execution in the training job
CVE-2025-23264NVIDIA / GPU stackcurated
Impact
Arbitrary code execution in the training job
Who can reach it
Malicious model/checkpoint or dataset
What to do
Bump Megatron-LM in training images; rebuild; treat checkpoints as untrusted input
Fleet impact
How widespread
Common - Megatron is the reference large-model training stack for customers doing pretraining on rented clusters
Cost to remediate
hot-patch (upgrade to v0.12.1, rebuild training images)
Why it hits the whole fleet
A malicious file supplied to the Python component triggers code injection; on a shared training cluster the attacker executes inside a job that already holds cluster-wide storage and fabric credentials
References
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.