GPU VulnDB

Database/NVIDIA / GPU stack

Linux kernel amdgpu kernel driver core (drm/amdgpu): Memory is handed to a consumer without being initialised or

CVE-2024-42228NVIDIA / GPU stackcurated

Impact

Memory is handed to a consumer without being initialised or cleared in the amdgpu kernel driver core. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: Using uninitialized value *size when calling amdgpu_vce_cs_reloc

Who can reach it

Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.

What to do

Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.

References

This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.