{"entries":[{"id":"CVE-2013-4782","cve":"CVE-2013-4782","aliases":[],"title":"Supermicro BMC (IPMI cipher suite 0): Authentication bypass — arbitrary IPMI commands with any password when cipher suite 0 is enabled. Still…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI cipher suite 0)","year":"2013","cvss_score":10,"severity":"critical","kev":false,"impact":"Authentication bypass — arbitrary IPMI commands with any password when cipher suite 0 is enabled. Still shipped enabled on some ODM builds a decade later","attack_vector":"Network / IPMI over LAN","remediation":"Disable cipher suite 0 in BMC config fleet-wide and assert it in a config scan; not fixable by firmware alone since it is a spec-permitted mode","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4782"],"status":"curated"},{"id":"CVE-2013-4783","cve":"CVE-2013-4783","aliases":[],"title":"Dell iDRAC (IPMI 1.5 cipher 0): Remote authentication bypass via cipher suite 0","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC (IPMI 1.5 cipher 0)","year":"2013","cvss_score":10,"severity":"critical","kev":false,"impact":"Remote authentication bypass via cipher suite 0","attack_vector":"Network / IPMI","remediation":"Disable cipher 0 in the iDRAC config baseline","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4783"],"status":"curated"},{"id":"CVE-2013-4784","cve":"CVE-2013-4784","aliases":[],"title":"HPE iLO (IPMI cipher 0): IPMI authentication bypass via cipher suite 0 on the iLO BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO (IPMI cipher 0)","year":"2013","cvss_score":10,"severity":"critical","kev":false,"impact":"IPMI authentication bypass via cipher suite 0 on the iLO BMC","attack_vector":"Network / IPMI","remediation":"Config baseline change disabling cipher 0","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4784"],"status":"curated"},{"id":"CVE-2014-2955","cve":"CVE-2014-2955","aliases":[],"title":"Raritan PX rack PDU (before firmware 1.5.11, DPXR20A-16 and related PX models): TENANT ISOLATION: the PDU's IPMI interface accepts 'cipher suite 0' (aka cipher zero), a known-broken IPMI…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Raritan PX rack PDU (before firmware 1.5.11, DPXR20A-16 and related PX models)","year":"2014","cvss_score":10,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: the PDU's IPMI interface accepts 'cipher suite 0' (aka cipher zero), a known-broken IPMI auth mode that accepts any password. An attacker who reaches the IPMI port can issue arbitrary IPMI commands with no valid credential — including power-control commands to cut or cycle outlets feeding whatever racks that PDU serves, tenant-owned or not.","attack_vector":"Fully remote and unauthenticated over the network-reachable IPMI port — the attacker just needs to request cipher suite 0 and supply any password; the PDU accepts it.","remediation":"Firmware upgrade to 1.5.11 or later, which disables cipher-zero support. If a firmware upgrade isn't immediately possible, a network-segmentation change to firewall off the IPMI port from anything but a trusted management VLAN is the compensating control — this is the same 'IPMI cipher zero' class of bug that hit many BMC vendors around the same era, so audit for other devices on the same segment too. Flash each PDU one at a time; outlets keep powering their load during the update, but remote power-control briefly drops.","references":["https://nvd.nist.gov/vuln/detail/CVE-2014-2955"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2015-5053","cve":"CVE-2015-5053","aliases":["GRID vGPU host memory mapping"],"title":"NVIDIA GPU driver / GRID vGPU and vSGA - host memory mapping path: MULTI-TENANT ISOLATION: the host memory mapping path did not restrict access to third-party device I/O…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU driver / GRID vGPU and vSGA - host memory mapping path","year":"2015","cvss_score":10,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: the host memory mapping path did not restrict access to third-party device I/O memory, giving a guest privilege escalation on the host. Scored 10.0. Included here despite its age because it is the archetype of the vGPU escape class and because GRID/vSGA stacks this old are still found running in long-lived VDI estates that were never in anyone's patch programme.","attack_vector":"A guest on a GRID vGPU or vSGA host. The mapping path is reachable from the guest driver.","remediation":"Upgrade the host driver to R346 346.87 / R352 352.41 (Linux) or R352 352.46 (GRID vGPU and vSGA) or, realistically, to a supported branch - anything on these versions is a decade out of support. Cost: full hypervisor host drain. If you find this in your estate the finding is not the CVE, it is that you have an unmanaged vGPU host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-5053"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2017-12542","cve":"CVE-2017-12542","aliases":[],"title":"HPE iLO4: Authentication bypass and remote code execution — the \"29 A's\" `Connection` header bug","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO4","year":"2017","cvss_score":10,"severity":"critical","kev":false,"impact":"Authentication bypass and remote code execution — the \"29 A's\" `Connection` header bug; trivially scriptable pre-auth root on the BMC","attack_vector":"Network, unauthenticated","remediation":"iLO4 firmware update to 2.53+; a node left unpatched here is fully owned by a single curl request","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-12542"],"status":"curated","fleet":{"ubiquity":"common - iLO is HPE's BMC across ProLiant/Apollo, incl. GPU-dense SKUs","remediation_pain":"firmware-flash to iLO >= 2.54 on every node, out-of-band; the exploit is trivial (a long header) and public, so exposure windows are measured in hours","pain_class":"firmware-flash","why_fleet_wide":"Unauthenticated remote auth bypass into the BMC yields administrator on the management processor, virtual-media boot of attacker media, and firmware-level persistence under the OS - identical across every HPE node of that generation."}},{"id":"CVE-2019-16649","cve":"CVE-2019-16649","aliases":[],"title":"Supermicro BMC virtual media (H11/H12/M11/X9/X10/X11): Virtual media service uses weak/absent encryption and authentication — credential capture and attaching an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC virtual media (H11/H12/M11/X9/X10/X11)","year":"2019","cvss_score":10,"severity":"critical","kev":false,"impact":"Virtual media service uses weak/absent encryption and authentication — credential capture and attaching an arbitrary virtual USB device to the host, i.e. arbitrary boot media","attack_vector":"Network, unauthenticated","remediation":"BMC firmware update across every affected generation; interim control is blocking the virtual-media ports (623/5900/5901) at the management-network boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16649"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-16650","cve":"CVE-2019-16650","aliases":["USBAnywhere"],"title":"Supermicro X10/X11 BMC (virtual media service): The BMC's virtual media service reuses socket file descriptors, so an unauthenticated attacker inherits an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro X10/X11 BMC (virtual media service)","year":"2019","cvss_score":10,"severity":"critical","kev":false,"impact":"The BMC's virtual media service reuses socket file descriptors, so an unauthenticated attacker inherits an existing client's privileges and can attach a virtual USB device to the server. In practice that means booting your node from the attacker's image, or dropping files onto a running host, without ever having a BMC credential. On a bare-metal GPU fleet this is a direct tenant-to-tenant and outsider-to-host compromise, and the implanted image outlives any OS reinstall.","attack_vector":"Anything with a network route to the BMC's virtual media port. No credentials, no user interaction. Supermicro boards are the whitebox default under a large share of neocloud GPU capacity, and BMCs on these boards are frequently found directly on a routable network.","remediation":"BMC firmware flash per node, out-of-band, with the usual Supermicro caveat that the fixed version differs per board SKU - you need a per-model inventory before you can plan the rollout. Immediate config-only mitigation that actually works: block the virtual media ports (623, 5900, 623/udp and the 623x range Supermicro uses) at the network edge and put every BMC behind a jump host on a dedicated management VLAN. Do the network control first; the flash campaign will take weeks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16650","https://eclypsium.com/blog/usbanywhere/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-5684","cve":"CVE-2019-5684","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime): MULTI-TENANT ISOLATION: a crafted shader reads out of bounds on an input texture array and gets code…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime)","year":"2019","cvss_score":10,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: a crafted shader reads out of bounds on an input texture array and gets code execution in the driver. VMware shipped its own advisory for this (VMSA-2019-0012) because in a virtualised graphics setup the shader comes from inside a guest VM - so this is a guest-to-host code execution path on a shared GPU host, which is why it carries a CVSS of 10.0. On a GPU cloud running vSGA/vGPU-adjacent graphics, one tenant's shader compromises the hypervisor host and therefore every other tenant on it.","attack_vector":"Anyone who can submit a shader to the host GPU: a tenant VM, a remote graphics session, or a local process. In the virtualised case, an unprivileged user inside any guest.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved. If you run VMware ESXi with virtualised graphics, also apply the fix VMware ships in VMSA-2019-0012 - the hypervisor-side package is separate from the in-guest driver, and both matter.","references":["http://www.vmware.com/security/advisories/VMSA-2019-0012.html","https://support.lenovo.com/us/en/product_security/LEN-28096","https://www.talosintelligence.com/vulnerability_reports/TALOS-2019-0779","https://nvd.nist.gov/vuln/detail/CVE-2019-5684"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1388","cve":"CVE-2021-1388","aliases":[],"title":"Cisco ACI Multi-Site Orchestrator (Application Services Engine): TENANT ISOLATION: complete unauthenticated authentication bypass on the controller that programs policy…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco ACI Multi-Site Orchestrator (Application Services Engine)","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: complete unauthenticated authentication bypass on the controller that programs policy across every ACI site. Whoever gets this owns the tenant separation model for the entire multi-site fabric — they can write EPG and contract policy that stitches any tenant to any other. It is a 10.0 for a reason.","attack_vector":"Unauthenticated, remote — anything that can reach the MSO API endpoint. If the orchestrator's management interface is on a flat ops network, that is a very large set of machines.","remediation":"Upgrade the MSO application. Application-level upgrade rather than a switch reload, so the data plane stays up — but treat any pre-patch exposure as a policy compromise and re-audit every contract and EPG binding afterwards, which is the real cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1388"],"status":"curated"},{"id":"CVE-2021-22205","cve":"CVE-2021-22205","aliases":[],"title":"GitLab: Image files passed unvalidated to a file parser (ExifTool)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2021","cvss_score":10,"severity":"critical","kev":true,"impact":"[KEV] Image files passed unvalidated to a file parser (ExifTool) -> unauthenticated remote command execution","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; assume compromise on any unpatched internet-facing instance","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-22205"],"status":"curated"},{"id":"CVE-2021-23281","cve":"CVE-2021-23281","aliases":[],"title":"Eaton Intelligent Power Manager (IPM) prior to 1.69: PHYSICAL. Unauthenticated remote code execution on Eaton's power-management platform, via unsanitised data…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Intelligent Power Manager (IPM) prior to 1.69","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"PHYSICAL. Unauthenticated remote code execution on Eaton's power-management platform, via unsanitised data reaching a Node.js code path. IPM is the software that monitors and shuts down infrastructure in response to power events across a site, so unauthenticated RCE here is total control of the power-response layer. A CVSS 10.0 on facility software is rare and this one earns it.","attack_vector":"Unauthenticated, remote, over the network to the IPM server.","remediation":"Upgrade IPM to 1.69 or later. Server-side upgrade, so cheap in maintenance terms - but assume compromise on any instance that was network-reachable and rebuild rather than patch. Rotate every device credential IPM held. IPM ships as part of several OEM power bundles, so check for it under other names before concluding you do not run it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23281"],"status":"curated"},{"id":"CVE-2021-26728","cve":"CVE-2021-26728","aliases":[],"title":"Lanner IAC-AST2500A BMC standard firmware 1.10.0: Arbitrary code execution as root on the BMC, at the maximum severity the scale allows. The operator outcome…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC standard firmware 1.10.0","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Arbitrary code execution as root on the BMC, at the maximum severity the scale allows. The operator outcome is complete out-of-band ownership of every node carrying this module: power control, boot-device selection, KVM into whatever the tenant has on screen, virtual-media boot of an attacker-supplied image, and an implant in BMC flash that survives host reimaging. Because the IAC-AST2500A is a module dropped into other people's board designs, an operator may be running it without the word 'Lanner' appearing anywhere in their asset inventory. Command injection plus stack buffer overflows in the KillDupUsr_func handler of spx_restservice, the REST service that fronts the BMC web interface. The IAC-AST2500A is a MegaRAC-derived BMC module resold into a wide range of whitebox and edge server designs.","attack_vector":"Network reachability to the BMC's REST service. Anything routable to the out-of-band management VLAN, which for whitebox and edge deployments is frequently less segmented than in a purpose-built datacenter.","remediation":"Firmware flash from Lanner, per module. This is the hardest remediation class in this database: Lanner's website returns a blanket Cloudflare 403 to automated clients, the advisories that exist are third-party (Nozomi Networks) rather than vendor-published, and firmware for a resold BMC module often has to be sourced through the board integrator rather than Lanner directly. For many operators the honest answer is that no fixed image is obtainable, and the only real mitigation is hard network isolation of the management VLAN with an explicit allowlist. Inventory first - identify which of your boards carry an IAC-AST2500A before assuming you are unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26728","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-26728/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26729","cve":"CVE-2021-26729","aliases":[],"title":"Lanner IAC-AST2500A BMC firmware 1.10.0: Root on the BMC without any credential at all, because the vulnerable handler is the login handler. This is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC firmware 1.10.0","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Root on the BMC without any credential at all, because the vulnerable handler is the login handler. This is the single worst entry in the ODM section: an attacker with a network path and nothing else takes complete out-of-band control of the node - power, console, virtual media, and persistent firmware residency below every layer the operator can reimage. On bare-metal rental this destroys the tenant handoff guarantee, because a node compromised this way stays compromised through wipe and reprovision. Command injection and multiple stack buffer overflows in the Login_handler_func function of spx_restservice, i.e. in the code that runs before anyone has authenticated.","attack_vector":"Anything that can reach the BMC's REST service over the network, unauthenticated. No credential, no host access, no prior foothold - only routability to the management interface.","remediation":"Firmware flash from Lanner or your board integrator. As with the rest of this cluster, obtaining a fixed image is the real obstacle: the vendor's site is WAF-blocked to automated access and the public advisories are third-party. Given the unauthenticated nature of this one, treat network isolation as mandatory and immediate rather than as a stopgap - these BMCs must not be reachable from anything except a small, explicitly allowlisted set of management hosts, and never from a tenant or general corporate network. If you cannot obtain fixed firmware, document the residual risk and plan the hardware out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26729","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-26729/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-32798","cve":"CVE-2021-32798","aliases":[],"title":"Jupyter Notebook (untrusted notebooks): Untrusted notebook executes JavaScript in the user's session on open","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Notebook (untrusted notebooks)","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"Untrusted notebook executes JavaScript in the user's session on open","attack_vector":"Customer-supplied `.ipynb` opened by another user or an operator","remediation":"Upgrade. Notebook files shared between tenants are active content","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32798"],"status":"curated"},{"id":"CVE-2021-39296","cve":"CVE-2021-39296","aliases":["GHSA-gg9x-v835-m48q"],"title":"OpenBMC phosphor-net-ipmid (IPMI 2.0 RMCP+ / IPMI over LAN): The headline OpenBMC bug. Crafted IPMI session-setup messages skip authentication entirely and hand the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC phosphor-net-ipmid (IPMI 2.0 RMCP+ / IPMI over LAN)","year":"2021","cvss_score":10,"severity":"critical","kev":false,"impact":"The headline OpenBMC bug. Crafted IPMI session-setup messages skip authentication entirely and hand the attacker full administrative control of the BMC - no credentials, no prior access, just UDP packets at the management interface. From there: power-cycle any node, mount virtual media, install BMC-resident firmware that survives host reimage, and pivot onto the host. On a GPU cluster where the management VLAN reaches every node, one packet source that can see that VLAN owns the fleet's out-of-band plane. CVSS 10.0 with scope change, which is rare and deserved. Google's security team reported it; Intel shipped it as SA-00737.","attack_vector":"Network access to the BMC's IPMI-over-LAN port (UDP 623). Unauthenticated. In practice: anyone who reaches the management VLAN - a misrouted tenant network, a jump host, a compromised switch, or a BMC accidentally exposed to the internet.","remediation":"Fixed in OpenBMC after 2.9. Getting the fix onto nodes is a BMC firmware flash: out-of-band, per node, ODM-rebase-lagged, brick risk. But the config-only mitigation here is strong and should be done first, today: disable IPMI over LAN entirely and use Redfish. OpenBMC's Redfish support has been production-ready for years and most fleets no longer need RMCP+. If you cannot disable it, ACL UDP 623 so only your management jump hosts can reach it. Check your fleet for BMCs reachable outside the management VLAN before doing anything else.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39296","https://github.com/google/security-research/security/advisories/GHSA-gg9x-v835-m48q","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00737.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-0543","cve":"CVE-2022-0543","aliases":[],"title":"Redis: Debian/Ubuntu packaging leaves a Lua sandbox escape","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2022","cvss_score":10,"severity":"critical","kev":true,"impact":"[KEV] Debian/Ubuntu packaging leaves a Lua sandbox escape -> RCE; mass-exploited by botnets","attack_vector":"Network (remote)","remediation":"Control-plane: distro package upgrade; never expose Redis to a tenant network","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0543"],"status":"curated"},{"id":"CVE-2022-21941","cve":"CVE-2022-21941","aliases":["ICSA-22-242-11"],"title":"Software House iSTAR Ultra door controller (before 6.8.9.CU01): Unauthenticated command injection giving root on the door controller. The iSTAR Ultra is the panel that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Software House iSTAR Ultra door controller (before 6.8.9.CU01)","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated command injection giving root on the door controller. The iSTAR Ultra is the panel that decides whether a door opens - including the doors into the datacenter hall and the cages inside it. Root on it means an attacker unlocks doors on demand, holds them unlocked, enrols their own credentials, and deletes or rewrites the access log so no record of the entry exists. What someone standing inside your cage can then do is the real impact: pull NVMe drives holding customer data and model weights, attach a laptop to a server's console or management port, plug into the out-of-band switch and reach every BMC in the row, install an interposer on a management link, or simply photograph the topology. For a bare-metal GPU provider that is also a tenant-handoff catastrophe - the boundary you sell your customers is precisely that nobody else can physically touch their machines, and this bug removes it silently. Root persistence on the controller means the compromise survives your incident response unless you re-image the panel.","attack_vector":"Unauthenticated, over the network, to the controller. iSTAR panels are typically on a dedicated physical-security VLAN, which sounds reassuring until you check who else is on it: the CCTV/VMS servers, the intercom system, the badge-office workstations, and the security integrator's remote-support path. Any of those is a stepping stone. Panels are also often mounted in unsecured back-of-house spaces, so physical access to the panel's Ethernet port is a parallel route.","remediation":"Firmware update to 6.8.9.CU01 or later, applied through the Software House/Johnson Controls integrator. Door controller firmware updates are disruptive in a specific way most operators underestimate: doors typically fall back to a fail-secure or fail-safe local mode during the flash, so you need security staff physically present at affected doors for the window. Plan it, do not skip it - a 10.0 on the panel guarding your GPUs is not something to defer. After patching, re-image rather than trust any panel you suspect was reachable while vulnerable, rotate the integrator's credentials, and put the physical-security VLAN behind a firewall with an explicit allow-list rather than treating it as inherently trusted.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-22-242-11","https://nvd.nist.gov/vuln/detail/CVE-2022-21941","https://www.johnsoncontrols.com/cyber-solutions/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-29165","cve":"CVE-2022-29165","aliases":[],"title":"Argo CD: Unauthenticated attacker forges JWTs and gains full Argo CD admin, which in a GitOps cluster means…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated attacker forges JWTs and gains full Argo CD admin, which in a GitOps cluster means cluster-admin","attack_vector":"Unauthenticated network reaching the Argo CD API","remediation":"Emergency Argo CD upgrade; rotate the signing key and all tokens; audit all Applications for tampering","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29165"],"status":"curated"},{"id":"CVE-2022-29226","cve":"CVE-2022-29226","aliases":[],"title":"Envoy: OAuth filter does not validate access tokens, so authentication can be skipped entirely","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"OAuth filter does not validate access tokens, so authentication can be skipped entirely","attack_vector":"Unauthenticated network","remediation":"Emergency Envoy upgrade; sidecar and gateway restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29226"],"status":"curated"},{"id":"CVE-2022-31481","cve":"CVE-2022-31481","aliases":["CVE-2022-31479","CVE-2022-31483","CVE-2022-31484","CVE-2022-31486","ICSA-22-153-01"],"title":"HID Mercury intelligent controllers sold by Carrier LenelS2 (LNL-X2210/X2220/X3300/X4420/4420, S2-LP-1501/1502/2500/4502): A full chain against the access-control panel generation that sits…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HID Mercury intelligent controllers sold by Carrier LenelS2","year":"2022","cvss_score":10,"severity":"critical","kev":false,"impact":"A full chain against the access-control panel generation that sits underneath a very large share of enterprise and datacenter badge systems, including Lenel OnGuard and S2 deployments: unauthenticated command execution via a crafted hostname, an unauthenticated buffer overflow in the update path, unauthenticated deletion of web-interface users, and authenticated path traversal and command injection. CISA's own summary is that an attacker gets device access, can monitor all communications to and from the panel, and can modify the onboard relays - the relays are the door strikes. So this is not badge-database tampering, it is direct electrical control of whether a door opens, plus visibility into every badge read at that panel. Someone who walks into your cage on the back of this can pull drives containing model weights and customer data, attach a console to a running node, or plug into the out-of-band switch and reach every BMC in the row. Because HID Mercury panels are OEM'd under multiple brands, many operators do not know they have them - check the board, not the badge software's vendor name.","attack_vector":"Unauthenticated network access to the panel for the worst of the set. Panels sit on the physical-security VLAN, usually in back-of-house electrical or comms closets, sometimes in the same closets tenants and contractors can reach. The multi-brand OEM situation widens exposure: a site can have Mercury boards behind three different vendors' software without a single asset record naming Mercury.","remediation":"Firmware update from the OEM whose badge is on your panel - Carrier LenelS2 published fixed firmware, and other Mercury OEMs issued their own. Identify the actual board model first (LNL-X2210 and friends, S2-LP-*), because the fix tracks the board, not the head-end software. Applying it is a security-integrator engagement with doors in local fallback during the flash, so it needs security staff on site. After patching, assume any panel exposed during the vulnerable window may have had relays or user accounts manipulated: audit the badge database, the panel user list, and the relay configuration against a known-good baseline. Long term, put the physical-security VLAN behind a firewall with an explicit allow-list from the head-end only, and enable port security on the switch ports serving panels.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-22-153-01","https://nvd.nist.gov/vuln/detail/CVE-2022-31481","https://nvd.nist.gov/vuln/detail/CVE-2022-31479","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-2825","cve":"CVE-2023-2825","aliases":[],"title":"GitLab: Unauthenticated path traversal reads arbitrary server files when an attachment sits under 5+ nested groups","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated path traversal reads arbitrary server files when an attachment sits under 5+ nested groups","attack_vector":"Network (remote)","remediation":"Control-plane: emergency upgrade; assume repository secrets were read","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2825"],"status":"curated"},{"id":"CVE-2023-31029","cve":"CVE-2023-31029","aliases":[],"title":"DGX A100 BMC: Full BMC compromise (heap buffer overflow) — worst-case out-of-band takeover","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Full BMC compromise (heap buffer overflow) — worst-case out-of-band takeover","attack_vector":"Network-adjacent unauthenticated on mgmt LAN","remediation":"Emergency: flash BMC 00.22.05+ out-of-band, audit BMC logs for compromise, rotate all mgmt credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31029","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2023-3765","cve":"CVE-2023-3765","aliases":[],"title":"MLflow: Absolute path traversal prior to 2.5.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Absolute path traversal prior to 2.5.0","attack_vector":"Unauthenticated network","remediation":"Upgrade to 2.5.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3765"],"status":"curated"},{"id":"CVE-2023-3939","cve":"CVE-2023-3939","aliases":["CVE-2023-3941","CVE-2023-3940","CVE-2023-3943","CVE-2023-3938","CVE-2023-3942"],"title":"ZKTeco-based OEM biometric access terminals (ZKTeco ProFace X, Smartec ST-FR043/ST-FR041ME and rebadged equivalents): OS command injection where every command runs as root, arbitrary file write as…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ZKTeco-based OEM biometric access terminals (ZKTeco ProFace X, Smartec ST-FR043/ST-FR041ME and rebadged equivalents)","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"OS command injection where every command runs as root, arbitrary file write as root via path traversal, arbitrary file read, a stack overflow with no stack canaries or PIE, and SQL injection allowing authentication as any user in the device database. This is the reader on the wall next to the door - the device that decides whether a face or a badge opens the hall. Root on it means an attacker opens the door at will, enrols their own biometric template, harvests the biometric templates and badge data of everyone who has ever entered (which is a privacy and regulatory problem on top of a security one), and keeps persistence on a device nobody ever patches. The rebadging is the trap: these terminals ship under many brand names, so an operator's asset list may say Smartec or a local integrator's label with no mention of ZKTeco anywhere. Once someone is through that door and into the cage, they reach drives holding model weights, server console ports, and the out-of-band management switch fronting every BMC in the row.","attack_vector":"Network access to the terminal for the injection and traversal paths - these devices sit on the physical-security VLAN and are frequently given a network address by whoever installed them with no ACL at all. Several of the flaws are also reachable by an attacker with brief physical proximity to the device, since the terminal is by definition mounted on the unsecured side of the door.","remediation":"Firmware from ZKTeco or the OEM that rebadged the device - and this is the problem, because rebadged terminals frequently never receive the upstream fix and the OEM may no longer exist. Start by physically identifying every biometric or badge terminal in the facility and determining the actual manufacturer of the board, not the label. Where a fixed firmware exists, flash it; where it does not, replace the terminal, which is a per-door hardware cost but a small one relative to what is behind the door. Regardless: isolate reader devices onto a segment that cannot reach anything else, never allow a reader to be routable from a tenant or corporate network, and if biometric templates are stored on-device, treat them as already exposed and notify accordingly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3939","https://nvd.nist.gov/vuln/detail/CVE-2023-3941","https://nvd.nist.gov/vuln/detail/CVE-2023-3943"],"status":"curated"},{"id":"CVE-2023-43654","cve":"CVE-2023-43654","aliases":["ShellTorch"],"title":"TorchServe: Unauthenticated SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe","year":"2023","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated SSRF → arbitrary model download → RCE on the serving host","attack_vector":"Unauthenticated network to an exposed TorchServe management port (8081), default config binds broadly","remediation":"Patch to 0.8.2+, restrict `allowed_urls`, and firewall 8080/8081/7070/7071 off any tenant-reachable network. Providers who publish TorchServe images must ship the hardened default","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-43654"],"status":"curated","fleet":{"ubiquity":"Common - TorchServe is the default PyTorch model server on SageMaker-style managed inference and on many neocloud inference offerings","remediation_pain":"`daemon-restart` - upgrade to TorchServe 0.8.2+ and bind `management_address` to 127.0.0.1; restarting the server drops in-flight inference but does not require a node reboot","pain_class":"node-reboot","why_fleet_wide":"The management API listened on 0.0.0.0 with no auth and accepted model URLs from *any* domain, so an unauthenticated attacker uploads a malicious model archive and gets RCE on every exposed inference host at once"}},{"id":"CVE-2023-7028","cve":"CVE-2023-7028","aliases":[],"title":"GitLab: Password reset email deliverable to an unverified address","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2023","cvss_score":10,"severity":"critical","kev":true,"impact":"[KEV] Password reset email deliverable to an unverified address -> full account takeover","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade + enforce 2FA + audit all reset events","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-7028"],"status":"curated"},{"id":"CVE-2024-11186","cve":"CVE-2024-11186","aliases":[],"title":"Arista CloudVision Portal (on-premise): An authenticated CloudVision user can take actions on managed EOS devices well beyond what their role should…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista CloudVision Portal (on-premise)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"An authenticated CloudVision user can take actions on managed EOS devices well beyond what their role should allow. CloudVision is the fabric's configuration and streaming-telemetry brain, so 'broader actions than intended' means pushing configlets to switches you were never granted. In a shared operations model — a neocloud with tenant-facing NOC accounts, or an MSP — this collapses the internal privilege model for the entire fabric.","attack_vector":"Any authenticated CloudVision Portal user on an on-premise deployment.","remediation":"Upgrade CloudVision Portal. Application upgrade on the CVP cluster; the switches keep forwarding. Afterwards review the CVP change log for configlet pushes that did not come from an authorized operator — that audit is the real work.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11186"],"status":"curated"},{"id":"CVE-2024-1709","cve":"CVE-2024-1709","aliases":[],"title":"ConnectWise ScreenConnect: Auth bypass via alternate path","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ConnectWise ScreenConnect","year":"2024","cvss_score":10,"severity":"critical","kev":true,"impact":"[KEV] Auth bypass via alternate path -> direct access to critical systems; trivially exploited at scale","attack_vector":"Network (remote)","remediation":"Control-plane: patch the RMM - compromise means code execution on every managed node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1709"],"status":"curated"},{"id":"CVE-2024-22216","cve":"CVE-2024-22216","aliases":["CVE-2023-51438"],"title":"Microchip maxView Storage Manager Redfish server (Adaptec SmartRAID / SmartHBA controllers), 3.00.23484 through 4.14.00.26064: In a default install where the Redfish server is enabled for remote…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Microchip maxView Storage Manager Redfish server (Adaptec SmartRAID / SmartHBA controllers), 3.00.23484 through…","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"In a default install where the Redfish server is enabled for remote management, the Redfish endpoint accepts unauthorized requests - read and write. The attacker gets the controller's own management API for Adaptec SmartRAID/SmartHBA: enumerate logical drives, change configuration, and reach controller update functions. Wiping or reconfiguring an array under a live training job is a fleet-availability event; the controller-update path is the worse one, because Adaptec controller firmware runs below the host OS with DMA and survives a tenant reimage, so an operator cannot honestly certify tenant handoff on a node that was exposed. CVSS 10.0 with a scope change is the vendor's own rating.","attack_vector":"Any host that can reach the maxView Redfish listener on the management network - no credentials. maxView is commonly installed by OEM server tooling and its Redfish service is on by default, so this is usually reachable from the provisioning VLAN rather than only from a hardened admin subnet.","remediation":"Upgrade maxView Storage Manager to 4.14.00.26068 or later (Microchip also back-patched 3.07.23980 and 4.07.00.25339). Software-only - restart the maxView service, no controller flash, no arrays offline. If you do not actually consume the Redfish interface, disable the maxView Redfish server outright; that is the faster fleet-wide mitigation and costs nothing. Note the OEM lag explicitly: Siemens shipped the same defect as CVE-2023-51438 in its industrial PCs a year later, and Dell/HPE/Supermicro rebadge maxView the same way - check the OEM bundle version, not just Microchip's.","references":["https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities/maxview-storage-manager-redfish-server-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2024-22216","https://nvd.nist.gov/vuln/detail/CVE-2023-51438"],"status":"curated"},{"id":"CVE-2024-22476","cve":"CVE-2024-22476","aliases":[],"title":"Intel Neural Compressor: An unauthenticated user can reach an input-validation failure in Neural Compressor and escalate. This is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Neural Compressor","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"An unauthenticated user can reach an input-validation failure in Neural Compressor and escalate. This is scored at the top of the scale, and Neural Compressor is a quantisation and optimisation service that teams commonly stand up as a shared internal endpoint next to their model registry - so an exposed instance is a pre-auth foothold beside your model weights.","attack_vector":"Network-reachable and unauthenticated where the service is exposed. Treat any internal deployment as reachable by anything else on the cluster network.","remediation":"Upgrade Intel Neural Compressor to 2.5.0 or later immediately, and put the service behind authentication and network policy regardless of version. Python package update, restart the service - no node reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22476","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01109.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-23652","cve":"CVE-2024-23652","aliases":[],"title":"BuildKit: \"Leaky Vessels\": RUN --mount empty-file removal can delete arbitrary host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"\"Leaky Vessels\": RUN --mount empty-file removal can delete arbitrary host files; complete builder compromise","attack_vector":"Malicious Dockerfile or frontend submitted to a shared builder","remediation":"Upgrade BuildKit/buildx everywhere; treat any shared tenant builder as compromised and rebuild it","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23652"],"status":"curated"},{"id":"CVE-2024-2912","cve":"CVE-2024-2912","aliases":[],"title":"BentoML: Insecure deserialization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Insecure deserialization → RCE from a crafted POST","attack_vector":"Unauthenticated network to the BentoML serving port","remediation":"Upgrade. A default BentoML service is an unauthenticated RCE endpoint pre-patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-2912"],"status":"curated"},{"id":"CVE-2024-3400","cve":"CVE-2024-3400","aliases":[],"title":"Palo Alto PAN-OS: GlobalProtect arbitrary file creation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2024","cvss_score":10,"severity":"critical","kev":true,"impact":"[KEV] GlobalProtect arbitrary file creation -> command injection, unauthenticated root on the firewall","attack_vector":"Network (remote)","remediation":"Control-plane: emergency hotfix; full forensics and rotate every secret on-box","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3400"],"status":"curated"},{"id":"CVE-2024-42479","cve":"CVE-2024-42479","aliases":[],"title":"llama.cpp (RPC backend): Unsafe `data` pointer in `rpc_tensor`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (RPC backend)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Unsafe `data` pointer in `rpc_tensor` → arbitrary address write","attack_vector":"Unauthenticated network to the llama.cpp RPC port on a distributed inference setup","remediation":"Rebuild; the RPC backend has no authentication and must never be tenant-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42479"],"status":"curated"},{"id":"CVE-2024-44984","cve":"CVE-2024-44984","aliases":[],"title":"Linux bnxt_en driver (XDP_REDIRECT double DMA unmap): A double DMA unmap in the XDP_REDIRECT path. Double-unmapping a DMA region is a memory-corruption primitive…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (XDP_REDIRECT double DMA unmap)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"A double DMA unmap in the XDP_REDIRECT path. Double-unmapping a DMA region is a memory-corruption primitive that touches the IOMMU mapping state, which is precisely the mechanism that keeps a device from reaching memory it should not. Anywhere you run XDP-based load balancing or packet steering in front of inference serving — a common pattern — this is live code.","attack_vector":"Traffic through the driver's XDP_REDIRECT path on a host with an XDP program attached.","remediation":"Kernel/driver upgrade plus host reboot. Interim: detach XDP programs from Broadcom NICs, which is a live change but costs you whatever the XDP program was doing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44984"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45409","cve":"CVE-2024-45409","aliases":[],"title":"GitLab (ruby-saml): Ruby-SAML does not properly verify the SAML Response signature","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab (ruby-saml)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Ruby-SAML does not properly verify the SAML Response signature -> forge an assertion as any user, incl. admin","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade; treat all SAML sessions as suspect","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45409"],"status":"curated"},{"id":"CVE-2024-6886","cve":"CVE-2024-6886","aliases":[],"title":"Gitea: Stored cross-site scripting in Gitea 1.22.0","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Gitea","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Stored cross-site scripting in Gitea 1.22.0 -> session theft from maintainers","attack_vector":"Network (remote)","remediation":"Control-plane: Gitea upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6886"],"status":"curated"},{"id":"CVE-2024-7591","cve":"CVE-2024-7591","aliases":[],"title":"Progress Kemp LoadMaster (including Multi-Tenancy edition): TENANT ISOLATION: a request handler fails to validate its input before passing it to a system call, giving an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Progress Kemp LoadMaster (including Multi-Tenancy edition)","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: a request handler fails to validate its input before passing it to a system call, giving an attacker OS command execution on the LoadMaster with no authentication at all. This explicitly affects the Multi-Tenancy edition — the product line built to let multiple tenants share one LoadMaster — so a single unauthenticated request can compromise the appliance underneath every tenant's virtual services on that instance.","attack_vector":"Fully remote and unauthenticated — a crafted request to the vulnerable API endpoint is sufficient, no login required.","remediation":"Software upgrade to the fixed LoadMaster/Multi-Tenancy release per Kemp's advisory. Patch immediately given the unauthenticated, maximum-severity nature of this bug; if this instance is shared across tenants, treat any exposure window as a potential full-tenant-boundary breach and audit for signs of compromise, not just apply the patch.","references":["https://insinuator.net/2024/11/vulnerability-disclosure-command-injection-in-kemp-loadmaster-load-balancer-cve-2024-7591","https://support.kemptechnologies.com/hc/en-us/articles/29196371689613-LoadMaster-Security-Vulnerability-CVE-2024-7591"],"status":"curated"},{"id":"CVE-2024-8525","cve":"CVE-2024-8525","aliases":["CVE-2024-8526","ICSA-24-326-01"],"title":"Automated Logic WebCTRL 7.0 / WebCTRL Premium Server / Carrier i-Vu building automation server: Unauthenticated file upload leading to remote command execution on the BAS server itself - the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Automated Logic WebCTRL 7.0 / WebCTRL Premium Server / Carrier i-Vu building automation server","year":"2024","cvss_score":10,"severity":"critical","kev":false,"impact":"Unauthenticated file upload leading to remote command execution on the BAS server itself - the machine that holds the graphics, the schedules, the trends and, crucially, the write authority over every controller in the building. A 10.0 here means an attacker who reaches the web port gets to be the building operator. They can command CRAH fans to zero, raise chilled-water setpoints, disable economizers, force valves shut, and edit the schedules so the change survives a reboot and looks intentional. Against a hall of 40-140 kW GPU racks that is minutes to thermal shutdown and a hardware-damage risk, and because the same server owns the trend and alarm database, the attacker can also rewrite what the operator sees while it happens. Owning WebCTRL is strictly better than owning any single controller, and it does not require a single credential from the compute network.","attack_vector":"Unauthenticated HTTP POST to the WebCTRL server. WebCTRL is a Windows/Tomcat application, and it is very commonly published beyond the facility VLAN because facilities staff and the controls contractor want browser access - that is the exposure that turns this from 'facility network' to 'internet-exposed via a badly-placed remote-access box or a public DNS entry'. If your site has a WebCTRL login page reachable from anything other than a jump host, treat that as already compromised until proven otherwise.","remediation":"Software upgrade on the BAS server - no controller firmware, no cooling downtime, so this one is genuinely fixable in a normal change window and there is no excuse to defer it. Move to a fixed WebCTRL/i-Vu release per Carrier's advisory. Then do the thing that should have been done first: take the WebCTRL server off any interface reachable from the internet or the corporate LAN, require VPN plus MFA to a jump host, and confirm the Tomcat service account is not a domain administrator. In a leased colo the WebCTRL server is the landlord's and often shared across the whole building - ask for its version and its network position, and treat 'it is behind our VPN' as an unverified claim until you see the rule.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-24-326-01","https://nvd.nist.gov/vuln/detail/CVE-2024-8525","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated"},{"id":"CVE-2025-0505","cve":"CVE-2025-0505","aliases":[],"title":"Arista CloudVision (Zero Touch Provisioning): Zero Touch Provisioning can be abused to obtain admin privileges on the CloudVision system itself. ZTP is by…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista CloudVision (Zero Touch Provisioning)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Zero Touch Provisioning can be abused to obtain admin privileges on the CloudVision system itself. ZTP is by design an unauthenticated-ish onboarding path — a new switch shows up and asks for its config — so this turns 'plug a device into the provisioning VLAN' into 'own the fabric controller'. For anyone doing rack-and-stack at scale, which is every GPU buildout, the ZTP network is live constantly.","attack_vector":"A device that can participate in ZTP against the CloudVision instance — i.e. anything on the provisioning network. Physical or logical access to that VLAN is the whole requirement.","remediation":"Upgrade CloudVision. Beyond the patch, treat the ZTP/provisioning VLAN as a privileged network: separate it from the general management network, keep it shut down when not actively provisioning, and require MAC/serial allowlisting. Those are config and process changes and they are what actually keeps this closed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0505"],"status":"curated"},{"id":"CVE-2025-15036","cve":"CVE-2025-15036","aliases":[],"title":"MLflow (`extract_archive_to_dir`): Path traversal in the dbconnect artifact cache","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (`extract_archive_to_dir`)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Path traversal in the dbconnect artifact cache","attack_vector":"Customer-supplied artifact archive","remediation":"Upgrade immediately — maximum severity","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-15036"],"status":"curated"},{"id":"CVE-2025-29270","cve":"CVE-2025-29270","aliases":[],"title":"Deep Sea Electronics DSE855 generator communications gateway v1.1.0-v1.1.26 (realtime.cgi): Incorrect access control on the realtime.cgi endpoint hands an attacker the admin panel and complete…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Deep Sea Electronics DSE855 generator communications gateway v1.1.0-v1.1.26 (realtime.cgi)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Incorrect access control on the realtime.cgi endpoint hands an attacker the admin panel and complete control of the device with no credentials. The DSE855 is the Ethernet gateway that fronts DSE generator and transfer-switch controllers, exposing them to the BMS and to remote monitoring. Control of it means control of the interface between the building's standby power and everything that watches it. Concretely: an attacker can reconfigure or disable the monitoring path so a genset failure goes unreported, and can reach the controller behind it. For a GPU datacenter this is the second half of the thermal problem - the cooling plant is the thing that stops when utility power drops and the generators do not pick up. A hall that loses chillers on a failed transfer is on the same minutes-to-thermal-shutdown clock as one whose CRAHs were commanded off, except now the UPS is draining under a 40-140 kW-per-rack load that it was probably not sized to ride through for long.","attack_vector":"Unauthenticated HTTP on the facility network - network-adjacent, no credentials. DSE855 units sit in generator yards, switchgear rooms and electrical closets on the building network. They are a classic 'installed by the generator contractor, never inventoried by IT' device, and because they exist to provide remote monitoring they are disproportionately likely to be reachable from outside the building through whatever remote-access arrangement the contractor set up.","remediation":"Firmware update from Deep Sea Electronics for the DSE855. This is a small standalone gateway, so the flash itself is quick and does not require taking generators out of service - but it does need someone with physical or network access to the unit and a maintenance window on the monitoring path, and the generator contractor usually owns that relationship rather than the datacenter operator. Alongside patching: put generator and switchgear network devices on an isolated segment, remove any internet path, and specifically audit the generator contractor's remote-access arrangement, which is frequently a consumer-grade router or a cellular modem nobody in IT knows about. In a leased site the gensets and their gateways are the landlord's - ask who can reach them remotely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29270","https://blog.byteray.co.uk/shadow-entry-discovery-of-authentication-bypass-vulnerability-in-dse855-communications-device-938e35d4b361"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-32444","cve":"CVE-2025-32444","aliases":[],"title":"vLLM (Mooncake ZMQ/TCP): Unsafe deserialization exposed on all interfaces","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (Mooncake ZMQ/TCP)","year":"2025","cvss_score":10,"severity":"critical","kev":false,"impact":"Unsafe deserialization exposed on all interfaces → RCE","attack_vector":"Unauthenticated network from any host on the cluster fabric","remediation":"Upgrade to 0.8.5+. Highest-severity vLLM issue; assume any pre-0.8.5 disaggregated deployment is compromised","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32444"],"status":"curated"},{"id":"CVE-2025-34028","cve":"CVE-2025-34028","aliases":[],"title":"Commvault Command Center: Unauthenticated ZIP upload + path traversal","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Commvault Command Center","year":"2025","cvss_score":10,"severity":"critical","kev":true,"impact":"[KEV] Unauthenticated ZIP upload + path traversal -> RCE via a malicious JSP","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately; the backup control plane is a crown-jewel target","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-34028"],"status":"curated"},{"id":"CVE-2026-10520","cve":"CVE-2026-10520","aliases":[],"title":"Ivanti Sentry: OS command injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Sentry","year":"2026","cvss_score":10,"severity":"critical","kev":true,"impact":"[KEV] OS command injection -> unauthenticated root-level remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch to R10.5.2/R10.6.2/R10.7.1; rebuild the appliance if it was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-10520"],"status":"curated"},{"id":"CVE-2026-45829","cve":"CVE-2026-45829","aliases":[],"title":"ChromaDB: Pre-authentication code injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB","year":"2026","cvss_score":10,"severity":"critical","kev":false,"impact":"Pre-authentication code injection → arbitrary code execution","attack_vector":"Unauthenticated network to the Chroma server","remediation":"Upgrade. Maximum severity, no auth required — any tenant-reachable Chroma is fully compromised","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45829"],"status":"curated"},{"id":"CVE-2026-74280","cve":"CVE-2026-74280","aliases":[],"title":"Linux crypto driver for Marvell OCTEON TX: The scatter-gather cleanup path in the Marvell OCTEON TX crypto driver uses the wrong loop index, so DMA…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux crypto driver for Marvell OCTEON TX","year":"2026","cvss_score":10,"severity":"critical","kev":false,"impact":"The scatter-gather cleanup path in the Marvell OCTEON TX crypto driver uses the wrong loop index, so DMA mappings are torn down against the wrong entries. Getting DMA cleanup wrong on an accelerator is exactly the primitive that undermines IOMMU-based device isolation — the mapping that should have been revoked stays live, or a mapping belonging to something else is revoked. OCTEON is used both as a DPU and as an inline crypto/offload engine in storage and network appliances.","attack_vector":"Local, through the crypto API on a host with an OCTEON TX accelerator.","remediation":"Kernel/driver upgrade plus host reboot. Nothing to flash. If you run OCTEON accelerators in a shared-tenancy role, verify IOMMU is enabled and enforcing on those devices as a standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74280"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74475","cve":"CVE-2026-74475","aliases":[],"title":"Linux VXLAN driver (neighbour hardware address read in route_shortcircuit): TENANT ISOLATION: `route_shortcircuit()` reads a neighbour's hardware address without taking the seqlock that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux VXLAN driver (neighbour hardware address read in route_shortcircuit)","year":"2026","cvss_score":10,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: `route_shortcircuit()` reads a neighbour's hardware address without taking the seqlock that protects it, so it can observe a torn or partially-updated MAC address while the neighbour subsystem is rewriting it. In an overlay, the destination MAC is what decides which VTEP — and therefore which tenant's segment — a frame is delivered to. A torn read there is not just a memory-safety problem; it is a frame going somewhere the forwarding logic did not intend.","attack_vector":"Concurrent neighbour updates alongside VXLAN forwarding on an affected host or software VTEP. Triggerable by ordinary overlay traffic combined with neighbour churn, which tenants generate routinely.","remediation":"Kernel upgrade plus host reboot, or NOS image upgrade plus switch reload on Linux-based switches. Bundle with the rest of the 2026 VXLAN batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74475"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-21872","cve":"CVE-2021-21872","aliases":["TALOS-2021-1312"],"title":"Lantronix PremierWave 2050 console server (Web Manager): An attacker who can log into the web management console gets a shell on the console server itself, running…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix PremierWave 2050 console server (Web Manager)","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"An attacker who can log into the web management console gets a shell on the console server itself, running with the same privileges as the web daemon. From there they can pivot to every serial-attached device the box terminates, and use the console server as a jump box into the rest of the OOB network.","attack_vector":"Needs an authenticated session to the Web Manager (any role), then sends a crafted HTTP request to the Diagnostics: Traceroute page. The traceroute target field isn't sanitized before being handed to a shell, so shell metacharacters turn it into arbitrary command execution.","remediation":"Firmware upgrade required (Lantronix has patched builds past 8.9.0.0R4) plus a reboot of each unit; no config workaround exists since the flaw is in the diagnostics handler itself. Roll out per-device, one console server at a time — each reboot drops active serial sessions to whatever racks it terminates.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2021-1312"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-21883","cve":"CVE-2021-21883","aliases":["TALOS-2021-1327"],"title":"Lantronix PremierWave 2050 console server (Web Manager): Same class of bug as the Traceroute injection on this device, reached through the Ping diagnostic instead —…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix PremierWave 2050 console server (Web Manager)","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Same class of bug as the Traceroute injection on this device, reached through the Ping diagnostic instead — full command execution on the console server, with downstream reach into every serial line it terminates.","attack_vector":"Any authenticated Web Manager user submits a crafted host value to the Diagnostics: Ping function; the value flows unsanitized into a shell command.","remediation":"Firmware upgrade past 8.9.0.0R4 and reboot. Same rollout cost as the Traceroute bug — treat both as one patch cycle per unit rather than two separate maintenance windows.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2021-1327"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-28476","cve":"CVE-2021-28476","aliases":[],"title":"Microsoft Hyper-V: vmswitch fails to validate guest OID requests - guest reads arbitrary host kernel memory or crashes the host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"vmswitch fails to validate guest OID requests - guest reads arbitrary host kernel memory or crashes the host; the highest-rated Hyper-V escape class","attack_vector":"Tenant VM guest","remediation":"Windows update + host reboot with live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28476"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-36782","cve":"CVE-2021-36782","aliases":[],"title":"Rancher: Cluster owners, members and even base users retrieve plaintext credentials via the Kubernetes API","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Cluster owners, members and even base users retrieve plaintext credentials via the Kubernetes API","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; rotate every credential Rancher stores, including cloud and registry credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36782"],"status":"curated"},{"id":"CVE-2021-36783","cve":"CVE-2021-36783","aliases":[],"title":"Rancher: Insufficiently protected credentials let project members read passwords and API tokens","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2021","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Insufficiently protected credentials let project members read passwords and API tokens","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; rotate all stored credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36783"],"status":"curated"},{"id":"CVE-2022-24768","cve":"CVE-2022-24768","aliases":[],"title":"Argo CD: Improper access control allows a low-privileged user to escalate to Argo CD admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Improper access control allows a low-privileged user to escalate to Argo CD admin","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24768"],"status":"curated"},{"id":"CVE-2022-43757","cve":"CVE-2022-43757","aliases":[],"title":"Rancher: Cleartext credential storage lets managed-cluster users read credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Cleartext credential storage lets managed-cluster users read credentials","attack_vector":"Any user on a managed cluster","remediation":"Upgrade Rancher; rotate all credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43757"],"status":"curated"},{"id":"CVE-2023-22647","cve":"CVE-2023-22647","aliases":[],"title":"Rancher: Standard users manipulate Kubernetes secrets in the local (management) cluster","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Standard users manipulate Kubernetes secrets in the local (management) cluster","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; audit local-cluster secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22647"],"status":"curated"},{"id":"CVE-2023-22651","cve":"CVE-2023-22651","aliases":[],"title":"Rancher: Update-logic failure misconfigures Rancher's admission webhook, disabling the validation that enforces tenant…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Update-logic failure misconfigures Rancher's admission webhook, disabling the validation that enforces tenant boundaries","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher; verify the webhook configuration explicitly after every upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22651"],"status":"curated"},{"id":"CVE-2023-32191","cve":"CVE-2023-32191","aliases":[],"title":"RKE / Rancher (k8s control plane): full-cluster-state configmap in kube-system readable by non-admins","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"RKE / Rancher (k8s control plane)","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"full-cluster-state configmap in kube-system readable by non-admins -> escalate to cluster-admin","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade RKE; restrict RBAC on kube-system configmaps","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32191"],"status":"curated"},{"id":"CVE-2023-40029","cve":"CVE-2023-40029","aliases":[],"title":"Argo CD: Cluster secrets stored in the last-applied-configuration annotation are readable by anyone with get access on…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Cluster secrets stored in the last-applied-configuration annotation are readable by anyone with get access on the secret","attack_vector":"Cluster user with namespace access to argocd","remediation":"Rolling Argo CD upgrade; rotate every cluster credential Argo CD holds","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40029"],"status":"curated"},{"id":"CVE-2024-20432","cve":"CVE-2024-20432","aliases":[],"title":"Cisco Nexus Dashboard Fabric Controller (REST API / web UI): A low-privileged NDFC user — the sort of read-mostly account you hand to an NOC or a tenant liaison — gets…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus Dashboard Fabric Controller (REST API / web UI)","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"A low-privileged NDFC user — the sort of read-mostly account you hand to an NOC or a tenant liaison — gets command injection on the fabric controller. NDFC holds the credentials for and pushes config to every switch it manages, so this is a straight path from a minor account to control of the whole leaf/spine build.","attack_vector":"Authenticated but low-privileged, remote. Any valid NDFC login is enough.","remediation":"Upgrade NDFC. Controller-side software upgrade, data plane unaffected. Afterwards rotate the device credentials NDFC stores, because those are what an attacker would have taken.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-ndfc-raci-T46k3jnN","https://nvd.nist.gov/vuln/detail/CVE-2024-20432"],"status":"curated"},{"id":"CVE-2024-24594","cve":"CVE-2024-24594","aliases":[],"title":"ClearML web server: XSS","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML web server","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"XSS → code execution in the operator's session","attack_vector":"Attacker-controlled experiment metadata rendered in the UI","remediation":"Upgrade; tenant-supplied experiment names become operator-plane payloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24594"],"status":"curated"},{"id":"CVE-2024-42327","cve":"CVE-2024-42327","aliases":[],"title":"Zabbix: SQL injection in CUser::addRelatedObjects reachable by ANY non-admin account with API access","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"SQL injection in CUser::addRelatedObjects reachable by ANY non-admin account with API access -> full DB compromise","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade; Zabbix commonly stores datacenter/IPMI credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42327"],"status":"curated"},{"id":"CVE-2024-9264","cve":"CVE-2024-9264","aliases":[],"title":"Grafana: SQL Expressions passes user input to duckdb unsanitized","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2024","cvss_score":9.9,"severity":"critical","kev":false,"impact":"SQL Expressions passes user input to duckdb unsanitized -> command injection and local file inclusion (VIEWER+)","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; remove the duckdb binary from the Grafana image","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9264"],"status":"curated"},{"id":"CVE-2025-10725","cve":"CVE-2025-10725","aliases":[],"title":"Red Hat OpenShift AI (notebook plane): A low-privileged data-scientist account can escalate to full cluster compromise","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Red Hat OpenShift AI (notebook plane)","year":"2025","cvss_score":9.9,"severity":"critical","kev":false,"impact":"A low-privileged data-scientist account can escalate to full cluster compromise","attack_vector":"Notebook user inside the managed AI platform","remediation":"Patch the platform. Direct tenant→provider escalation on a managed GPU platform","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10725"],"status":"curated"},{"id":"CVE-2025-49844","cve":"CVE-2025-49844","aliases":[],"title":"Redis: \"RediShell\" - authenticated user crafts a Lua script to trigger a use-after-free","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2025","cvss_score":9.9,"severity":"critical","kev":false,"impact":"\"RediShell\" - authenticated user crafts a Lua script to trigger a use-after-free -> RCE on the host","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade of every Redis (scheduler/queue/session); restrict EVAL via ACL","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49844"],"status":"curated"},{"id":"CVE-2025-54381","cve":"CVE-2025-54381","aliases":[],"title":"BentoML (file upload): SSRF in the file-upload path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (file upload)","year":"2025","cvss_score":9.9,"severity":"critical","kev":false,"impact":"SSRF in the file-upload path","attack_vector":"Authenticated or unauthenticated request depending on deployment","remediation":"Upgrade to 1.4.19+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54381"],"status":"curated"},{"id":"CVE-2026-7374","cve":"CVE-2026-7374","aliases":[],"title":"KubeVirt: Improper symlink validation in virt-handler lets a user with edit rights in one namespace escape to the host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":9.9,"severity":"critical","kev":false,"impact":"Improper symlink validation in virt-handler lets a user with edit rights in one namespace escape to the host; the most severe KubeVirt issue","attack_vector":"Cluster user with namespace access","remediation":"Emergency KubeVirt upgrade; virt-handler DaemonSet rollout on every node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7374"],"status":"curated"},{"id":"CVE-2015-8011","cve":"CVE-2015-8011","aliases":[],"title":"lldpd (lldp_decode, management addresses): Buffer overflow in lldpd's LLDP decoder via large management addresses and TLV boundaries, allowing daemon…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"lldpd (lldp_decode, management addresses)","year":"2015","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Buffer overflow in lldpd's LLDP decoder via large management addresses and TLV boundaries, allowing daemon crash and possibly code execution. Old, but included because lldpd is one of those daemons that ships inside embedded switch and appliance images and stays frozen at whatever version the vendor picked years ago — the CVE date tells you nothing about whether your fabric is running it.","attack_vector":"Unauthenticated, adjacent — a crafted LLDP frame.","remediation":"Upgrade lldpd past 0.8.0 and restart. On embedded NOSes and appliances, check the shipped lldpd version explicitly rather than assuming a modern image implies a modern lldpd. Companion crash issue: CVE-2015-8012.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-8011","https://nvd.nist.gov/vuln/detail/CVE-2015-8012"],"status":"curated"},{"id":"CVE-2016-4325","cve":"CVE-2016-4325","aliases":["VU#785823"],"title":"Lantronix xPrintServer: The device ships with a hardcoded root account baked into every unit of a given firmware line. Anyone who can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix xPrintServer","year":"2016","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The device ships with a hardcoded root account baked into every unit of a given firmware line. Anyone who can reach the management interface gets root on the box without needing to guess or phish a credential — it's the same key for every deployed unit.","attack_vector":"No authentication needed beyond network reachability to the device; the credential is embedded in firmware and identical across all units running the affected build.","remediation":"Firmware flash to 5.0.1-65 or later on every affected unit — this isn't something a config change or password rotation fixes, since the account is compiled into the image. Budget one flash-and-reboot cycle per device; xPrintServer's main job (serial/print bridging) is unavailable during the flash.","references":["http://www.kb.cert.org/vuls/id/785823"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2016-4375","cve":"CVE-2016-4375","aliases":[],"title":"HPE iLO3 / iLO4: Multiple unspecified flaws allowing remote information disclosure, data modification and DoS on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO3 / iLO4","year":"2016","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Multiple unspecified flaws allowing remote information disclosure, data modification and DoS on the management controller","attack_vector":"Network","remediation":"iLO firmware update (iLO3 <1.88, iLO4 <2.44); the \"unspecified\" advisory style means operators cannot risk-assess individual issues and must patch blind","references":["https://nvd.nist.gov/vuln/detail/CVE-2016-4375"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2017-16748","cve":"CVE-2017-16748","aliases":["ICSA-18-191-03","ICSA-19-022-01"],"title":"Tridium Niagara AX (<=3.8) and Niagara 4 (<=4.4) framework: Log into the Niagara platform with a disabled account name and a blank password and you get administrator.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tridium Niagara AX (<=3.8) and Niagara 4 (<=4.4) framework","year":"2017","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Log into the Niagara platform with a disabled account name and a blank password and you get administrator. Niagara is the integration layer a very large share of datacenters use to tie CRAC/CRAH, chillers, generators, metering and sometimes access control into one supervisory system, so administrator on the Niagara station is administrator over the whole facility control surface at once. That means arbitrary writes to cooling setpoints and fan commands across every integrated subsystem, the ability to disable alarms and schedules, and access to the credential store Niagara keeps for its downstream drivers - which is how a BMS compromise turns into compromise of the chiller, the ATS controller and the metering network in one step. For a GPU operator the headline is simple: one blank password and the hall's thermal envelope is under attacker control, with minutes of margin before accelerators shut down.","attack_vector":"Unauthenticated login against the Niagara station's web or Fox interface on the facility network. Niagara stations are one of the most consistently internet-exposed classes of building controller in existence - integrators publish them for remote support and forget - so treat internet exposure as likely rather than exceptional until you have checked your own external attack surface for the Niagara Fox port and the station web UI.","remediation":"Patchable via a Niagara framework upgrade (AX 3.8U1 / Niagara 4.4U1 or later), performed by the systems integrator who owns the station. This is a software upgrade on the supervisor plus a JACE controller update, so it needs a maintenance window and a contractor but not a cooling outage. Do it, then audit the account list for disabled-but-present accounts, and pull the station off any internet-facing interface. Because Niagara stations are so often integrator-managed rather than operator-managed, the harder task is organisational: find out who actually holds the platform credentials for your station, whether the integrator has a permanent remote path in, and whether that path is MFA'd. In a leased colo the station belongs to the landlord and typically serves the entire building - demand the framework version and the remote-access architecture in writing.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-18-191-03","https://nvd.nist.gov/vuln/detail/CVE-2017-16748"],"status":"curated"},{"id":"CVE-2017-5689","cve":"CVE-2017-5689","aliases":["Silent Bob is Silent"],"title":"Intel Active Management Technology / Standard Manageability: An authentication bypass in the AMT web interface: sending an empty response hash is accepted as valid, so an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Active Management Technology / Standard Manageability","year":"2017","cvss_score":9.8,"severity":"critical","kev":true,"impact":"An authentication bypass in the AMT web interface: sending an empty response hash is accepted as valid, so an unauthenticated network attacker gets full AMT administrative control. AMT provides out-of-band power control, KVM and virtual media below the OS, so this is total control of the machine from the management network, invisible to anything running on the host. Listed in CISA's Known Exploited Vulnerabilities catalog.","attack_vector":"Any attacker who can reach the AMT ports (16992/16993/16994/16995, and 623/664) on a provisioned machine. If your management network is flat or reachable from tenant VLANs, that is everyone.","remediation":"Update Intel CSME/AMT firmware via the OEM. Where firmware is unavailable, unprovision AMT and block the AMT ports at the network layer - that is the mitigation that actually deploys on the same day. Firmware update requires an OEM package, a drain and a reboot. Any machine exposed while vulnerable should be treated as compromised at the firmware level, since AMT access permits persistent implantation below the OS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5689","https://downloadmirror.intel.com/26754/eng/INTEL-SA-00075%20Mitigation%20Guide-Rev%201.1.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2017-8979","cve":"CVE-2017-8979","aliases":[],"title":"HPE iLO2: Authentication bypass and code execution in iLO2 firmware 2.29","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO2","year":"2017","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass and code execution in iLO2 firmware 2.29","attack_vector":"Network, unauthenticated","remediation":"iLO2 is EOL — remediation is decommissioning or hard network isolation, not patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-8979"],"status":"curated"},{"id":"CVE-2018-0314","cve":"CVE-2018-0314","aliases":[],"title":"Cisco NX-OS / FXOS (Cisco Fabric Services): Unauthenticated remote code execution as root through Cisco Fabric Services, the inter-switch distribution…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS / FXOS (Cisco Fabric Services)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution as root through Cisco Fabric Services, the inter-switch distribution protocol. CFS is on by default on many platforms and speaks between switches, so a single compromised switch or a host that can spoof CFS reaches every other switch that trusts the same distribution domain. Part of a 2018 cluster (CVE-2018-0304/0308/0310/0312/0314) that shares this exposure.","attack_vector":"Unauthenticated, remote — the attacker must be able to send CFS messages, which in an unsegmented fabric means any host on a VLAN where CFS-over-IP is enabled.","remediation":"NX-OS upgrade plus reload. Immediately: disable CFS distribution and CFS-over-IP where you do not use it (`no cfs distribute`, `no cfs ipv4 distribute`) — live config, no reload, and it closes the whole 2018 family at once.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0314","https://nvd.nist.gov/vuln/detail/CVE-2018-0304","https://nvd.nist.gov/vuln/detail/CVE-2022-20624"],"status":"curated"},{"id":"CVE-2018-12031","cve":"CVE-2018-12031","aliases":[],"title":"Eaton Intelligent Power Manager v1.6 - node_upgrade_srv.js firmware parameter: Local file inclusion through directory traversal on the firmware parameter of the node upgrade service. The…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Intelligent Power Manager v1.6 - node_upgrade_srv.js firmware parameter","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Local file inclusion through directory traversal on the firmware parameter of the node upgrade service. The affected code path is the one that distributes firmware to managed power devices, so this is an attacker standing in the middle of your UPS firmware supply chain.","attack_vector":"Remote to the IPM server, unauthenticated per the advisory.","remediation":"Upgrade IPM well past 1.6 - anything on 1.6 is also carrying the 2020 and 2021 unauthenticated RCEs. Gate firmware distribution to power devices behind change control regardless of the platform you use.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12031"],"status":"curated"},{"id":"CVE-2018-1207","cve":"CVE-2018-1207","aliases":[],"title":"Dell iDRAC7/8: CGI injection giving unauthenticated remote code execution as root on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7/8","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"CGI injection giving unauthenticated remote code execution as root on the BMC","attack_vector":"Network, unauthenticated","remediation":"Firmware update to 2.52.52.52+; nodes still on older iDRAC7/8 firmware are the most common legacy hole in a mixed-vintage GPU fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1207"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-12171","cve":"CVE-2018-12171","aliases":["INTEL-SA-00149"],"title":"Intel Baseboard Management Controller firmware before 1.43.91f76955 (Intel server boards and systems): An unprivileged user can execute arbitrary code or force denial of service on the BMC. Owning…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Baseboard Management Controller firmware before 1.43.91f76955 (Intel server boards and systems)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An unprivileged user can execute arbitrary code or force denial of service on the BMC. Owning the BMC is owning the node: virtual media boot, host power control, KVM, SOL, firmware update paths for BIOS and ME, and the SMBus. It is below-the-OS persistence in the strongest sense - the BMC has its own flash and its own OS, so nothing you do to the host image removes an implant there, and the node carries it into the next tenant. It also gives an attacker a fleet-wide physical-consequence lever: mass power-off of every node they can reach on the management network.","attack_vector":"Network access to the BMC without valid credentials. Anyone who can route to the out-of-band management network - which in practice includes anything that reaches an internet-exposed or flat-VLAN BMC, and any tenant if the OOB network is not fully separated.","remediation":"BMC firmware update to 1.43.91f76955 or later, from Intel or the board ODM (Quanta, Wiwynn, Supermicro on Intel reference designs). BMC updates do not need a host reboot, so this is one you can roll without draining jobs - do it fleet-wide. In parallel: never expose BMCs to the internet, put them behind a jump host on a dedicated VRF, rotate to unique per-node credentials, and disable the host-side KCS interface.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12171","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00149.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-12327","cve":"CVE-2018-12327","aliases":[],"title":"ntpq / ntpdc (NTP 4.2.8p11 client utilities): Stack buffer overflow in the ntpq and ntpdc command-line tools via a long argument, giving code execution or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ntpq / ntpdc (NTP 4.2.8p11 client utilities)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow in the ntpq and ntpdc command-line tools via a long argument, giving code execution or privilege escalation. The interesting case for a cluster operator is automation: monitoring scripts that shell out to ntpq with a hostname taken from inventory turn an inventory-poisoning bug into code execution on the monitoring host.","attack_vector":"Local, via a long argument to ntpq/ntpdc — reachable wherever these tools are invoked with externally influenced arguments.","remediation":"Upgrade the ntp package. No service restart needed for the client tools; nothing to reboot. Audit any monitoring or automation that passes untrusted strings to ntpq.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12327"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-18202","cve":"CVE-2018-18202","aliases":[],"title":"QLogic 4Gb Fibre Channel 5.5.2.6.0 and 4/8Gb SAN 7.10.1.20.0 switch modules for IBM BladeCenter: Three undocumented accounts - support, diags and prom - each with a fixed password, baked into the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"QLogic 4Gb Fibre Channel 5.5.2.6.0 and 4/8Gb SAN 7.10.1.20.0 switch modules for IBM BladeCenter","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Three undocumented accounts - support, diags and prom - each with a fixed password, baked into the FC switch module firmware. Anyone who knows them owns the switch module: zoning, port state, firmware. The reason this belongs in a modern GPU-datacenter database is not BladeCenter itself but the pattern - embedded FC switch modules and SAS/FC expander boards inherited with second-hand chassis carry vendor service accounts that no amount of operator password hygiene touches, and nothing in a normal build process ever looks for them.","attack_vector":"Anyone who can reach the module's management interface (telnet/SSH/web) on the chassis management network. Credentials are public.","remediation":"Unfixable by configuration - the accounts are in firmware. Either upgrade to a firmware release where they are removed, if one exists for your module, or retire the module. Practically, for inherited or second-hand chassis: treat every embedded switch/expander module as carrying vendor backdoor accounts until proven otherwise, keep chassis management on an isolated segment no tenant or workload VLAN can route to, and make 'scan for vendor service accounts' part of hardware intake rather than something you do after an incident.","references":["http://misteralfa-hack.blogspot.com/2018/10/ibm-bladecenter-qlogic-4g-fibre-channel.html","https://nvd.nist.gov/vuln/detail/CVE-2018-18202"],"status":"curated"},{"id":"CVE-2018-20687","cve":"CVE-2018-20687","aliases":[],"title":"Raritan CommandCenter Secure Gateway (CC-SG), before 8.0.0: TENANT ISOLATION: CC-SG is Raritan's single-pane-of-glass gateway that brokers KVM and serial console access…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Raritan CommandCenter Secure Gateway (CC-SG), before 8.0.0","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: CC-SG is Raritan's single-pane-of-glass gateway that brokers KVM and serial console access across a whole fleet of Dominion KX/SX switches. An unauthenticated attacker who reaches its web-services endpoint can send a crafted XML request with a malicious DTD, read arbitrary files off the gateway, or force it to make outbound requests (SSRF) into whatever internal network segments the gateway can reach — which by design is every rack it brokers console access to.","attack_vector":"Fully unauthenticated, remote — a crafted XML request to the CommandCenterWebServices endpoint is enough; no login required.","remediation":"Software upgrade to CC-SG 8.0.0 or later. This is a single appliance (or small HA pair) rather than a per-rack device, so the upgrade footprint is small, but treat it as urgent — the gateway sits in front of console/KVM access to the entire managed fleet, so an unauthenticated file-read/SSRF bug there is a direct path toward every rack it manages.","references":["http://packetstormsecurity.com/files/155359/Raritan-CommandCenter-Secure-Gateway-XML-Injection.html","http://seclists.org/fulldisclosure/2019/Nov/11"],"status":"curated"},{"id":"CVE-2018-7243","cve":"CVE-2018-7243","aliases":["SEVD-2018-074-01"],"title":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS): The card's integrated web server has a broken authorization check, letting a remote attacker get full…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The card's integrated web server has a broken authorization check, letting a remote attacker get full administrative access to the UPS/STS management interface without valid credentials — including whatever load-shedding and shutdown controls that UPS exposes.","attack_vector":"Remote, over the network, to the card's web server on port 80/443 — no valid credentials required to bypass the authorization check.","remediation":"Firmware flash of the Network Management Card required; Schneider's SEVD-2018-074-01 advisory has the fixed build. Roll out per card — each flash briefly drops remote monitoring/management of that UPS (the UPS itself keeps powering its load through the flash).","references":["https://www.schneider-electric.com/en/download/document/SEVD-2018-074-01/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-7246","cve":"CVE-2018-7246","aliases":["SEVD-2018-074-01"],"title":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS): On default settings without SSL enabled, repeatedly requesting the card's Access Control page leaks the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric MGE Network Management Card Transverse (MGE UPS / MGE STS)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"On default settings without SSL enabled, repeatedly requesting the card's Access Control page leaks the administrative account credentials in plaintext to anyone who can sniff the traffic — handing over full UPS management access.","attack_vector":"Requires network position to observe traffic to/from the card's web server (or direct access to the unencrypted HTTP endpoint) while an admin session touches the Access Control page.","remediation":"Firmware flash to the fixed build, and as an immediate compensating step, force SSL/TLS on for the card's web interface rather than leaving it on plaintext HTTP. Same per-card rollout as the authorization-bypass companion CVE — do both in the same maintenance pass.","references":["https://www.schneider-electric.com/en/download/document/SEVD-2018-074-01/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-7820","cve":"CVE-2018-7820","aliases":[],"title":"APC UPS Network Management Card 2 (AOS 6.5.6): When Remote Monitoring is turned on and then off again, the credentials used for remote monitoring stay…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC UPS Network Management Card 2 (AOS 6.5.6)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"When Remote Monitoring is turned on and then off again, the credentials used for remote monitoring stay viewable in plaintext on the card. Anyone who gets a look at the card's config (via the web UI, a config export, or a support dump) picks up a working credential to the UPS's remote-monitoring channel.","attack_vector":"Requires some access to the card's configuration or web interface (e.g. a lower-privileged account, an exported config file, or a support bundle) — not a fully unauthenticated remote exploit, but a credential-exposure path.","remediation":"Credential rotation for the affected remote-monitoring account is the immediate fix; pair it with the AOS firmware update from APC that stops persisting the credential in plaintext once monitoring is disabled. Rotate credentials across the whole NMC2 fleet, not just the units you know were exposed.","references":["https://www.apc.com/salestools/CCON-BFQMXC/CCON-BFQMXC_R0_EN.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-9091","cve":"CVE-2018-9091","aliases":[],"title":"Kemp LoadMaster (LMOS): A flaw in session management lets a remote, unauthenticated attacker bypass the LoadMaster's security…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Kemp LoadMaster (LMOS)","year":"2018","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A flaw in session management lets a remote, unauthenticated attacker bypass the LoadMaster's security protections entirely and run elevated shell commands (ls, ps, cat, etc.) — enough to pull certificates, private keys, and other sensitive data off the load balancer.","attack_vector":"Fully remote and unauthenticated against the LoadMaster's management interface.","remediation":"Software upgrade to LMOS 7.1.35.5 (LTS) or 7.2.41.2+ (mainline) per Kemp's mitigation article. Upgrade and reboot; if this LoadMaster fronts live inference traffic, fail over to a standby instance during the update. Also rotate any certificates/keys that were on the device, since the bug allowed reading them.","references":["https://support.kemptechnologies.com/hc/en-us/articles/360001982452-Mitigation-for-Remote-Access-Execution-Vulnerability"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-0153","cve":"CVE-2019-0153","aliases":[],"title":"Intel CSME 12.0.0-12.0.34: A buffer overflow in a CSME subsystem reachable over the network by an unauthenticated attacker, giving…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME 12.0.0-12.0.34","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A buffer overflow in a CSME subsystem reachable over the network by an unauthenticated attacker, giving privilege escalation on the management engine. CSME sits below the OS with its own network stack, so a network-reachable overflow there is control of the platform outside anything the host can observe or defend.","attack_vector":"Unauthenticated network access to the affected CSME service. Whether that is reachable depends entirely on how isolated your management network is.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines. Until firmware lands, isolate the management network - this class of bug is unreachable if the management plane is not routable from anywhere a tenant can be.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0153","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00213.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-1010275","cve":"CVE-2019-1010275","aliases":[],"title":"Helm: Improper certificate validation allows unauthorized clients to connect to Tiller","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Improper certificate validation allows unauthorized clients to connect to Tiller","attack_vector":"Unauthenticated network","remediation":"Migrate off Helm 2; there is no Tiller in Helm 3","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1010275"],"status":"curated"},{"id":"CVE-2019-11171","cve":"CVE-2019-11171","aliases":["INTEL-SA-00313","CVE-2019-11168","CVE-2019-11170","CVE-2019-11178"],"title":"Intel Baseboard Management Controller firmware (Intel server boards and systems) - web/network services: Heap corruption in the BMC's network-facing code, unauthenticated, yielding information…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Baseboard Management Controller firmware (Intel server boards and systems) - web/network services","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Heap corruption in the BMC's network-facing code, unauthenticated, yielding information disclosure, privilege escalation and denial of service; the same advisory bundles an outright authentication bypass and a session-validation failure. A compromised BMC is permanent control of the node underneath the OS - virtual media, power, KVM, and the write path to BIOS and ME firmware. It survives tenant reimage by construction, and a fleet-wide compromise of BMCs is a fleet-wide power-off button, which is a hall-level physical event, not a per-node one.","attack_vector":"Unauthenticated network access to the BMC's services. Any host on the OOB management network, anything that reaches a BMC exposed through a misconfigured route or a flat provisioning VLAN, and - where the KCS host interface is enabled - a tenant with root on the node.","remediation":"BMC firmware update from Intel / the board ODM (Quanta, Wiwynn, Supermicro on Intel designs). No host reboot needed, so it can be rolled without draining training jobs, but it must be staged per board family and verified per node. Structurally: dedicated OOB network with no tenant path, jump-host-only access, unique credentials per node, disable KCS on bare-metal SKUs, and re-flash the BMC as part of node reclaim between tenants rather than trusting its current image.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11171","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00313.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-13553","cve":"CVE-2019-13553","aliases":["CVE-2019-13549","ICSA-19-297-01"],"title":"Rittal SK 3232-series chiller web interface (built on Carel pCOWeb firmware A1.5.3-B1.2.4): Whoever can reach the chiller's web card owns the chilled-water loop. The hard-coded credentials give…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal SK 3232-series chiller web interface (built on Carel pCOWeb firmware A1.5.3-B1.2.4)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Whoever can reach the chiller's web card owns the chilled-water loop. The hard-coded credentials give direct control over the two operations that matter physically: switching the cooling unit off, and moving the water temperature setpoint. That is a thermal attack, not an IT one. A GB200 or H100 rack drawing 40-140 kW has essentially zero thermal mass at the die; pull chilled water away from the CDUs or rear-door heat exchangers feeding it and inlet temperature crosses the GPU thermal-shutdown threshold in single-digit minutes. Every job in the affected hall dies, unsaved training checkpoints are lost, and repeated thermal cycling degrades VRMs, HBM stacks and pump seals. The subtler and nastier variant is not shutdown but a slow setpoint drift: raise supply water two degrees and the fleet silently throttles, which shows up as unexplained tokens-per-second regression and blown SLA credits long before anyone looks at the chiller.","attack_vector":"Unauthenticated HTTP on whatever network the chiller's pCOWeb card is plugged into. In practice that is the facility/mechanical VLAN, which in a leased colo is the landlord's network, not the tenant's - so a GPU operator may have no visibility into it at all and no idea whether it is flat with the building's office LAN. In owned or built-to-suit sites this card is usually on the same mechanical VLAN as the CRAHs, the BMS front end and the vendor's remote-support jump box. Internet exposure is real but not the common case; the common case is that anyone who lands on any building-systems subnet can reach it, and the credentials are published.","remediation":"There is no meaningful patch path - this is Carel pCOWeb OEM firmware embedded in a Rittal chiller, and the credentials are hard-coded. Realistic fix is network isolation: the chiller card goes on its own VLAN with an allow-list to exactly the BMS front end and nothing else, plus egress deny. If you lease space, you cannot patch this yourself; it is the landlord's mechanical plant. Put it in the contract - demand a network diagram for the mechanical VLAN, an attestation that no chiller/CRAH controller is reachable from any tenant or corporate network, and the right to have a third party validate it. Also demand that thermal-shutdown behaviour be tested: you need to know how many minutes you actually have, not a vendor's brochure number.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13553","https://www.cisa.gov/news-events/ics-advisories/icsa-19-297-01","https://seclists.org/fulldisclosure/2019/Oct/45"],"status":"curated"},{"id":"CVE-2019-14271","cve":"CVE-2019-14271","aliases":[],"title":"Docker / moby: Code injection into `docker cp` via nsswitch loading a library from the container chroot","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Code injection into `docker cp` via nsswitch loading a library from the container chroot; host root","attack_vector":"Malicious image","remediation":"Upgrade Docker Engine; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14271"],"status":"curated"},{"id":"CVE-2019-18658","cve":"CVE-2019-18658","aliases":[],"title":"Helm: Malicious chart includes sensitive host content such as /etc/passwd, or triggers DoS, when loaded as a…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Malicious chart includes sensitive host content such as /etc/passwd, or triggers DoS, when loaded as a directory","attack_vector":"Malicious chart from a tenant or third-party repo","remediation":"Upgrade Helm; never run `helm package`/`helm lint` on untrusted charts on a privileged host","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18658"],"status":"curated"},{"id":"CVE-2019-18801","cve":"CVE-2019-18801","aliases":[],"title":"Envoy: HTTP/2 request writes to the heap outside request buffers when the upstream is HTTP/1","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"HTTP/2 request writes to the heap outside request buffers when the upstream is HTTP/1; potential RCE in the proxy","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; for a mesh this means restarting every sidecar, which restarts tenant pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18801"],"status":"curated"},{"id":"CVE-2019-18802","cve":"CVE-2019-18802","aliases":[],"title":"Envoy: Header whitespace handling enables request smuggling and authorization bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Header whitespace handling enables request smuggling and authorization bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; sidecar restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18802"],"status":"curated"},{"id":"CVE-2019-18960","cve":"CVE-2019-18960","aliases":[],"title":"Firecracker: vsock buffer overflow producing potentially exploitable crashes","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"vsock buffer overflow producing potentially exploitable crashes","attack_vector":"Any tenant guest VM","remediation":"Upgrade Firecracker; restart microVMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18960"],"status":"curated"},{"id":"CVE-2019-20427","cve":"CVE-2019-20427","aliases":[],"title":"Lustre ptlrpc module (server-side client packet validation): TENANT ISOLATION: a Lustre client can send a crafted RPC that overflows a buffer in the server's ptlrpc…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre ptlrpc module (server-side client packet validation)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: a Lustre client can send a crafted RPC that overflows a buffer in the server's ptlrpc module, panicking it and possibly achieving remote code execution on the storage server. On a Lustre cluster the client is the tenant's GPU node — Lustre's security model assumes clients are trusted, and this is what that assumption costs. Code execution on an MDS or OSS is access to every tenant's data on the filesystem, and a panic takes the shared filesystem down for every running job.","attack_vector":"Any Lustre client — i.e. any compute node with the filesystem mounted, which in a rented GPU cluster means the tenant's own machine. No privilege escalation needed on the client beyond the ability to send RPCs.","remediation":"Upgrade Lustre servers to 2.12.3 or later. This is a coordinated storage-cluster upgrade: MDS and OSS nodes need the new build and a restart, and while Lustre supports failover pairs, most sites take an I/O pause. Structurally, treat the Lustre network (LNet) as a boundary: put it on a dedicated fabric that tenant workloads cannot address arbitrarily, and use Lustre nodemap/Shared-Secret Key authentication rather than relying on client trust.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-20427"],"status":"curated"},{"id":"CVE-2019-3705","cve":"CVE-2019-3705","aliases":[],"title":"Dell iDRAC7/8: Stack buffer overflow in the iDRAC web server — unauthenticated RCE on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7/8","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow in the iDRAC web server — unauthenticated RCE on the BMC","attack_vector":"Network, unauthenticated","remediation":"iDRAC7/8 are EOL on many fleets; remediation may require a chassis refresh rather than a patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3705"],"status":"curated"},{"id":"CVE-2019-3706","cve":"CVE-2019-3706","aliases":[],"title":"Dell iDRAC9: Authentication bypass in the iDRAC9 web interface — full out-of-band control of the server","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass in the iDRAC9 web interface — full out-of-band control of the server","attack_vector":"Network, unauthenticated","remediation":"iDRAC firmware update via Lifecycle Controller or racadm; can be scripted fleet-wide but requires a reboot window on older iDRAC lines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3706"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-3707","cve":"CVE-2019-3707","aliases":[],"title":"Dell iDRAC9: Authentication bypass via the WS-MAN interface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass via the WS-MAN interface","attack_vector":"Network, unauthenticated","remediation":"Same iDRAC firmware update; also disable WS-MAN if unused","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3707"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-5685","cve":"CVE-2019-5685","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime): A crafted shader overruns a shader-local temporary array, giving code execution in the graphics driver. Same…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX driver (shader compiler/runtime)","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A crafted shader overruns a shader-local temporary array, giving code execution in the graphics driver. Same delivery surface as the texture-array bug: wherever attacker-authored shaders reach your GPUs, this is remote code execution against the host driver.","attack_vector":"Anyone able to submit shaders - remote/VDI session users, tenant VMs, or local processes.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://www.talosintelligence.com/vulnerability_reports/TALOS-2019-0812","https://nvd.nist.gov/vuln/detail/CVE-2019-5685"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-6260","cve":"CVE-2019-6260","aliases":["Pantsdown"],"title":"ASPEED AST2400 / AST2500 BMC SoC: Arbitrary read/write of the BMC's entire physical address space **from the host CPU** — host-to-BMC boundary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED AST2400 / AST2500 BMC SoC","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary read/write of the BMC's entire physical address space **from the host CPU** — host-to-BMC boundary collapse. A tenant with host root can implant the BMC; survives reimaging and node reallocation","attack_vector":"Local, from the host OS via iLPC2AHB / PCIe VGA / X-DMA bridges","remediation":"Fix is a BMC firmware build that disables the AHB bridges (OpenBMC has it; many ODM builds do not). On multi-tenant bare metal this is the single most important control — otherwise every tenant handoff is a potential persistent implant","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-6260"],"status":"curated","fleet":{"ubiquity":"universal - ASPEED is effectively the sole-source BMC SoC in x86 server boards, including GPU servers","remediation_pain":"firmware-flash per node; on some boards the only mitigation is disabling the LPC/PCIe P2A bridges in an OEM image respin, and several SKUs remain unpatchable-mitigate-only","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"Host-side root can read/write the BMC's entire physical address space over LPC/PCIe, so any tenant that gets host root pivots into the always-on management processor - below the hypervisor, persistent across reimaging, on every node of an ASPEED-based fleet."}},{"id":"CVE-2019-7276","cve":"CVE-2019-7276","aliases":["CVE-2019-7279","CVE-2019-7274","CVE-2019-7273"],"title":"Optergy Proton / Enterprise building management platform: A backdoor console giving remote root code execution, alongside hard-coded credentials, CSRF, and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Optergy Proton / Enterprise building management platform","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A backdoor console giving remote root code execution, alongside hard-coded credentials, CSRF, and authenticated file upload that also runs as root. Optergy is a full BMS platform - it aggregates HVAC, metering and often access control for a site - so root on it is root over the facility's control surface and its credential store. An attacker gets to command cooling equipment directly, alter schedules and setpoints, silence alarms, and read out whatever downstream device passwords the platform holds. For a GPU operator the physical consequence is the familiar one: cooling commanded away from a hall of 40 kW+ racks means thermal shutdown in minutes, lost checkpoints and thermal-cycling damage. The backdoor is the part that should decide the response - a deliberate hidden access path means you cannot reason about who has been in the system historically.","attack_vector":"Remote, unauthenticated, over the platform's web interface on the facility network. Optergy deployments are frequently published for remote access because the product is sold on browser-based management, so internet exposure is a realistic assumption rather than an edge case. Hard-coded credentials mean even a 'secured' instance is open to anyone who read the advisory.","remediation":"Vendor firmware/software update - a platform upgrade rather than controller flashing, so no cooling downtime, but it requires the integrator and a version that removes the backdoor console. Given a deliberate backdoor was present, patching alone is not sufficient: rebuild or re-image the platform, rotate every credential it ever stored, and rotate credentials on every downstream device it integrated. Then remove all internet exposure and put it behind a jump host with MFA. If the platform is under an integrator's remote-support contract, audit that path specifically.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-7276","https://nvd.nist.gov/vuln/detail/CVE-2019-7279","https://applied-risk.com/resources/advisories"],"status":"curated"},{"id":"CVE-2019-9569","cve":"CVE-2019-9569","aliases":["McAfee ATR HVACking"],"title":"Delta Controls enteliBUS Manager (eBMGR) V3.40_B-571848, dactetra service: Unauthenticated remote code execution on a central plant controller. The enteliBUS Manager is the box that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Delta Controls enteliBUS Manager (eBMGR) V3.40_B-571848, dactetra service","year":"2019","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution on a central plant controller. The enteliBUS Manager is the box that sequences chillers, pumps, boilers and air handling for a whole building; the research that produced this bug demonstrated full control over HVAC and, on the same hardware family, over the access-control and life-safety points wired into it. Owning it gives an attacker root-equivalent control over the mechanical plant with no credentials and no user interaction, plus the ability to lie to the supervisory system above it. Against a GPU hall this is the textbook scenario: command the plant off or drive setpoints, and a 40 kW+ rack crosses thermal shutdown in minutes while the BMS graphic still shows normal operation. Because the compromise is code execution rather than a config change, it also persists across reboots and survives the operator's instinctive 'restart the controller' response.","attack_vector":"Unauthenticated network access to the controller's service port on the facility network - no credentials, no interaction. enteliBUS controllers are commonly reachable from anywhere on the building VLAN and, in sites where the controls contractor set up remote support, from the internet through a poorly-placed remote-access appliance.","remediation":"Firmware update from Delta Controls, applied through the certified Delta dealer who owns the site - operators generally cannot obtain or apply Delta firmware directly, which is itself the problem. That means a scheduled contractor visit and a plant controller offline during the flash: real cooling risk, real cost, and a window most operators will not take without an incident to justify it. Treat segmentation as the primary control: dedicated VLAN, deny-by-default with an allow-list from the supervisor only, and no path from the internet. Verify by scanning your own facility VLAN for the controller's service port rather than trusting the dealer's assurance.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9569","https://www.mcafee.com/blogs/other-blogs/mcafee-labs/hvacking-understanding-the-delta-between-security-and-reality/","https://www.deltacontrols.com/products/hvac-controls/central-plant-controllers/entelibus"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11483","cve":"CVE-2020-11483","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): Hard-coded credentials in the DGX BMC firmware. Anyone who can reach the BMC's management interface…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Hard-coded credentials in the DGX BMC firmware. Anyone who can reach the BMC's management interface authenticates with credentials baked into a firmware image that is publicly downloadable - meaning no credential rotation on your side ever helped. That is full out-of-band control of a DGX-1 or DGX-2: power, console, virtual media, and a path to reflashing the host. Affects DGX-1 before BMC 3.38.30 and DGX-2 before 1.06.06.","attack_vector":"Anyone with network reach to the BMC. If the management network is flat, shared with tenants, or accidentally routable, that is effectively anyone inside the datacenter - and internet-exposed BMCs make it anyone at all.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11483"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11486","cve":"CVE-2020-11486","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): File upload into the BMC that gets automatically processed, yielding remote code execution on the baseboard…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"File upload into the BMC that gets automatically processed, yielding remote code execution on the baseboard management controller of a DGX-1. Code on the BMC is below the host OS: it survives host reinstall, sees the host's memory and storage paths, controls power and firmware, and is invisible to everything running on the node. This is the worst outcome in the DGX-1 BMC set. DGX-1 before BMC 3.38.30.","attack_vector":"Anyone with network reach to the BMC's management interface. Combined with the hard-coded credentials in the same bulletin, that is unauthenticated in practice.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11486"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11951","cve":"CVE-2020-11951","aliases":[],"title":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10): PHYSICAL. A backdoor root account in the PDU firmware. Not a weak default that an operator could change - an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"PHYSICAL. A backdoor root account in the PDU firmware. Not a weak default that an operator could change - an undocumented account shipped in the image. Anyone who knows it owns the device that switches power to the rack, and nothing in your provisioning process would ever have noticed it.","attack_vector":"Network access to the PDU management interface, with credentials that are public knowledge once the advisory is out.","remediation":"Firmware update to a build that removes the account. This is a per-PDU flash, two per rack, and the update must be verified rather than assumed - check that the account is actually gone on a sample. Treat undocumented accounts as a procurement question going forward: require vendors to attest that no non-configurable accounts exist before you buy a PDU SKU at fleet scale.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11951"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11956","cve":"CVE-2020-11956","aliases":[],"title":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10): Least-privilege violation: low-privilege users on the PDU get far more capability than the role implies, up…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal PDU-3C002DEC rack PDU firmware (through 5.17.10)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Least-privilege violation: low-privilege users on the PDU get far more capability than the role implies, up to and including control of the device. Read-only monitoring accounts handed to a DCIM system or an NOC contractor become power control.","attack_vector":"Any authenticated user on the PDU, including monitoring accounts.","remediation":"Firmware update per PDU. Audit third-party accounts on the power estate at the same time - the monitoring integration is usually how the low-privilege credential got into someone else's hands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11956"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-12812","cve":"CVE-2020-12812","aliases":["FG-IR-19-283"],"title":"Fortinet FortiOS SSL-VPN: A logic flaw lets a user who changes their login case (e.g. 'User' vs 'user') complete SSL-VPN authentication…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Fortinet FortiOS SSL-VPN","year":"2020","cvss_score":9.8,"severity":"critical","kev":true,"impact":"A logic flaw lets a user who changes their login case (e.g. 'User' vs 'user') complete SSL-VPN authentication without ever being prompted for their second factor — a clean bypass of FortiToken MFA on the VPN gateway. Confirmed in CISA's KEV catalog as actively exploited; if this FortiGate is the VPN entry point into a cluster's management network, MFA was supposed to be the thing stopping a stolen password from being enough.","attack_vector":"Requires a valid username/password (e.g. phished or reused) but no second factor — the attacker just varies the case of the username at login to skip the FortiToken prompt.","remediation":"Firmware upgrade to the fixed FortiOS release per Fortinet PSIRT FG-IR-19-283. Given confirmed active exploitation, patch ahead of routine cycles, and afterward force a credential rotation for any accounts that authenticated to SSL-VPN during the vulnerable window in case MFA was bypassed on them already.","references":["https://fortiguard.com/psirt/FG-IR-19-283","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2020-12812"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-13092","cve":"CVE-2020-13092","aliases":[],"title":"scikit-learn / joblib: `joblib.load()` executes commands from an untrusted file via `__reduce__`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"scikit-learn / joblib","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`joblib.load()` executes commands from an untrusted file via `__reduce__`","attack_vector":"Customer-supplied `.joblib`/`.pkl` model","remediation":"No fix — this is pickle semantics. Reject joblib artifacts from untrusted sources; use skops or ONNX","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13092"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2020-15373","cve":"CVE-2020-15373","aliases":[],"title":"Brocade Fabric OS REST API: Multiple buffer overflows in the Fabric OS REST API reachable by an unauthenticated remote attacker. The REST…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS REST API","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Multiple buffer overflows in the Fabric OS REST API reachable by an unauthenticated remote attacker. The REST API is what modern SAN automation drives, so it is enabled in exactly the environments that also automate zoning — meaning the vulnerable interface is the one wired into your provisioning pipeline.","attack_vector":"Unauthenticated, remote to the FOS REST API on v8.2.1 through v8.2.1d, and 8.2.2 before v8.2.2c.","remediation":"Fabric OS upgrade plus reboot, per fabric. If you do not use the REST API, disabling it is a live config change that removes the exposure without a maintenance window. Related: CVE-2020-15374, CVE-2020-15371.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15373","https://nvd.nist.gov/vuln/detail/CVE-2020-15371"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-15639","cve":"CVE-2020-15639","aliases":["CVE-2020-15640","CVE-2020-15641","CVE-2020-15642","CVE-2020-15643","CVE-2020-15644","CVE-2020-15645","CVE-2020-17387","CVE-2020-17388","CVE-2020-17389","CVE-2020-5803","CVE-2020-5804","CVE-2020-5805"],"title":"Marvell QConvergeConsole GUI 5.5.0.64 - 5.5.0.74 (QLogic HBA management): The earlier cluster on the same console: unauthenticated RCE via decryptFile, unauthenticated file disclosure…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Marvell QConvergeConsole GUI 5.5.0.64 - 5.5.0.74 (QLogic HBA management)","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The earlier cluster on the same console: unauthenticated RCE via decryptFile, unauthenticated file disclosure via getFileUploadBytes, several RCE paths whose authentication requirement is defeated by a bypass in the console's own auth mechanism, path traversal that deletes arbitrary files as SYSTEM or root, and Tomcat credentials stored in cleartext in tomcat-users.xml. Included here because operators routinely find 5.5.0.6x/7x consoles still running on legacy management servers that were never inventoried - the same fleet-wide HBA firmware control as the 2025 cluster, on hosts nobody is patching.","attack_vector":"Any host reachable to the QConvergeConsole web port. Unauthenticated for the RCE and disclosure primitives; the cleartext tomcat-users.xml additionally gives any local OS user on the console host a working login.","remediation":"Do not patch this generation - retire it. Inventory for QConvergeConsole installs by port and by package, uninstall from all hosts, and replace with CLI-driven HBA management. If a console must stay, upgrade to the current release, rotate the Tomcat credentials and put the port behind an admin-only ACL. No HBA firmware flash and no storage downtime is required to remove the console.","references":["https://www.marvell.com/content/dam/marvell/en/public-collateral/fibre-channel/marvell-fibre-channel-security-advisory-2020-07.pdf","https://www.zerodayinitiative.com/advisories/ZDI-20-967/","https://nvd.nist.gov/vuln/detail/CVE-2020-15639"],"status":"curated"},{"id":"CVE-2020-27745","cve":"CVE-2020-27745","aliases":[],"title":"Slurm: RPC buffer overflow in the PMIx MPI plugin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RPC buffer overflow in the PMIx MPI plugin","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm; restart slurmctld and slurmd. Draining a Slurm partition means killing running training jobs","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27745"],"status":"curated"},{"id":"CVE-2020-3992","cve":"CVE-2020-3992","aliases":[],"title":"VMware ESXi (OpenSLP): Use-after-free in OpenSLP on port 427 - unauthenticated remote code execution on the hypervisor","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi (OpenSLP)","year":"2020","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Use-after-free in OpenSLP on port 427 - unauthenticated remote code execution on the hypervisor [KEV]","attack_vector":"Unauthenticated network on the management segment","remediation":"Patch and disable the SLP service entirely (VMware's own recommendation). Any ESXi with 427 reachable should be treated as already compromised","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-3992"],"status":"curated"},{"id":"CVE-2020-5344","cve":"CVE-2020-5344","aliases":[],"title":"Dell iDRAC9: Stack-based buffer overflow via crafted remote input — pre-auth code execution on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack-based buffer overflow via crafted remote input — pre-auth code execution on the BMC","attack_vector":"Network, unauthenticated","remediation":"iDRAC firmware update; requires a rolling out-of-band update campaign across the fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5344"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-7521","cve":"CVE-2020-7521","aliases":["SEVD-2020-224-01"],"title":"APC Easy UPS On-Line Software (SFAPV9601) FileUploadServlet: Path traversal in a file upload servlet allows writing an executable anywhere on the host, giving code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS On-Line Software (SFAPV9601) FileUploadServlet","year":"2020","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Path traversal in a file upload servlet allows writing an executable anywhere on the host, giving code execution on the machine that manages UPS shutdown across the site. Same PHYSICAL end state as the newer Easy UPS bugs - an attacker gains the ability to command an orderly power-down of everything the software manages.","attack_vector":"Unauthenticated network access to the software's web servlet.","remediation":"Upgrade past v2.0. Server-side upgrade, cheap. If the host was exposed, rebuild rather than patch - and rotate every UPS and host credential it stored.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7521"],"status":"curated"},{"id":"CVE-2021-1104","cve":"CVE-2021-1104","aliases":[],"title":"RISC-V ISA (MTVEC register) as used in NVIDIA GPU microcontrollers: A documented ambiguity in the RISC-V specification leaves the machine trap vector base address register in an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"RISC-V ISA (MTVEC register) as used in NVIDIA GPU microcontrollers","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A documented ambiguity in the RISC-V specification leaves the machine trap vector base address register in an undefined state at reset, which fault injection can exploit to redirect trap handling and break the secure boot chain of an embedded microcontroller. NVIDIA's own offensive-security team assigned this; it is relevant because modern NVIDIA GPU and platform microcontrollers are RISC-V based, so it describes a weakness in the root of trust below the GPU driver. Exploitation needs physical glitching, not remote access - so it belongs in your supply-chain and physical-security threat model, not your patch queue.","attack_vector":"An attacker with physical access to the board who can perform voltage or clock glitching. Not reachable from software, local or remote.","remediation":"Architectural ambiguity in the RISC-V ISA specification as implemented in embedded microcontrollers, including NVIDIA's. There is no operator-installable patch: the fix is in silicon and in hardened boot firmware from the chip vendor. For an operator this is unpatchable in the field - mitigate by treating physical access to a GPU as game over, keeping fault-injection-capable access (open chassis, exposed board) out of shared-tenancy racks, and preferring hardware generations that NVIDIA states carry the hardened boot ROM.","references":["https://riscv.org/news/2021/08/video-glitching-risc-v-chips-mtvec-corruption-for-hardening-isa-adam-zabrocki-and-alex-matrosov-def-con-29/","https://nvd.nist.gov/vuln/detail/CVE-2021-1104"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2021-1361","cve":"CVE-2021-1361","aliases":[],"title":"Cisco Nexus 3000/9000 (internal file management service): Unauthenticated remote file write, read and delete as root on the switch, over a service that listens on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus 3000/9000 (internal file management service)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote file write, read and delete as root on the switch, over a service that listens on the management interface by default. An attacker on the management network can drop a file, replace a config, or wipe the box without ever having a credential. This is the single worst pre-auth exposure in the Nexus line and it is a good argument for treating the switch management VLAN as production, not as 'internal'.","attack_vector":"Unauthenticated, remote — anything that can reach TCP/9075 on the switch's mgmt0 interface. No credentials required.","remediation":"NX-OS image upgrade and switch reload. As a stopgap, an interface ACL on mgmt0 restricting the affected port materially reduces exposure and can be applied live with no reload. Rollout: one reload per switch, drain-and-patch per MLAG pair.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1361"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-27797","cve":"CVE-2021-27797","aliases":[],"title":"Brocade Fabric OS (hard-coded credentials): Documented hard-coded credentials in Brocade Fabric OS — the operating system on the Fibre Channel directors…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS (hard-coded credentials)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Documented hard-coded credentials in Brocade Fabric OS — the operating system on the Fibre Channel directors and switches that carry NVMe-over-FC and FC-attached storage into AI clusters. Hard-coded credentials on a SAN switch mean anyone who can reach its management interface owns the zoning configuration, and zoning is the only thing keeping one host's initiators from seeing another's LUNs.","attack_vector":"Anyone with network reachability to the FOS management interface. The credentials are public.","remediation":"Upgrade Fabric OS to 8.2.1c / 8.1.2h or later — all 8.0.x and 7.x versions are affected with no fixed release, so those switches need replacing or hard isolation. A FOS upgrade is a firmware install plus switch reboot; on a redundant dual-fabric SAN you do one fabric at a time and multipathing covers it. Until then, put FOS management interfaces on an isolated OOB network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27797"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-28235","cve":"CVE-2021-28235","aliases":[],"title":"etcd: Authentication flaw via the debug function","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"etcd","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication flaw via the debug function -> remote privilege escalation","attack_vector":"Network (remote)","remediation":"Control-plane: CRITICAL - etcd holds every k8s secret; upgrade + rotate all stored secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28235"],"status":"curated"},{"id":"CVE-2021-28503","cve":"CVE-2021-28503","aliases":[],"title":"Arista EOS (eAPI certificate auth): Certificate-based eAPI authentication skips credential re-evaluation — authentication bypass on the switch's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (eAPI certificate auth)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Certificate-based eAPI authentication skips credential re-evaluation — authentication bypass on the switch's programmatic API, which is exactly the interface a neocloud's fabric automation uses","attack_vector":"Network","remediation":"EOS upgrade with fabric failover; also rotate any eAPI client certificates issued while vulnerable","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28503"],"status":"curated"},{"id":"CVE-2021-29212","cve":"CVE-2021-29212","aliases":["HPESBGN04189","ZDI-21-1278"],"title":"HPE iLO Amplifier Pack (unauthenticated directory traversal): Unauthenticated directory traversal on the iLO Amplifier Pack appliance, letting an attacker read arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO Amplifier Pack (unauthenticated directory traversal)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated directory traversal on the iLO Amplifier Pack appliance, letting an attacker read arbitrary files off it. Amplifier Pack is the fleet-scale iLO management appliance - it inventories every iLO it manages and holds the credentials it uses to reach them. Reading its filesystem without authentication is therefore a credential harvest for the entire out-of-band plane, not a single-host information leak. One unauthenticated request against one appliance can yield the keys to every BMC behind it. Affects versions 1.80, 1.81, 1.90 and 1.95.","attack_vector":"Anything routable to the Amplifier Pack appliance, unauthenticated. In most deployments that means the management network - so the question is which hosts, VPN pools and monitoring systems share a segment with your fleet management appliance.","remediation":"Upgrade iLO Amplifier Pack past the affected 1.80-1.95 range - a single appliance upgrade, no per-node work, no reboots and no job drain, so rollout cost is near zero. What is not near zero: if this appliance was reachable while unpatched, assume the iLO credentials it stored are burned and rotate them fleet-wide. That credential rotation is the expensive part, and skipping it leaves the actual exposure open after the appliance is patched.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbgn04189en_us","https://www.zerodayinitiative.com/advisories/ZDI-21-1278/","https://nvd.nist.gov/vuln/detail/CVE-2021-29212"],"status":"curated"},{"id":"CVE-2021-31884","cve":"CVE-2021-31884","aliases":["CVE-2021-31886","CVE-2021-31887","CVE-2021-31888","CVE-2017-9946","SSA-044112","SSA-114589"],"title":"Siemens APOGEE PXC / MEC / MBC and TALON TC BACnet and P2 automation controllers: A cluster of critical flaws in the Siemens field controllers that actually drive air handling units, VAV…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Siemens APOGEE PXC / MEC / MBC and TALON TC BACnet and P2 automation controllers","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A cluster of critical flaws in the Siemens field controllers that actually drive air handling units, VAV boxes and chilled-water valves, including unauthenticated paths to the integrated web server and memory-safety issues reachable over the network. Compromise of an APOGEE or TALON controller is compromise of the physical actuator layer: fan enable, damper position, valve position, setpoint. There is no supervisory system to overrule it because the controller is the thing that closes the loop. The older companion issue (CVE-2017-9946) lets an attacker bypass authentication on the embedded web server and pull configuration off the device, which hands over the point map needed to make the attack surgical. In a GPU hall, an attacker who owns the AHU controllers can drive inlet temperatures past the accelerator shutdown threshold in minutes and, by holding valve position, keep them there.","attack_vector":"Network access to the controller's HTTP/HTTPS ports (80/443) on the facility network. These are wall-mounted controllers in mechanical rooms and ceiling spaces, on a flat building VLAN with no port security in the overwhelming majority of sites. Physical access to the panel is also a real vector because these are often in unlocked or shared mechanical spaces that tenant security policy does not cover.","remediation":"Firmware upgrade per Siemens SSAs - APOGEE PXC Compact/Modular to V3.5.4+ (BACnet) or V2.8.19+ (P2), TALON TC accordingly. That is a per-controller flash by a Siemens-certified technician, with each AHU losing automatic control during its flash, so it is a genuine maintenance-window project measured in technician-days across a large site and one that most operators will keep deferring. Because of that, the load-bearing control is network isolation: controllers on a dedicated VLAN, HTTP/HTTPS permitted only from the supervisory station, and physical locks on mechanical rooms and control panels. If you lease, these are the landlord's controllers - ask for the firmware baseline, and if the answer is 'all versions' assume the cluster applies.","references":["https://cert-portal.siemens.com/productcert/pdf/ssa-044112.pdf","https://cert-portal.siemens.com/productcert/pdf/ssa-148078.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-31884","https://nvd.nist.gov/vuln/detail/CVE-2017-9946"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-31921","cve":"CVE-2021-31921","aliases":[],"title":"Istio: With AUTO_PASSTHROUGH gateways, an external client reaches arbitrary in-cluster services, bypassing all…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"With AUTO_PASSTHROUGH gateways, an external client reaches arbitrary in-cluster services, bypassing all authorization","attack_vector":"Unauthenticated network via the ingress gateway","remediation":"Emergency istiod and gateway upgrade; audit for AUTO_PASSTHROUGH gateways","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-31921"],"status":"curated"},{"id":"CVE-2021-32974","cve":"CVE-2021-32974","aliases":["ICSA-21-187-01"],"title":"Moxa NPort IAW5000A-I/O serial device server: The built-in web server doesn't validate input properly, letting a remote unauthenticated attacker run…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Moxa NPort IAW5000A-I/O serial device server","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The built-in web server doesn't validate input properly, letting a remote unauthenticated attacker run arbitrary commands on the device server — full takeover of the box bridging serial equipment onto the network.","attack_vector":"Remote, unauthenticated — a crafted HTTP request to the web management interface is enough.","remediation":"Firmware upgrade to the version Moxa published in its security advisory; requires a flash and reboot on every affected unit, which briefly drops the serial sessions it's carrying. No compensating config change exists since the flaw is in input handling, not a feature you can disable.","references":["https://www.cisa.gov/uscert/ics/advisories/icsa-21-187-01","https://www.moxa.com/en/support/product-support/security-advisory/nport-iaw5000a-io-serial-device-server-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-39226","cve":"CVE-2021-39226","aliases":[],"title":"Grafana: Unauthenticated access to snapshots via /api/snapshots/:key","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2021","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Unauthenticated access to snapshots via /api/snapshots/:key -> view and delete snapshot data","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the monitoring tier","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39226"],"status":"curated"},{"id":"CVE-2021-39815","cve":"CVE-2021-39815","aliases":["PowerVR pinned-memory UAF"],"title":"Imagination PowerVR GPU driver - pinned memory lifecycle: MULTI-TENANT ISOLATION: an unprivileged app allocates pinned GPU memory, unpins it so the page can be freed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Imagination PowerVR GPU driver - pinned memory lifecycle","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: an unprivileged app allocates pinned GPU memory, unpins it so the page can be freed and reallocated elsewhere, and keeps using it in GPU calls - a use-after-free that reads and writes pages now belonging to something else. Scored 9.8. This is the GPU-memory-lifecycle bug class that matters most for shared accelerators: the GPU keeps a mapping the OS thinks it revoked.","attack_vector":"Unprivileged local application with GPU access. Re-filed as CVE-2022-20122 in a later Android bulletin.","remediation":"Update the vendor GPU driver. The transferable lesson for a GPU fleet is to ask, for whichever accelerator you run, whether unpinning actually tears down the device-side mapping - the same design mistake is what makes GPU memory reuse dangerous on any vendor.","references":["https://source.android.com/security/bulletin/2022-04-01","https://nvd.nist.gov/vuln/detail/CVE-2021-39815"],"status":"curated"},{"id":"CVE-2021-41842","cve":"CVE-2021-41842","aliases":["VU#796611"],"title":"Insyde InsydeH2O (AtaLegacySmm SMM driver): The SMI handler in the legacy ATA driver does not validate the CommBuffer it is handed, so a caller from the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AtaLegacySmm SMM driver)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The SMI handler in the legacy ATA driver does not validate the CommBuffer it is handed, so a caller from the OS can steer SMM into executing attacker-supplied code. That is full ring -2 control: persistence that survives OS reinstall and disk wipe, the ability to disable or forge Secure Boot and TPM measurements, and a vantage point beneath any hypervisor. The highest-scored entry in the whole 2021-2022 Insyde/Binarly batch.","attack_vector":"Local admin or root on the host OS, then a software SMI invoking the vulnerable handler. On bare-metal GPU rental this is exactly the privilege a tenant already has on their leased node.","remediation":"InsydeH2O kernel fix (5.0 / 05.08.46, 5.1 / 05.16.46, 5.2 / 05.26.46, 5.3 / 05.35.46, 5.4 / 05.43.46, 5.5 / 05.51.45) - but you cannot apply that. You need the BIOS image your server OEM built on top of it, and the rebase lag from Insyde's kernel drop to a shipping Dell/HPE/Lenovo/Supermicro payload ran into many months for this batch. Firmware flash plus one reboot per node. No config workaround: SMM cannot be turned off. If you rent bare metal to untrusted tenants on unpatched firmware, treat every returned node as compromised and re-flash rather than reimage.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41842","https://kb.cert.org/vuls/id/796611","https://www.insyde.com/security-pledge"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-42774","cve":"CVE-2021-42774","aliases":["CVE-2021-42772","CVE-2021-42773","CVE-2021-42775"],"title":"Broadcom Emulex HBA Manager / OneCommand Manager (Fibre Channel and FC-NVMe HBAs), before 11.4.425.0 and 12.8.542.31: Unless the agent was installed in Strictly Local Management mode, the Emulex…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Broadcom Emulex HBA Manager / OneCommand Manager (Fibre Channel and FC-NVMe HBAs), before 11.4.425.0 and 12.8.542.31","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unless the agent was installed in Strictly Local Management mode, the Emulex management daemon accepts unauthenticated remote commands. Two of them matter: the remote firmware-download path has a buffer overflow reachable pre-auth, and the same path lets an attacker place or replace an arbitrary file on the host. That is remote, unauthenticated code execution on the host plus a direct route to writing HBA firmware. An Emulex HBA is a PCIe device with DMA and its own processor that owns the node's path to shared storage - firmware planted there survives an OS reimage and can silently mirror or corrupt a later tenant's FC traffic. The companion GetDumpFile issues additionally let an unauthenticated caller pull arbitrary files off the host.","attack_vector":"Any host that can reach the HBA Manager remote management listener on the host's management interface. No credentials in non-secure (default remote) mode. This listener frequently ends up on the same flat provisioning VLAN as the BMCs.","remediation":"Upgrade HBA Manager to 11.4.425.0 or 12.8.542.31 and above on every host with an Emulex HBA. The stronger and faster fix is to reinstall the agent in Strictly Local Management mode, which removes the remote listener entirely and leaves only local hbacmd - do this by default on bare-metal tenant nodes, since remote HBA management is rarely worth the exposure. Neither step requires an HBA firmware flash or an array outage; it is an agent reinstall and service restart. If a node was exposed, treat the HBA firmware as untrusted and reflash from the vendor image before returning it to the pool.","references":["https://docs.broadcom.com/doc/elx_HBAManager-Lin-RN12811-101.pdf","https://www.broadcom.com/products/storage/fibre-channel-host-bus-adapters/emulex-hba-manager","https://nvd.nist.gov/vuln/detail/CVE-2021-42774"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-43215","cve":"CVE-2021-43215","aliases":["CVE-2017-0104"],"title":"Microsoft iSNS Server service (Internet Storage Name Service for iSCSI discovery): Memory corruption in the iSNS Server service leading to remote code execution, with an earlier…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Microsoft iSNS Server service (Internet Storage Name Service for iSCSI discovery)","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Memory corruption in the iSNS Server service leading to remote code execution, with an earlier integer-overflow variant in the same service. iSNS is the discovery registry for an iSCSI fabric: it tells initiators which targets and portals exist and which discovery domains they belong to. Code execution there gives an attacker the ability to re-point initiators at targets of their choosing, which is a fabric-wide man-in-the-middle on block storage without ever touching a target or a host. Discovery domains are also the iSCSI analogue of FC zoning, so controlling iSNS is controlling who can see whose LUNs.","attack_vector":"Any host that can reach the iSNS server's registration port (TCP 3205) on the storage or management network. No authentication - iSNS has essentially none in common deployments.","remediation":"Patch the Windows host running the iSNS Server role and reboot. Better: most iSCSI deployments in a GPU datacenter do not need iSNS at all - initiators are configured with explicit target portals by the provisioning system. If that describes you, remove the iSNS Server role and strip the iSNS configuration from initiators; that removes an unauthenticated fabric-control service permanently and costs nothing operationally. If you do need it, put port 3205 behind an ACL that only initiator hosts can traverse.","references":["https://msrc.microsoft.com/update-guide/vulnerability/CVE-2021-43215","https://nvd.nist.gov/vuln/detail/CVE-2021-43215"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47215","cve":"CVE-2021-47215","aliases":["net/mlx5e kTLS crash in RX resync flow"],"title":"Linux kernel mlx5_core kTLS RX offload: TLS RX resync list corruption: entries are moved by the resync handler while still in use in NAPI, corrupting…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core kTLS RX offload","year":"2021","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TLS RX resync list corruption: entries are moved by the resync handler while still in use in NAPI, corrupting the list from the receive softirq. Remote and unauthenticated on any node using mlx5 hardware kTLS RX offload.","attack_vector":"Remote sender over a TLS connection handled by mlx5 kTLS RX offload, able to trigger resync conditions (out-of-order or retransmitted TLS records).","remediation":"Upgrade the host kernel to 5.16 or the 5.15.5 stable backport. Rolling reboot. Interim: disable kTLS RX offload (ethtool -K <dev> tls-hw-rx-offload off) as a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47215","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47215.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-1388","cve":"CVE-2022-1388","aliases":["K23605346"],"title":"F5 BIG-IP (iControl REST): An unauthenticated attacker can send undisclosed requests to the iControl REST management API and bypass…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"F5 BIG-IP (iControl REST)","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"An unauthenticated attacker can send undisclosed requests to the iControl REST management API and bypass authentication entirely, then run arbitrary system commands, create/delete files, or disable services — full compromise of the load balancer. This is one of the most widely exploited F5 bugs on record and is confirmed in CISA's KEV catalog; mass scanning for it started within days of disclosure.","attack_vector":"Remote and unauthenticated, but only if the iControl REST management interface is reachable from the attacker's network position — the standard defense-in-depth advice from F5 is that this interface should never be internet-facing, only reachable from a management network.","remediation":"Software upgrade to the fixed BIG-IP version per F5 K23605346, then reboot; if immediate patching isn't possible, restricting iControl REST access to a trusted management network (or disabling it on the self-IP/external interfaces) is the documented interim mitigation. Given confirmed mass exploitation, treat any unpatched device with an internet-reachable management plane as likely already compromised, not just vulnerable.","references":["https://support.f5.com/csp/article/K23605346","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2022-1388"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-20861","cve":"CVE-2022-20861","aliases":[],"title":"Cisco Nexus Dashboard (web UI / CSRF): One of a batch of unauthenticated flaws in Nexus Dashboard that together allow remote command execution…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus Dashboard (web UI / CSRF)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"One of a batch of unauthenticated flaws in Nexus Dashboard that together allow remote command execution, reading and uploading container images, and CSRF. Nexus Dashboard is the pane of glass over your whole data-center fabric, so an attacker who lands here can push configuration and container workloads to every managed switch. Uploading a container image is the persistence path — it survives the dashboard being patched.","attack_vector":"Unauthenticated, remote to the Nexus Dashboard web interface. CSRF variant needs an admin to visit an attacker page while logged in.","remediation":"Upgrade the Nexus Dashboard cluster software. Application upgrade, no switch reload; the managed fabric keeps forwarding throughout. Re-verify every managed device's running image afterwards, since image upload was in scope.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-20861","https://nvd.nist.gov/vuln/detail/CVE-2022-20858","https://nvd.nist.gov/vuln/detail/CVE-2022-20857"],"status":"curated"},{"id":"CVE-2022-23676","cve":"CVE-2022-23676","aliases":[],"title":"ArubaOS-Switch (HPE Aruba wired switches): Remote arbitrary code execution on ArubaOS-Switch devices, affecting every version of the 15.xx and early…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ArubaOS-Switch (HPE Aruba wired switches)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Remote arbitrary code execution on ArubaOS-Switch devices, affecting every version of the 15.xx and early 16.xx trains. Aruba wired switches show up in GPU-cluster builds as the out-of-band management and provisioning fabric rather than the GPU data path — which makes RCE here a foothold on the network that reaches every BMC in the building.","attack_vector":"Remote to the switch's services. No credentials in the affected paths.","remediation":"ArubaOS-Switch firmware upgrade plus switch reload. Management switches are usually not redundant the way the GPU fabric is, so a reload means the OOB network drops — schedule it when you do not need remote console. Restrict management-plane reachability first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23676"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-27518","cve":"CVE-2022-27518","aliases":[],"title":"Citrix ADC/Gateway: SAML SP/IdP config","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix ADC/Gateway","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] SAML SP/IdP config -> unauthenticated remote arbitrary code execution (nation-state exploited)","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27518"],"status":"curated"},{"id":"CVE-2022-28163","cve":"CVE-2022-28163","aliases":[],"title":"Brocade SANnav Management Portal - Zone management endpoints, before SANnav 2.2.0: SQL injection in multiple endpoints associated with zone management, allowing arbitrary SQL against the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade SANnav Management Portal - Zone management endpoints, before SANnav 2.2.0","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SQL injection in multiple endpoints associated with zone management, allowing arbitrary SQL against the SANnav database. SANnav is the single management plane for an entire Fibre Channel estate; its database holds the fabric inventory, the zoning configuration and the stored switch credentials. Arbitrary SQL there means reading every switch password SANnav holds and manipulating the zoning configuration it pushes - so one flaw in a management appliance converts into cross-tenant LUN exposure across every fabric it manages, without ever touching a switch directly.","attack_vector":"A user who can reach the SANnav web application. The zone-management endpoints sit behind the portal login, so realistically this is a low-privilege operator account, a stolen session, or an attacker who first used one of the SANnav authentication-bypass defects.","remediation":"Upgrade the SANnav Management Portal to 2.2.0 or later. This is an appliance/VM upgrade, not a switch firmware flash - no fabric downtime and no arrays offline, but it does take the management plane out for the duration. Afterwards, rotate every switch credential stored in SANnav, because the point of the bug is that they were readable. Keep SANnav off any network a tenant or a BMC can reach.","references":["https://www.broadcom.com/support/fibre-channel-networking/security-advisories/brocade-security-advisory-2022-1842","https://nvd.nist.gov/vuln/detail/CVE-2022-28163"],"status":"curated"},{"id":"CVE-2022-29264","cve":"CVE-2022-29264","aliases":[],"title":"coreboot 4.13-4.16 (SMM handling on application processors): Arbitrary code execution in System Management Mode on application processors - every core other than the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"coreboot 4.13-4.16 (SMM handling on application processors)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution in System Management Mode on application processors - every core other than the bootstrap processor. Scored critical. On a many-core GPU host that is a large number of entry points into ring -2, and SMM compromise means firmware persistence beneath the OS and the hypervisor, forged or suppressed boot measurements, and an implant that a reimage between tenants does not touch. Relevant to operators running open firmware stacks (OCP-style, Open System Firmware, or coreboot-based management and storage nodes) rather than vendor BIOS.","attack_vector":"Local attacker on the host able to reach SMM on a non-bootstrap core. Requires code execution on the node, not remote access.","remediation":"Rebuild and reflash coreboot at 4.17 or later - which for coreboot-based fleets is your own build pipeline rather than an OEM download, so the rebase lag is yours to control and can be much shorter than the IBV-to-OEM path. Firmware flash, one reboot per node. No config workaround. Note that coreboot does not publish a CVE-indexed advisory page and has no GitHub security advisories, so tracking its security fixes means watching commits and release notes directly rather than waiting for an advisory feed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29264","https://github.com/advisories/GHSA-3295-v7jw-7p5g"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-29502","cve":"CVE-2022-29502","aliases":[],"title":"Slurm: Incorrect access control leading to privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Incorrect access control leading to privilege escalation","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm; restart the controller and all node daemons","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29502"],"status":"curated"},{"id":"CVE-2022-32295","cve":"CVE-2022-32295","aliases":["AMP-SB-0002","Altra SPI-NOR SMC protection"],"title":"Ampere Altra and Altra Max UEFI reference design before SRP 1.09 - SMC interface exposing SPI-NOR flash: The OS or hypervisor can reach the SPI-NOR boot flash through an insufficiently protected…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra and Altra Max UEFI reference design before SRP 1.09 - SMC interface exposing SPI-NOR flash","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The OS or hypervisor can reach the SPI-NOR boot flash through an insufficiently protected SMC. That means whoever owns the kernel on an Altra box can rewrite platform firmware - a persistent, below-the-OS implant that survives reimaging, disk wipe and tenant handoff. For a bare-metal Arm GPU provider this is the canonical tenant-persistence failure: rent a node for an hour, own it for its service life. It also destroys any attestation story you have told customers.","attack_vector":"Any code at host kernel or hypervisor level on an Altra / Altra Max system - which, on bare-metal rental, means the tenant by design. No physical access needed.","remediation":"Update to Altra SRP 1.09 or later from the board OEM (the fix hardens the SMC so the non-secure world can no longer drive SPI-NOR). Flash + reboot + drain per node, and the OEM has to ship an SRP build for your specific board - Ampere publishes the reference, your ODM integrates it, so the lag is on them. Independently and more importantly: on bare-metal, verify boot flash contents against a golden image at every tenant handoff. Assume any node rented before the SRP update may already carry an implant and reflash it from an out-of-band path rather than trusting in-band verification.","references":["https://amperecomputing.com/products/security-bulletins/altra-spi-nor-smc.html","https://nvd.nist.gov/vuln/detail/CVE-2022-32295","https://amperecomputing.com/products/product-security"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33186","cve":"CVE-2022-33186","aliases":[],"title":"Brocade Fabric OS (unauthenticated remote code execution): Unauthenticated remote code execution on a Fibre Channel switch running Fabric OS. Code execution on a SAN…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS (unauthenticated remote code execution)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution on a Fibre Channel switch running Fabric OS. Code execution on a SAN switch is total control of the storage fabric: zoning, LUN masking enforcement, and the path every host takes to its data. Affects v9.1.1, v9.0.1e, v8.2.3c, v7.4.2j and earlier.","attack_vector":"Unauthenticated, remote to the switch's management services.","remediation":"Fabric OS upgrade plus switch reboot, one fabric at a time so multipathing keeps hosts online. Restrict FOS management reachability to a dedicated OOB network as an immediate config control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33186"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-40684","cve":"CVE-2022-40684","aliases":[],"title":"Fortinet FortiOS/FortiProxy: Auth bypass via an alternate path","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS/FortiProxy","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Auth bypass via an alternate path -> unauthenticated operations on the admin interface","attack_vector":"Network (remote)","remediation":"Control-plane: firmware + audit for attacker-added admin accounts and SSH keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40684"],"status":"curated"},{"id":"CVE-2022-42475","cve":"CVE-2022-42475","aliases":[],"title":"Fortinet FortiOS: SSL-VPN heap-based buffer overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS","year":"2022","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] SSL-VPN heap-based buffer overflow -> unauthenticated RCE, exploited as a zero-day","attack_vector":"Network (remote)","remediation":"Control-plane: firmware upgrade; full IOC sweep","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42475"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42970","cve":"CVE-2022-42970","aliases":["SEVD-2022-256-01"],"title":"APC Easy UPS Online Monitoring Software (Windows and Windows Server): PHYSICAL. Critical functions in the UPS monitoring server are exposed with no authentication at all. This…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS Online Monitoring Software (Windows and Windows Server)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"PHYSICAL. Critical functions in the UPS monitoring server are exposed with no authentication at all. This software is what issues graceful-shutdown commands to hosts and controls UPS behaviour - so an attacker who reaches it can trigger a coordinated shutdown of everything the software manages. On a GPU fleet that is a clean, deniable way to kill every running job at once, no memory corruption required.","attack_vector":"Unauthenticated, over the network to the monitoring server. This software typically runs on a Windows box on the facility or management network with wide reachability, because it needs to talk to every UPS and every managed host.","remediation":"Upgrade the monitoring software (SEVD-2022-256-01). This is a server-side upgrade rather than device firmware, so it is cheap - a reboot of one Windows host, not a maintenance window on the power train. The harder work is the network position: this box should not be reachable from the compute network or from tenant-facing subnets, and its shutdown-command channel should be authenticated.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42970"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-42971","cve":"CVE-2022-42971","aliases":["SEVD-2022-256-01"],"title":"APC Easy UPS Online Monitoring Software (Windows and Windows Server): Unrestricted file upload leads to remote code execution by dropping a JSP payload. Full control of the host…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS Online Monitoring Software (Windows and Windows Server)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unrestricted file upload leads to remote code execution by dropping a JSP payload. Full control of the host that orchestrates UPS shutdowns across the site - which is the same PHYSICAL outcome as above, plus a persistent Windows foothold sitting on the facility network.","attack_vector":"Unauthenticated network access to the monitoring server's web component.","remediation":"Software upgrade per SEVD-2022-256-01, then treat the host as potentially compromised and rebuild it rather than patching in place if it was ever internet-reachable. Verify what the box could reach - it usually holds credentials for every UPS and every managed host it can shut down.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42971"],"status":"curated"},{"id":"CVE-2022-45907","cve":"CVE-2022-45907","aliases":[],"title":"PyTorch (`torch.jit.annotations.parse_type_line`): Arbitrary code execution via unsafe `eval` in TorchScript type parsing","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.jit.annotations.parse_type_line`)","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution via unsafe `eval` in TorchScript type parsing","attack_vector":"Customer-supplied TorchScript model file","remediation":"Patch torch in base images; TorchScript ingest of untrusted models should be sandboxed regardless","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-45907"],"status":"curated"},{"id":"CVE-2022-46892","cve":"CVE-2022-46892","aliases":["AMP-SB-0006","root complex OS re-enable"],"title":"Ampere Altra and Altra Max before firmware 2.10c - PCIe root complex access control: The OS can re-initialise a PCIe root complex the platform firmware deliberately disabled. Operators disable…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra and Altra Max before firmware 2.10c - PCIe root complex access control","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The OS can re-initialise a PCIe root complex the platform firmware deliberately disabled. Operators disable root complexes to fence off slots - a management NIC, a device belonging to another tenant partition, a port that should not exist in this SKU. Re-enabling it hands the tenant a live PCIe path to hardware they were never allocated, and a PCIe device is a DMA master. In a GPU-passthrough fleet this is the difference between 'the tenant has the two GPUs we gave them' and 'the tenant can talk to whatever else is on the fabric'.","attack_vector":"Host kernel or hypervisor code on an Altra / Altra Max node. On bare-metal rental the tenant already has this. On a virtualised Arm host it needs a prior host-kernel compromise.","remediation":"Update Altra / Altra Max platform firmware to 2.10c or later via the board OEM. Flash + reboot + drain. There is no software workaround - the access control lives in firmware. Compensating control for anyone who cannot patch quickly: stop relying on 'firmware disabled that root complex' as an isolation boundary and physically depopulate or electrically isolate slots that must not be reachable.","references":["https://amperecomputing.com/products/security-bulletins/root-complex-OS-re-enable","https://nvd.nist.gov/vuln/detail/CVE-2022-46892"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50717","cve":"CVE-2022-50717","aliases":["nvmet-tcp Transfer Tag bounds check","NVMe/TCP H2C ttag out-of-bounds"],"title":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c: TENANT ISOLATION: The NVMe/TCP target used the host-supplied Transfer Tag directly as an array index to look…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c","year":"2022","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: The NVMe/TCP target used the host-supplied Transfer Tag directly as an array index to look up the command structure, with no bounds check. A connected initiator sets an arbitrary ttag in an H2C Data PDU and the target reads and operates on memory outside the command array - remote out-of-bounds access on the storage node with full confidentiality, integrity and availability impact. Because NVMe/TCP requires no authentication by default, the 'connected initiator' bar is effectively 'anyone who can reach port 4420'. This one has been in shipping kernels since NVMe/TCP target support landed and was only assigned a CVE retroactively, so long-lived storage nodes are the ones to check.","attack_vector":"Open an NVMe/TCP connection to the target and send an H2CData PDU with an out-of-range Transfer Tag. No authentication needed unless DH-HMAC-CHAP has been explicitly configured. Reachable across any routed path to the target port.","remediation":"Host reboot / kernel upgrade - and specifically check long-lived storage nodes, since the fix was backported late and a node that has not been rebooted in a year may still be exposed. Interim: firewall NVMe/TCP 4420 to known initiators and enable in-band DH-HMAC-CHAP on a patched kernel so the connection itself requires credentials. Roll targets in waves behind multipath so tenants see no outage.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2022/CVE-2022-50717.json","https://nvd.nist.gov/vuln/detail/CVE-2022-50717"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0052","cve":"CVE-2023-0052","aliases":["CVE-2023-0053","ICSA-23-012-05"],"title":"SAUTER Controls Nova 200-220 series (firmware <=3.3-006) with BACnetstac <=4.2.1: Commands execute with no credentials at all, and the only management protocols the device offers are Telnet…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SAUTER Controls Nova 200-220 series (firmware <=3.3-006) with BACnetstac <=4.2.1","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Commands execute with no credentials at all, and the only management protocols the device offers are Telnet and FTP - both cleartext. An unauthorized user can log in, change the device configuration and modify the control program. These are HVAC automation stations; whoever holds them holds the air handling and chilled-water sequences they run. The cleartext management path compounds it: any credentials that do exist elsewhere in the building, entered through these devices, are recoverable by passive sniffing on the facility VLAN, which turns one weak controller into a credential source for the rest of the BMS. In a GPU hall the direct consequence is loss of thermal control with a hardware-damage tail; the indirect one is that your entire building-controls credential set should be considered compromised if these devices are present and the network is shared.","attack_vector":"Unauthenticated Telnet/FTP from anywhere on the facility network. There is nothing to bypass. Passive sniffing on the same segment additionally yields any credentials in flight. Internet exposure of Telnet on building controllers is a recurring Shodan finding, so check your external surface for port 23 as well.","remediation":"SAUTER's guidance is upgrade to fixed firmware where available and otherwise disable the affected services - but on this generation, disabling Telnet and FTP removes the only management path the device has, which is why sites leave them on. Treat this as effectively unpatchable in place: the durable fix is replacing the controller generation, a capital project with a contractor and per-device downtime. Interim controls: strict VLAN isolation with an allow-list from the supervisor only, switch ACLs blocking 21/23 from everything else, and physical security on the panels. Leased colo: this is landlord equipment, so the honest remediation is contractual - require disclosure of controller make/model/firmware across the mechanical plant and the right to audit the segment.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-012-05","https://nvd.nist.gov/vuln/detail/CVE-2023-0052","https://nvd.nist.gov/vuln/detail/CVE-2023-0053"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-20596","cve":"CVE-2023-20596","aliases":[],"title":"AMD SMM Supervisor (AMD-SB-7011): MULTI-TENANT ISOLATION: The highest-scored AMD platform CVE in this database at 9.8 critical. A flaw in the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SMM Supervisor (AMD-SB-7011)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: The highest-scored AMD platform CVE in this database at 9.8 critical. A flaw in the AMD SMM Supervisor - the component that is supposed to *constrain* what SMM code can do - yields full compromise of confidentiality, integrity and availability. SMM sits above the hypervisor and can reach all physical memory; owning it means owning every VM, container and confidential guest on the node, persistently and invisibly to anything running above.","attack_vector":"NVD scores this as network-reachable with no privileges required, which is unusually severe for an SMM issue and worth treating at face value until you can prove otherwise for your platform. AMD's own framing is narrower. Given the disagreement, patch first and reconcile the vector later.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Treat as the top of your AMD firmware queue purely on score and blast radius. If your OEM has not shipped a BIOS carrying it, escalate with them rather than waiting on the normal cycle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20596","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7011.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-2780","cve":"CVE-2023-2780","aliases":[],"title":"MLflow: Path traversal prior to 2.3.1","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Path traversal prior to 2.3.1","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2780"],"status":"curated"},{"id":"CVE-2023-27997","cve":"CVE-2023-27997","aliases":["FG-IR-23-097","XORtigate"],"title":"Fortinet FortiOS / FortiProxy SSL-VPN: A heap-based buffer overflow in the SSL-VPN daemon lets a remote, unauthenticated attacker run arbitrary code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Fortinet FortiOS / FortiProxy SSL-VPN","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"A heap-based buffer overflow in the SSL-VPN daemon lets a remote, unauthenticated attacker run arbitrary code on the FortiGate — full device compromise. This is the 'XORtigate' bug, confirmed in CISA's KEV catalog as actively exploited; if this FortiGate is the VPN gateway into your cluster's management network, an attacker doesn't need any credentials to get a foothold there.","attack_vector":"Remote, unauthenticated — a specifically crafted request to the SSL-VPN service is sufficient, no login required.","remediation":"Firmware upgrade of FortiOS/FortiProxy to the fixed release per Fortinet PSIRT FG-IR-23-097. Given confirmed active exploitation, patch immediately rather than waiting for a scheduled window, and assume compromise on any internet-facing unit that was unpatched during the exploitation window — a reboot alone doesn't remediate a box that was already popped.","references":["https://fortiguard.com/psirt/FG-IR-23-097","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2023-27997"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-29374","cve":"CVE-2023-29374","aliases":[],"title":"LangChain (`LLMMathChain`): Prompt injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (`LLMMathChain`)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Prompt injection → arbitrary Python execution","attack_vector":"Untrusted prompt or retrieved document reaching an agent running on the GPU node","remediation":"No patch for the pattern — LLM output feeding `exec` is the design. Sandbox agent execution; provider isolates the node's credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29374"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-31024","cve":"CVE-2023-31024","aliases":[],"title":"DGX A100 BMC: RCE on BMC (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE on BMC (stack buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash BMC 00.22.05+ out-of-band immediately; isolate BMC VLAN","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31024","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2023-31030","cve":"CVE-2023-31030","aliases":[],"title":"DGX A100 BMC: RCE on BMC (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE on BMC (stack buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash BMC 00.22.05+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31030","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2023-32484","cve":"CVE-2023-32484","aliases":[],"title":"Dell Enterprise SONiC OS (input validation): Improper input validation on Dell Networking switches running Enterprise SONiC, exploitable by a remote…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (input validation)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Improper input validation on Dell Networking switches running Enterprise SONiC, exploitable by a remote unauthenticated attacker. Affects 4.1.0, 4.0.5, 3.5.4 and below — i.e. essentially every SONiC release before the 2024 hardening pass. If you bought Dell switches with SONiC for a cost-optimised GPU buildout in 2022-2023 and have not touched the NOS since, this is live.","attack_vector":"Unauthenticated, remote to the switch.","remediation":"NOS image upgrade and switch reboot. On SONiC that is a full image install, so the switch is down for minutes — stage across MLAG pairs. Put SONiC management interfaces on an isolated OOB network as the standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32484"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-32485","cve":"CVE-2023-32485","aliases":[],"title":"Dell SmartFabric Storage Software: Improper input validation in Dell SmartFabric Storage Software 1.3 and lower, exploitable by a remote…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell SmartFabric Storage Software","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Improper input validation in Dell SmartFabric Storage Software 1.3 and lower, exploitable by a remote unauthenticated attacker. SmartFabric Storage Software is the NVMe-over-TCP fabric controller — it performs the discovery and zoning that decides which host initiators can see which NVMe subsystems. Compromising it is compromising the storage access-control layer for the whole cluster. Companion unauthenticated command injection: CVE-2022-31232.","attack_vector":"Unauthenticated, remote to the SmartFabric Storage Software service.","remediation":"Upgrade the SmartFabric Storage Software appliance/VM past 1.3 (and past 1.4 for the CVE-2023-4306x set). Application upgrade with a service restart; NVMe-oF sessions reconnect. Afterwards, re-verify the zoning database against intent, because an attacker with control here would change exactly that.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32485","https://nvd.nist.gov/vuln/detail/CVE-2022-31232"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2023-3265","cve":"CVE-2023-3265","aliases":["ZDI-23-1147"],"title":"CyberPower PowerPanel Enterprise DCIM - username handling: Authentication bypass: appending a non-printable character to the built-in 'cyberpower' username logs an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM - username handling","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass: appending a non-printable character to the built-in 'cyberpower' username logs an attacker straight in. Unauthenticated to full DCIM administrator, with no exploit development required. PowerPanel Enterprise manages UPS and PDU estates, so this is direct PHYSICAL exposure of the power layer.","attack_vector":"Unauthenticated, remote, against the PowerPanel Enterprise login. Anyone who can reach the web interface.","remediation":"Upgrade PowerPanel Enterprise. Software upgrade on one host - genuinely cheap. Then check whether the default 'cyberpower' account exists at all and remove it. Get the DCIM off any network a tenant workload can route to.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3265"],"status":"curated"},{"id":"CVE-2023-3266","cve":"CVE-2023-3266","aliases":["ZDI-23-1148"],"title":"CyberPower PowerPanel Enterprise DCIM - LDAP authentication path: If LDAP authentication is selected, the authentication mechanism is incomplete and every check can be…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM - LDAP authentication path","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"If LDAP authentication is selected, the authentication mechanism is incomplete and every check can be bypassed. The operators most likely to be affected are the mature ones - the ones who wired their DCIM into corporate directory rather than using local accounts. Doing the responsible thing is what turns the bug on.","attack_vector":"Unauthenticated, remote, on any PowerPanel Enterprise instance configured for LDAP.","remediation":"Upgrade PowerPanel Enterprise immediately. As an interim, switching off LDAP mode removes the vulnerable path but costs you central account control - a genuinely unpleasant trade, so prioritise the upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3266"],"status":"curated"},{"id":"CVE-2023-34048","cve":"CVE-2023-34048","aliases":[],"title":"VMware vCenter: Out-of-bounds write in the DCERPC implementation - unauthenticated remote code execution","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Out-of-bounds write in the DCERPC implementation - unauthenticated remote code execution; exploited as a zero-day by UNC3886 since at least 2021 [KEV]","attack_vector":"Unauthenticated network to the management plane","remediation":"vCenter patch + service restart. Assume compromise on any vCenter that was internet- or tenant-reachable before Oct 2023","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34048"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2023-34362","cve":"CVE-2023-34362","aliases":[],"title":"Progress MOVEit Transfer: Unauthenticated SQL injection into the web app","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Progress MOVEit Transfer","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Unauthenticated SQL injection into the web app -> DB access and RCE (Cl0p mass exploitation)","attack_vector":"Network (remote)","remediation":"Control-plane: patch or decommission; assume data exfiltration if it was internet-facing","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34362"],"status":"curated"},{"id":"CVE-2023-3519","cve":"CVE-2023-3519","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Unauthenticated remote code execution on the gateway appliance","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Unauthenticated remote code execution on the gateway appliance","attack_vector":"Network (remote)","remediation":"Control-plane: emergency patch plus a webshell hunt","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3519"],"status":"curated"},{"id":"CVE-2023-35861","cve":"CVE-2023-35861","aliases":[],"title":"Supermicro BMC email/SMTP alert notification handler (H12DST-B): Command execution as root on the BMC, reached through a feature nearly every operator turns on because they…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC email/SMTP alert notification handler (H12DST-B)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Command execution as root on the BMC, reached through a feature nearly every operator turns on because they want temperature and PSU alerts. Root on the BMC is total control of the node's out-of-band plane: power, boot device, serial and graphical console, and the ability to write persistent code that outlives any host reinstall. On a rented bare-metal GPU node it also means the previous tenant's implant can be watching the next tenant's console. Firmware 03.10.35 and present across the same BMC codebase on other Supermicro boards. User-controlled notification fields reach a shell without sanitisation.","attack_vector":"Reachable over the network to the BMC. The alerting configuration surface is exactly the kind of thing left enabled and reachable from the monitoring VLAN, so an attacker who compromises a monitoring or DCIM host is already in position.","remediation":"Firmware flash to 03.10.35 or later per board SKU, from Supermicro's June 2023 SMTP advisory. As an immediate config-only stopgap you can disable BMC email alerting entirely, which removes the vulnerable path at the cost of losing hardware alerts - acceptable for a few days, not as a permanent posture. Because this bug is in shared Supermicro BMC code rather than one board's, audit the whole fleet rather than only H12DST-B.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-35861","https://blog.freax13.de/cve/cve-2023-35861","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/35xxx/CVE-2023-35861.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-36258","cve":"CVE-2023-36258","aliases":[],"title":"LangChain (PALChain): Arbitrary code execution via `os.system`/`exec` in generated code","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (PALChain)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution via `os.system`/`exec` in generated code","attack_vector":"Untrusted prompt input","remediation":"Upgrade past 0.0.236; the fix was bypassed twice (see below)","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36258"],"status":"curated"},{"id":"CVE-2023-36281","cve":"CVE-2023-36281","aliases":[],"title":"LangChain (`load_prompt`): Arbitrary code execution from a JSON prompt file","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (`load_prompt`)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution from a JSON prompt file","attack_vector":"Customer-supplied prompt file","remediation":"Upgrade; prompt files are executable content","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36281"],"status":"curated"},{"id":"CVE-2023-36845","cve":"CVE-2023-36845","aliases":[],"title":"Juniper Junos OS J-Web (EX/SRX): **[KEV]** Unauthenticated remote code execution by setting `PHPRC` through a crafted J-Web request","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS J-Web (EX/SRX)","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"**[KEV]** Unauthenticated remote code execution by setting `PHPRC` through a crafted J-Web request; chained with the other J-Web bugs for full device takeover","attack_vector":"Network, unauthenticated","remediation":"Junos upgrade or disabling J-Web entirely; on a management-plane switch the fastest mitigation is turning J-Web off, which removes the GUI ops staff depend on","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36845"],"status":"curated"},{"id":"CVE-2023-38408","cve":"CVE-2023-38408","aliases":[],"title":"OpenSSH (ssh-agent): Remote code execution in ssh-agent PKCS#11 support when agent forwarding reaches a hostile host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (ssh-agent)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Remote code execution in ssh-agent PKCS#11 support when agent forwarding reaches a hostile host","attack_vector":"Unauthenticated network (against a forwarded agent)","remediation":"Package update; restart agents. Policy fix: forbid agent forwarding into tenant-reachable bastions","references":["https://access.redhat.com/security/cve/CVE-2023-38408"],"status":"curated"},{"id":"CVE-2023-39281","cve":"CVE-2023-39281","aliases":["INSYDE-SA-2023054"],"title":"Insyde InsydeH2O (AsfSecureBootDxe): Stack buffer overflow leading to arbitrary code execution during the DXE phase - and it lives in the driver…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AsfSecureBootDxe)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow leading to arbitrary code execution during the DXE phase - and it lives in the driver responsible for Secure Boot handling for ASF (Alert Standard Format, the out-of-band manageability path). Code execution in DXE means running before the OS with firmware privileges and with Secure Boot policy still under the attacker's influence. Scored critical, and the placement inside the Secure Boot path is what makes it worse than the raw score suggests.","attack_vector":"Attacker able to supply the oversized input the DXE driver parses during boot. Given the ASF/manageability association, treat anything that can reach the platform's out-of-band alerting path as in scope alongside local OS-level access.","remediation":"OEM BIOS update on the fixed Insyde kernel (5.0-5.5 affected). Firmware flash, reboot per node. No config workaround inside firmware; on the network side, keep the BMC and manageability interfaces on an isolated management VLAN with no route from tenant or provisioning networks - which is good practice independent of this CVE and materially reduces who can reach the ASF path.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39281","https://www.insyde.com/security-pledge/SA-2023054"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-41910","cve":"CVE-2023-41910","aliases":[],"title":"lldpd (CDP PDU parser, cdp_decode): A crafted CDP PDU with specific CDP_TLV_ADDRESSES TLVs forces lldpd into an out-of-bounds heap read. lldpd is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"lldpd (CDP PDU parser, cdp_decode)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A crafted CDP PDU with specific CDP_TLV_ADDRESSES TLVs forces lldpd into an out-of-bounds heap read. lldpd is what runs on Linux-based switch OSes and on servers that advertise link topology, it runs as root, and it accepts input from any directly attached device with no authentication whatsoever. In a GPU cluster where LLDP is used to verify rail-optimized cabling, lldpd is running on every node and every switch.","attack_vector":"Unauthenticated, adjacent — a single crafted frame from a directly connected device. Any tenant bare-metal node can attack the switch or the neighbours it is cabled to.","remediation":"Upgrade lldpd to 1.0.17 or later and restart the daemon — a package upgrade with a service restart, no reboot, no switch reload. On appliance NOSes this arrives as a NOS image update instead. Cheap fix; the reason it lingers is that nobody inventories lldpd versions.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-41910"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2023-42793","cve":"CVE-2023-42793","aliases":[],"title":"JetBrains TeamCity: Authentication bypass leading to remote code execution on TeamCity Server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"JetBrains TeamCity","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Authentication bypass leading to remote code execution on TeamCity Server","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; treat previously built artifacts as suspect","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-42793"],"status":"curated"},{"id":"CVE-2023-4323","cve":"CVE-2023-4323","aliases":["CVE-2023-4324","CVE-2023-4325","CVE-2023-4326","CVE-2023-4329","CVE-2023-4331","CVE-2023-4332","CVE-2023-4333","CVE-2023-4345"],"title":"Broadcom LSI Storage Authority (LSA) / Intel RAID Web Console 3 (RWC3) - management service for MegaRAID and LSI HBA controllers, all versions before 7.017.011.000: LSA is the agent that fronts…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Broadcom LSI Storage Authority (LSA) / Intel RAID Web Console 3 (RWC3) - management service for MegaRAID and LSI HBA…","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"LSA is the agent that fronts the MegaRAID/LSI controller on every server that has one, and it runs as root/SYSTEM with the ability to create, delete and re-initialize virtual drives and to push controller firmware. Sessions are not invalidated properly in Gateway (multi-node) setup, so an attacker who reaches the LSA web service can ride an admin session and drive the RAID controller on any managed node: wipe or re-create arrays under a running tenant, or stage a controller firmware update. Because the controller sits below the OS with DMA to host memory, a firmware push through this path is below-the-OS persistence that a tenant reimage does not remove - it breaks tenant handoff. The sibling issues in the same disclosure are the supporting weaknesses: a bundled vulnerable libcurl, no CSP, no SameSite on the session cookie, SHA-1 ciphersuites and obsolete TLS versions on the management listener, and world-readable log files.","attack_vector":"Any host that can reach the LSA HTTPS listener (default TCP 2463) on the management or provisioning network. In Gateway mode one LSA instance manages many nodes, so one reachable management endpoint fans out to the whole fleet. No valid credentials are needed to abuse the session-handling flaw; the weak TLS and cookie defaults widen it to on-path and browser-side attackers.","remediation":"Software-only: upgrade LSA / Intel RWC3 to 7.017.011.000 or later on every managed node and on the Gateway. No controller firmware flash and no array downtime - restarting the LSA service is enough. The real cost is that this agent is installed by OEM tooling on every server with a Broadcom controller, and Dell/HPE/Supermicro/Intel ship their own rebadged LSA build months behind Broadcom, so you often cannot take the upstream package. Interim control: bind LSA to loopback or a dedicated management VLAN and firewall 2463 off the tenant and provisioning networks; if you do not use the web UI, uninstall LSA and drive the controller with StorCLI from a config-managed path instead.","references":["https://www.broadcom.com/support/resources/product-security-center","https://nvd.nist.gov/vuln/detail/CVE-2023-4323","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00926.html"],"status":"curated"},{"id":"CVE-2023-43845","cve":"CVE-2023-43845","aliases":[],"title":"ATEN PE6208 switched PDU: The PDU ships with a default telnet account and never forces the operator to change it on first login. Anyone…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATEN PE6208 switched PDU","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The PDU ships with a default telnet account and never forces the operator to change it on first login. Anyone who finds an un-rotated unit gets an administrator telnet session — full outlet control, including turning power off to whatever racks that PDU feeds.","attack_vector":"Network reachability to the PDU's telnet service plus knowledge of the published default credential; no exploit development needed.","remediation":"Credential rotation is the immediate fix — audit every deployed PE6208 for the default telnet account and change it now. ATEN's firmware update additionally forces a credential change on first login for new deployments, so pair the audit with a firmware upgrade where feasible.","references":["https://github.com/setersora/pe6208"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-44467","cve":"CVE-2023-44467","aliases":[],"title":"langchain-experimental (PALChain): Bypass of the CVE-2023-36258 fix","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental (PALChain)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Bypass of the CVE-2023-36258 fix","attack_vector":"Untrusted prompt input","remediation":"Upgrade past 0.0.306","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-44467"],"status":"curated"},{"id":"CVE-2023-45249","cve":"CVE-2023-45249","aliases":[],"title":"Acronis Cyber Infrastructure: Default passwords","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Acronis Cyber Infrastructure","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Default passwords -> unauthenticated remote command execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch + change all default credentials on the storage/backup appliance","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45249"],"status":"curated"},{"id":"CVE-2023-48022","cve":"CVE-2023-48022","aliases":["ShadowRay"],"title":"Ray (job submission API): Unauthenticated RCE — the Jobs API accepts arbitrary code by design","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (job submission API)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE — the Jobs API accepts arbitrary code by design","attack_vector":"Unauthenticated network to the Ray dashboard/Jobs API (default 8265)","remediation":"**No patch — vendor disputes it.** The only remediation is network isolation and an auth proxy. Exploited in the wild against GPU clusters; a neocloud must never let 8265 reach a tenant or public network","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48022"],"status":"curated","fleet":{"ubiquity":"Very common - Ray is the default distributed-compute layer for RL, tuning and batch inference on GPU clusters; ~230k exposed instances observed","remediation_pain":"`node-drain` + network re-architecture - there is no vendor patch (Anyscale calls it intended behavior), so remediation means putting auth in front of every dashboard/Jobs API and restarting every Ray cluster in the fleet","pain_class":"node-drain","why_fleet_wide":"The Jobs API has no authorization: anyone who can reach the dashboard submits arbitrary code to the whole cluster, so one exposed head node hands over every GPU and every dataset attached to it"}},{"id":"CVE-2023-48788","cve":"CVE-2023-48788","aliases":[],"title":"Fortinet FortiClient EMS: Unauthenticated SQL injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiClient EMS","year":"2023","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Unauthenticated SQL injection -> command execution as SYSTEM","attack_vector":"Network (remote)","remediation":"Control-plane: patch the endpoint-management server","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48788"],"status":"curated"},{"id":"CVE-2023-49934","cve":"CVE-2023-49934","aliases":[],"title":"Slurm: SQL injection against the SlurmDBD accounting database","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SQL injection against the SlurmDBD accounting database","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm to 23.11.1+; audit accounting data","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49934"],"status":"curated"},{"id":"CVE-2023-49937","cve":"CVE-2023-49937","aliases":[],"title":"Slurm: Double free allowing denial of service or possibly arbitrary code execution","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Double free allowing denial of service or possibly arbitrary code execution","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49937"],"status":"curated"},{"id":"CVE-2023-6014","cve":"CVE-2023-6014","aliases":[],"title":"MLflow: Arbitrary account creation bypassing authentication","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary account creation bypassing authentication","attack_vector":"Unauthenticated network to the tracking server","remediation":"Upgrade; the basic-auth plugin is not a boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6014"],"status":"curated","fleet":{"ubiquity":"Common - same MLflow footprint; this is the incomplete-fix follow-on","remediation_pain":"`daemon-restart` - server upgrade, plus rotation of everything the server could reach","pain_class":"daemon-restart","why_fleet_wide":"Basic-auth bypass on the tracking server, so the one control that was supposed to contain the previous LFI does not hold; full model-registry and artifact-store compromise"}},{"id":"CVE-2023-6018","cve":"CVE-2023-6018","aliases":[],"title":"MLflow: Overwrite any file on the MLflow host without authentication","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Overwrite any file on the MLflow host without authentication","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6018"],"status":"curated"},{"id":"CVE-2023-6019","cve":"CVE-2023-6019","aliases":[],"title":"Ray (dashboard `cpu_profile`): Command injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (dashboard `cpu_profile`)","year":"2023","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Command injection → OS command execution, unauthenticated","attack_vector":"Unauthenticated network to the Ray dashboard","remediation":"Upgrade to 2.8.1+ and isolate the dashboard","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6019"],"status":"curated"},{"id":"CVE-2024-0012","cve":"CVE-2024-0012","aliases":[],"title":"Palo Alto PAN-OS: Management web interface authentication bypass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Management web interface authentication bypass -> PAN-OS administrator privileges","attack_vector":"Network (remote)","remediation":"Control-plane: patch + remove the management interface from the internet","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0012"],"status":"curated"},{"id":"CVE-2024-0138","cve":"CVE-2024-0138","aliases":[],"title":"Base Command Manager (CMDaemon): Unauthenticated RCE on the cluster manager","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Base Command Manager (CMDaemon)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE on the cluster manager -> full cluster takeover","attack_vector":"Network-adjacent unauthenticated attacker reaching CMDaemon","remediation":"Emergency: patch Base Command Manager, restrict CMDaemon to the mgmt network, audit for compromise; cluster-wide credential rotation","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=Base%20Command%20Manager"],"status":"curated","fleet":{"ubiquity":"Common - Base Command / Bright Cluster Manager is the control plane on many enterprise and neocloud GPU clusters","remediation_pain":"`daemon-restart` of the cluster control plane, which is itself a scheduling outage for the whole cluster","pain_class":"daemon-restart","why_fleet_wide":"Missing authentication in CMDaemon, remotely exploitable with no user interaction or privileges: compromising the cluster manager means owning provisioning for every node in the cluster at once"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-862"]},{"id":"CVE-2024-11041","cve":"CVE-2024-11041","aliases":[],"title":"vLLM (MessageQueue / ZMQ): `pickle.loads` on socket data","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (MessageQueue / ZMQ)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`pickle.loads` on socket data → unauthenticated RCE","attack_vector":"Unauthenticated network to the vLLM internal ZMQ socket, reachable by a co-tenant","remediation":"Upgrade; bind ZMQ to loopback and enforce per-tenant network policy. Internal IPC sockets must never cross the tenant boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11041"],"status":"curated"},{"id":"CVE-2024-1305","cve":"CVE-2024-1305","aliases":[],"title":"OpenVPN (tap-windows6): Unchecked write size","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN (tap-windows6)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unchecked write size -> memory buffer overflow and potential arbitrary code execution in kernel space","attack_vector":"Network (remote)","remediation":"Control-plane: operator endpoint driver update","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1305"],"status":"curated"},{"id":"CVE-2024-21652","cve":"CVE-2024-21652","aliases":[],"title":"Argo CD: Chained DoS plus in-memory data manipulation lets an attacker bypass authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Chained DoS plus in-memory data manipulation lets an attacker bypass authentication","attack_vector":"Unauthenticated network reaching the Argo CD API","remediation":"Rolling Argo CD upgrade to 2.8.13/2.9.9/2.10.4+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21652"],"status":"curated"},{"id":"CVE-2024-21762","cve":"CVE-2024-21762","aliases":[],"title":"Fortinet FortiOS: SSL-VPN out-of-bounds write","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] SSL-VPN out-of-bounds write -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware; rotate all VPN user credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21762"],"status":"curated"},{"id":"CVE-2024-2221","cve":"CVE-2024-2221","aliases":[],"title":"Qdrant (snapshot upload): Path traversal + arbitrary file upload via `/collections/{c}/snapshots/upload`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Qdrant (snapshot upload)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Path traversal + arbitrary file upload via `/collections/{c}/snapshots/upload`","attack_vector":"Network user able to upload a snapshot","remediation":"Upgrade; snapshot upload is an arbitrary-write primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-2221"],"status":"curated"},{"id":"CVE-2024-23653","cve":"CVE-2024-23653","aliases":[],"title":"BuildKit: Interactive-container API lacks entitlement checks, so a build can run a privileged container and escape the…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Interactive-container API lacks entitlement checks, so a build can run a privileged container and escape the builder","attack_vector":"Anyone who can submit a build to a shared builder","remediation":"Upgrade BuildKit; rebuild builder nodes; never share one BuildKit daemon across tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23653"],"status":"curated"},{"id":"CVE-2024-23897","cve":"CVE-2024-23897","aliases":[],"title":"Jenkins: CLI parser expands `@file` into argument contents","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] CLI parser expands `@file` into argument contents -> unauthenticated arbitrary file read, chains to RCE","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT patch; disable the CLI; rotate every credential in the Jenkins store","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23897"],"status":"curated"},{"id":"CVE-2024-2421","cve":"CVE-2024-2421","aliases":["CVE-2024-2420","CVE-2024-2422","ICSA-24-151-01"],"title":"LenelS2 NetBox access control and event monitoring system (<=5.6.1): Unauthenticated remote code execution with elevated privileges, plus hardcoded credentials that bypass…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"LenelS2 NetBox access control and event monitoring system (<=5.6.1)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution with elevated privileges, plus hardcoded credentials that bypass authentication outright, plus a second authenticated RCE. NetBox is the head-end - the system that holds the cardholder database, the access rules, the door schedules and the event log for the whole site. Owning it is strictly more powerful than owning a single panel: an attacker can grant themselves a credential that works on every door, schedule doors to unlock, and erase the events showing they did. Physical entry to the hall and the cage follows, and from inside the cage the attacker reaches drives holding customer data and model weights, server console ports, and the out-of-band management switch that fronts every BMC in the row. For a bare-metal GPU provider, an attacker with head-end control can also target one specific tenant's cage on demand, which turns a security incident into a customer-trust and contractual event. Hardcoded credentials mean the exposure predates any breach you can detect.","attack_vector":"Unauthenticated over the network for the RCE and the hardcoded-credential bypass. NetBox is a web-managed appliance; sites routinely make it reachable from the corporate network so security staff can administer badges, and internet-exposed NetBox instances have been observed. Any of those makes this a direct, no-credential path from outside to physical door control.","remediation":"Upgrade NetBox past 5.6.1 to the fixed release per Carrier's advisory - a head-end software update in a normal change window, no door hardware touched, so there is no operational excuse to defer. Because hardcoded credentials were present, patching alone does not establish trust: rotate all system and integration credentials, audit the cardholder database and access rules against a known-good export, review the event log for gaps, and check whether any credentials were added during the exposure window. Then remove NetBox from any internet-facing or general corporate reachability and put it behind a jump host with MFA.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-24-151-01","https://nvd.nist.gov/vuln/detail/CVE-2024-2421","https://nvd.nist.gov/vuln/detail/CVE-2024-2420","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated"},{"id":"CVE-2024-24592","cve":"CVE-2024-24592","aliases":[],"title":"ClearML fileserver: No authentication — arbitrary read/write/delete of all stored artifacts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML fileserver","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"No authentication — arbitrary read/write/delete of all stored artifacts","attack_vector":"Unauthenticated network to the fileserver","remediation":"Upgrade and front with auth. All experiment data and models are exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24592"],"status":"curated"},{"id":"CVE-2024-27198","cve":"CVE-2024-27198","aliases":[],"title":"JetBrains TeamCity: Alternative-path authentication bypass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"JetBrains TeamCity","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Alternative-path authentication bypass -> unauthenticated admin actions on the CI server","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade; rotate all build secrets and signing keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27198"],"status":"curated"},{"id":"CVE-2024-27444","cve":"CVE-2024-27444","aliases":[],"title":"langchain-experimental: Second bypass of CVE-2023-44467","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Second bypass of CVE-2023-44467","attack_vector":"Untrusted prompt input","remediation":"Upgrade past 0.1.8. Three CVEs for one sandbox — code-generating chains have no safe configuration","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27444"],"status":"curated"},{"id":"CVE-2024-29849","cve":"CVE-2024-29849","aliases":[],"title":"Veeam Backup Enterprise Manager: Unauthenticated users can log in as any user to the Enterprise Manager web interface","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup Enterprise Manager","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated users can log in as any user to the Enterprise Manager web interface","attack_vector":"Network (remote)","remediation":"Control-plane: patch or decommission Enterprise Manager","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29849"],"status":"curated"},{"id":"CVE-2024-32053","cve":"CVE-2024-32053","aliases":[],"title":"CyberPower PowerPanel platform - hardcoded database, service and cloud credentials: Hardcoded credentials used by the platform to authenticate to its database, to other services and to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel platform - hardcoded database, service and cloud credentials","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Hardcoded credentials used by the platform to authenticate to its database, to other services and to CyberPower's cloud. The cloud element is what makes this worth calling out separately: the credential is shared across deployments, so an attacker who extracts it once has a position against many operators' installations at the vendor's back end, not just yours.","attack_vector":"Anyone with the shipped software. The cloud credential in particular is reachable from the internet by design.","remediation":"Vendor upgrade is the only fix. Ask the vendor directly whether the cloud-side credential was rotated on their end - your upgrade does not do that. If the answer is unsatisfying, disable the cloud integration.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32053"],"status":"curated"},{"id":"CVE-2024-32735","cve":"CVE-2024-32735","aliases":[],"title":"CyberPower PowerPanel Enterprise prior to v2.8.3 - PDNU REST APIs: Certain utility REST APIs have no authentication at all, giving an unauthenticated remote attacker a direct…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise prior to v2.8.3 - PDNU REST APIs","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Certain utility REST APIs have no authentication at all, giving an unauthenticated remote attacker a direct route into the application. Yet another unauthenticated path into a system that controls power distribution - the pattern across this product line is that authentication is applied per-endpoint and repeatedly missed.","attack_vector":"Unauthenticated, remote, to the PowerPanel Enterprise API surface.","remediation":"Upgrade to v2.8.3 or later. Given the density of unauthenticated findings in this product across 2023 and 2024, an operator should also decide whether it belongs in the design at all, or whether the power estate should be monitored through something with a better track record.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32735"],"status":"curated"},{"id":"CVE-2024-33625","cve":"CVE-2024-33625","aliases":[],"title":"CyberPower PowerPanel business application - JWT signing key: The JWT signing key is hardcoded in the application, so an attacker forges any token they like and becomes…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel business application - JWT signing key","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The JWT signing key is hardcoded in the application, so an attacker forges any token they like and becomes any user. Same shape as the hardcoded credentials: a secret that is not secret and cannot be rotated by the operator.","attack_vector":"Unauthenticated, remote. Requires only the shipped software to extract the key.","remediation":"Vendor upgrade. Nothing an operator can configure fixes a hardcoded signing key. Isolate the host until patched, and treat any PowerPanel instance that was internet-reachable as compromised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-33625"],"status":"curated"},{"id":"CVE-2024-34025","cve":"CVE-2024-34025","aliases":[],"title":"CyberPower PowerPanel business application - hardcoded authentication credentials: A hardcoded credential set compiled into the application gives administrator access to anyone who reads the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel business application - hardcoded authentication credentials","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A hardcoded credential set compiled into the application gives administrator access to anyone who reads the binary. There is no configuration that removes it and no password rotation that helps. On a platform that controls power distribution, this is a permanent unauthenticated back door until the vendor ships a build without it.","attack_vector":"Unauthenticated, remote, using credentials extractable from the shipped software by anyone.","remediation":"Upgrade to a build that removes the credentials - configuration changes cannot help. Until upgraded, the only real control is network isolation: the PowerPanel host must be unreachable from anything but a management jump box.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-34025"],"status":"curated"},{"id":"CVE-2024-35198","cve":"CVE-2024-35198","aliases":[],"title":"TorchServe: `allowed_urls` bypass","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`allowed_urls` bypass → arbitrary model load → RCE","attack_vector":"Unauthenticated network to the management API","remediation":"Patch; the earlier ShellTorch fix is insufficient. Treat the management API as never tenant-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35198"],"status":"curated"},{"id":"CVE-2024-36435","cve":"CVE-2024-36435","aliases":[],"title":"Supermicro BMC firmware web/management service (X11/X12/X13/H12/H13/B12/B13, CMM6): An attacker who never authenticates gets code execution inside the BMC's own firmware OS. On a GPU fleet that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware web/management service (X11/X12/X13/H12/H13/B12/B13, CMM6)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"An attacker who never authenticates gets code execution inside the BMC's own firmware OS. On a GPU fleet that means out-of-band power control over every affected node, the ability to mount an attacker-supplied ISO as virtual media and reboot a node into it, KVM access to whatever a tenant has on the console, and a persistent implant living beneath the hypervisor and the host OS that survives every reimage, every OS patch and every tenant handoff. Because the BMC also drives the CMM6 in blade chassis, a single compromised management module gives leverage over an entire enclosure rather than one node. H13, B12 and B13 motherboards plus CMM6 blade chassis management modules - i.e. essentially the whole current Supermicro server line including the GPU chassis and SuperBlade enclosures.","attack_vector":"Anything that can open a TCP connection to the BMC's management interface, with no credentials at all. In practice that is anyone with a route to the out-of-band management VLAN - a jump host, a misconfigured L3 leaf, a compromised DCIM or monitoring box, or a BMC that ended up with a public address. No host-side foothold and no tenant workload access is required.","remediation":"Firmware flash, per node, out of band. Supermicro shipped fixed BMC images in its July 2024 BMC/IPMI advisory batch, but the fixed version differs per motherboard SKU, so an operator with a mixed X11/X12/X13 fleet has to build a board-to-image matrix before flashing anything. Budget a BMC reboot per node - the host stays up during a BMC flash on most Supermicro boards but a failed flash can leave the BMC unresponsive and requires a physical recovery, so stage it rack by rack. Until every node is flashed, the only mitigation that actually holds is hard network isolation: BMCs on a dedicated VLAN with no route from tenant networks and an explicit allowlist for the handful of management hosts that need them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36435","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/36xxx/CVE-2024-36435.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-3660","cve":"CVE-2024-3660","aliases":[],"title":"Keras / TensorFlow: Arbitrary code injection in Keras < 2.13 via Lambda-layer model loading","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras / TensorFlow","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code injection in Keras < 2.13 via Lambda-layer model loading","attack_vector":"Customer-supplied `.h5` model","remediation":"Rebuild images off TF/Keras < 2.13; no runtime mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3660"],"status":"curated"},{"id":"CVE-2024-37079","cve":"CVE-2024-37079","aliases":[],"title":"VMware vCenter: Heap overflow in the DCERPC implementation - unauthenticated remote code execution on vCenter","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Heap overflow in the DCERPC implementation - unauthenticated remote code execution on vCenter","attack_vector":"Unauthenticated network to the management plane","remediation":"vCenter patch + service restart. Never expose vCenter to a tenant-reachable network segment","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37079"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-3817","cve":"CVE-2024-3817","aliases":[],"title":"Terraform (go-getter): Argument injection when go-getter shells out to Git for remote branch discovery","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Terraform (go-getter)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Argument injection when go-getter shells out to Git for remote branch discovery -> RCE in the IaC runner","attack_vector":"Network (remote)","remediation":"Control-plane: bump go-getter/Terraform in the CI runner image","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3817"],"status":"curated"},{"id":"CVE-2024-38812","cve":"CVE-2024-38812","aliases":[],"title":"VMware vCenter: Heap overflow in DCERPC - unauthenticated remote code execution on vCenter","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Heap overflow in DCERPC - unauthenticated remote code execution on vCenter [KEV]","attack_vector":"Unauthenticated network to the management plane","remediation":"vCenter patch + service restart; the first patch was incomplete, so verify the build number rather than trusting the advisory date","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38812"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-39236","cve":"CVE-2024-39236","aliases":[],"title":"Gradio: Code injection via `gradio/component_meta.py`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Code injection via `gradio/component_meta.py`","attack_vector":"Attacker-influenced component definition","remediation":"Upgrade past 4.36.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39236"],"status":"curated"},{"id":"CVE-2024-40711","cve":"CVE-2024-40711","aliases":[],"title":"Veeam Backup & Replication: Deserialization of untrusted data","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Deserialization of untrusted data -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT - the backup server owns the restore path; patch, isolate, rotate","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40711"],"status":"curated"},{"id":"CVE-2024-41660","cve":"CVE-2024-41660","aliases":["GHSA-wmgv-jffg-v3xr"],"title":"OpenBMC slpd-lite (Service Location Protocol daemon, UDP 427): slpd-lite is a small SLP responder that OpenBMC installs by default, and it overflows memory on crafted UDP…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC slpd-lite (Service Location Protocol daemon, UDP 427)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"slpd-lite is a small SLP responder that OpenBMC installs by default, and it overflows memory on crafted UDP packets. Because it is in the default build, this is not a niche configuration - if your image came from an OpenBMC tree and nobody explicitly removed the package, the daemon is listening. Unauthenticated remote memory corruption in a BMC-resident network daemon is the entry point that everything else in this cluster builds on: get code running on the BMC, then use the LPC-control or crypto kernel bugs to reach BMC root, then write flash and persist across tenant handover.","attack_vector":"Unauthenticated, network, UDP port 427 on the BMC's management interface. Nothing on the host and no credentials required. SLP is a discovery protocol nobody in a modern GPU fleet actually uses, which makes the exposure pure cost.","remediation":"Two moves and the cheap one is very cheap. Config-only: block UDP 427 at the management-VLAN boundary and, better, remove slpd-lite from the image or stop and mask the service. It provides nothing an operator needs - Redfish discovery does not depend on it. The durable fix is upstream in the slpd-lite repository and reaches nodes through a BMC firmware flash: per node, out-of-band, ODM-lagged. Do the ACL and service disable now; batch the flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41660","https://github.com/openbmc/slpd-lite/security/advisories/GHSA-wmgv-jffg-v3xr"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-4323","cve":"CVE-2024-4323","aliases":[],"title":"Fluent Bit: \"Linguistic Lumberjack\" - memory corruption parsing trace requests in the embedded HTTP server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fluent Bit","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"\"Linguistic Lumberjack\" - memory corruption parsing trace requests in the embedded HTTP server -> RCE","attack_vector":"Network (remote)","remediation":"Data-plane: DaemonSet on every GPU node - fleet rollout + disable /api/v1/traces","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4323"],"status":"curated"},{"id":"CVE-2024-43864","cve":"CVE-2024-43864","aliases":["net/mlx5e fix CT entry update leaks of modify header context"],"title":"Linux kernel mlx5_core TC connection tracking offload: Updating a connection-tracking entry allocates a replacement modify-header context","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core TC connection tracking offload","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Updating a connection-tracking entry allocates a replacement modify-header context; if that allocation fails - which it does once you exceed the firmware's maximum, a state an attacker can drive by opening many connections - the error pointer is stored and later dereferenced on free, panicking the kernel, and the old context is leaked. Remote traffic volume alone is enough to reach it on a node doing hardware conntrack offload.","attack_vector":"Remote, unauthenticated: open enough tracked connections through an mlx5 host doing CT offload to exhaust the firmware's modify-header capacity.","remediation":"Upgrade the host kernel to 6.11 or a stable backport (6.6.45, 6.10.4). Rolling reboot. Interim: disable hardware connection-tracking offload on the mlx5 interfaces or cap conntrack table size, both live config changes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43864","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-43864.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-44970","cve":"CVE-2024-44970","aliases":["net/mlx5e SHAMPO invalid WQ linked list unlink"],"title":"Linux kernel mlx5_core RX datapath (SHAMPO): SHAMPO can deliver completion entries with zero consumed strides for a work queue entry that was already…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (SHAMPO)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SHAMPO can deliver completion entries with zero consumed strides for a work queue entry that was already fully consumed and unlinked, causing a second unlink that corrupts the receive work queue linked list. Remote, unauthenticated corruption of RX ring state on the host NIC - the node's networking dies and the kernel is in an inconsistent state.","attack_vector":"Unauthenticated remote sender of network traffic to an mlx5 host with SHAMPO/HW-GRO active.","remediation":"Upgrade the host kernel to 6.11 or a stable backport (6.1.105, 6.6.46, 6.10.5). Rolling reboot fleet-wide. Same interim mitigation as the other SHAMPO bugs: turn off rx-gro-hw via ethtool, a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44970","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-44970.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45410","cve":"CVE-2024-45410","aliases":[],"title":"Traefik: Traefik-added X-Forwarded-* headers can be spoofed by the client and are trusted by the backend","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Traefik-added X-Forwarded-* headers can be spoofed by the client and are trusted by the backend; authentication bypass at the application layer","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade; explicitly configure trusted forwarded headers","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45410"],"status":"curated"},{"id":"CVE-2024-46717","cve":"CVE-2024-46717","aliases":["net/mlx5e SHAMPO incorrect page release"],"title":"Linux kernel mlx5_core RX datapath (SHAMPO / HW-GRO): A remote sender can make the mlx5 receive path release a SHAMPO header page while it is still in use, giving…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (SHAMPO / HW-GRO)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A remote sender can make the mlx5 receive path release a SHAMPO header page while it is still in use, giving a double release and DMA into freed memory. This runs in NAPI softirq on every inbound frame, before any socket lookup or credential check - so an unauthenticated peer anywhere that can route packets to the node corrupts host kernel memory on the primary datacenter NIC. Note that NVD rates this 5.5 while the assigning kernel.org CNA rates it 9.8; the CNA score is the one that reflects reachability.","attack_vector":"Unauthenticated remote attacker able to send network traffic to a node running mlx5 with HW-GRO/SHAMPO enabled. No account, no VM, no adjacency needed.","remediation":"Upgrade the host kernel to 6.11 or a stable backport (6.1.109, 6.6.50, 6.10.9). Rolling reboot of every ConnectX/BlueField host. Immediate mitigation without reboot: disable HW-GRO on the mlx5 interfaces (ethtool -K <dev> rx-gro-hw off), which takes the SHAMPO path out of service at some throughput cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46717","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-46717.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46946","cve":"CVE-2024-46946","aliases":[],"title":"langchain-experimental: Arbitrary code execution in 0.1.17–0.3.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution in 0.1.17–0.3.0","attack_vector":"Untrusted prompt input","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46946"],"status":"curated"},{"id":"CVE-2024-47167","cve":"CVE-2024-47167","aliases":[],"title":"Gradio: SSRF from the file-upload/proxy path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SSRF from the file-upload/proxy path","attack_vector":"Unauthenticated network","remediation":"Upgrade to 4.44+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47167"],"status":"curated"},{"id":"CVE-2024-47575","cve":"CVE-2024-47575","aliases":[],"title":"Fortinet FortiManager: \"FortiJump\" - missing authentication in fgfmd","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiManager","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] \"FortiJump\" - missing authentication in fgfmd -> unauthenticated RCE on the fleet manager","attack_vector":"Network (remote)","remediation":"Control-plane: patch; FortiManager compromise means whole-fabric compromise","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47575"],"status":"curated"},{"id":"CVE-2024-47943","cve":"CVE-2024-47943","aliases":[],"title":"Rittal IoT Interface and CMC III Processing Unit - firmware upgrade signature check: The admin web interface verifies patch files with an HMAC-style check whose key is a long string hard-coded…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal IoT Interface and CMC III Processing Unit - firmware upgrade signature check","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The admin web interface verifies patch files with an HMAC-style check whose key is a long string hard-coded in the firmware - and the firmware is freely downloadable, so anyone can extract the key and sign their own malicious update. That is a complete bypass of firmware integrity on the device that monitors cabinet temperature, humidity, door state, smoke and access for a rack or a row. An attacker who lands code there gets persistence on facility hardware nobody re-images, the ability to falsify environmental telemetry so a real thermal excursion in a GPU row never raises an alarm, and control of whatever access and door outputs the CMC III drives. The falsified-telemetry angle is the one operators underestimate: your temperature monitoring is the thing that is supposed to tell you a cooling attack is underway, and this lets the attacker turn it into a liar while the racks cook. The IoT Interface variant is a general-purpose gateway, so the same implant is a pivot point onto the facility network.","attack_vector":"Access to the admin web interface to upload the crafted patch - so an authenticated or otherwise-reachable path to the device on the facility VLAN. CMC III processing units and IoT Interfaces are typically on the monitoring network alongside PDU cards and environmental sensors, reachable from the DCIM collector and often from anything else on that segment. Given the same product family's history of backdoor accounts and injection flaws, reaching the admin interface is not a high bar.","remediation":"Rittal firmware update - a small device, quick flash, no cooling impact, so this is one of the more tractable items on this list. Do it across every CMC III processing unit and IoT Interface, and note that the fleet count in a large hall is per-row or per-cabinet, so it is a technician-day of work rather than a single change. Because the flaw defeats firmware integrity, any unit you believe was reachable by an untrusted host should be re-flashed from vendor media rather than trusted after an in-place update. Then isolate: monitoring devices on their own VLAN, admin interfaces reachable only from a jump host, and no route from tenant or corporate networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47943","https://r.sec-consult.com/rittaliot","https://seclists.org/fulldisclosure/2024/Oct/4"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-48063","cve":"CVE-2024-48063","aliases":[],"title":"PyTorch (`torch.distributed` RemoteModule / RPC): Deserialization RCE across the distributed RPC channel (vendor-disputed as intended behavior)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.distributed` RemoteModule / RPC)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Deserialization RCE across the distributed RPC channel (vendor-disputed as intended behavior)","attack_vector":"Any host that can reach the RPC port of a multi-node training job — i.e. a co-tenant on the same fabric","remediation":"No patch — vendor position is that `torch.distributed` is a trusted-network protocol. Remediation is network isolation: per-tenant VRF/VLAN on the training fabric, never expose RPC ports across tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48063"],"status":"curated","fleet":{"ubiquity":"Very common - `torch.distributed` RPC / RemoteModule underpins multi-node training and RL rollout fanout on GPU clusters","remediation_pain":"Config/network change, not a patch - PyTorch disputes it, so remediation is isolating the RPC plane per tenant, which is an architectural change across the fleet","pain_class":"other","why_fleet_wide":"`RemoteModule` deserializes attacker-controlled data, so anyone who can reach a training job's RPC port executes code on every rank - one reachable worker owns the whole distributed job"}},{"id":"CVE-2024-4985","cve":"CVE-2024-4985","aliases":[],"title":"GitHub Enterprise Server: Forged SAML response with encrypted assertions enabled","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitHub Enterprise Server","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Forged SAML response with encrypted assertions enabled -> provision/gain site-admin access","attack_vector":"Network (remote)","remediation":"Control-plane: GHES upgrade immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4985"],"status":"curated"},{"id":"CVE-2024-53138","cve":"CVE-2024-53138","aliases":["net/mlx5e kTLS incorrect page refcounting"],"title":"Linux kernel mlx5_core kTLS TX offload: The kTLS TX path mixes get_page() and page_ref_inc() when acquiring references but only put_page() when…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core kTLS TX offload","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The kTLS TX path mixes get_page() and page_ref_inc() when acquiring references but only put_page() when releasing. With large folios the page is dereferenced too many times, producing a use-after-free. Found in the wild via sendfile() with zero-copy over NFS. Any node terminating TLS in hardware on ConnectX is exposed - storage gateways and object-store frontends in a GPU cloud typically are.","attack_vector":"Remote, unauthenticated, over a TLS connection served by mlx5 hardware kTLS TX offload with large-folio pages in play.","remediation":"Upgrade the host kernel to 6.12 or one of the many stable backports (5.4.287, 5.10.231, 5.15.174, 6.1.119, 6.6.63, 6.11.10) - the wide backport range means most distros already ship a fix. Rolling reboot. Interim: disable kTLS TX offload on mlx5 interfaces (ethtool -K <dev> tls-hw-tx-offload off), a live config change at a CPU cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53138","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-53138.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-54085","cve":"CVE-2024-54085","aliases":[],"title":"AMI MegaRAC SPx (Redfish Host Interface): **[KEV]** Unauthenticated auth bypass, full BMC takeover, malicious firmware flash. Persists below the OS and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (Redfish Host Interface)","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"**[KEV]** Unauthenticated auth bypass, full BMC takeover, malicious firmware flash. Persists below the OS and survives reimaging","attack_vector":"Network / Redfish host interface, unauthenticated","remediation":"Out-of-band BMC flash on every node; requires an ODM rebase of the AMI fix (Supermicro / Lenovo / HPE / ASRock each ship their own build, availability lags AMI by months), bricking risk on interrupted flash","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-54085"],"status":"curated","fleet":{"ubiquity":"universal - AMI MegaRAC is the OEM BMC stack shipped under most Supermicro/ASRock Rack/Gigabyte/Quanta GPU servers","remediation_pain":"firmware-flash - BMC firmware image per node, applied out-of-band, and the OEM must first rebase AMI's fix into its own build; realistically a rolling node-drain because a bad flash bricks the board","pain_class":"firmware-flash","why_fleet_wide":"Unauthenticated Redfish auth bypass by spoofing the X-Server-Addr/Host header gives full BMC takeover on every node running the same OEM image; the BMC sits below the hypervisor, so an implant survives OS reimaging and GPU node rebuilds. First BMC CVE ever added to CISA KEV (2025-06-25)."}},{"id":"CVE-2024-5452","cve":"CVE-2024-5452","aliases":[],"title":"PyTorch Lightning: RCE via deserialization of untrusted checkpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch Lightning","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via deserialization of untrusted checkpoint","attack_vector":"Customer-supplied `.ckpt` file","remediation":"Tenant-owned library. Provider action is scanning uploaded checkpoints at the object-store boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5452"],"status":"curated"},{"id":"CVE-2024-55591","cve":"CVE-2024-55591","aliases":[],"title":"Fortinet FortiOS/FortiProxy: Auth bypass via crafted Node.js websocket requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiOS/FortiProxy","year":"2024","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Auth bypass via crafted Node.js websocket requests -> super-admin privileges","attack_vector":"Network (remote)","remediation":"Control-plane: firmware; review all admin and local users","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-55591"],"status":"curated"},{"id":"CVE-2024-5660","cve":"CVE-2024-5660","aliases":["TFV-12","Hardware Page Aggregation bypass","HPA erratum"],"title":"Arm Neoverse V1 / V2 / V3 / V3AE / N2 and Cortex-A77/A78/A710/X1-X925 cores with Hardware Page Aggregation enabled; Trusted Firmware-A v2.2-v2.12 and LTS 2.8/2.10 branches: A guest can defeat…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Neoverse V1 / V2 / V3 / V3AE / N2 and Cortex-A77/A78/A710/X1-X925 cores with Hardware Page Aggregation enabled…","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A guest can defeat Stage-2 translation, meaning the hypervisor's memory isolation for that VM stops holding. On a multi-tenant Arm host this is a tenant-to-hypervisor escape, and from there a read of every other VM resident on the socket. It also breaks Granule Protection Table checks, so if you are relying on Arm CCA / Realm confidentiality to sell 'the operator cannot see your model', that guarantee is void on affected silicon. Neoverse V2 is the Grace core, V1/V2 are Graviton3/Graviton4, N2 ships in several Arm server parts, so this touches most of the Arm side of an AI fleet.","attack_vector":"Unprivileged or kernel-level code inside any guest VM on an affected Arm host, with no special device access. No physical access, no network reachability, nothing but a VM on the box.","remediation":"This is silicon errata, so the fix is firmware discipline, not a code patch: EL3 firmware must set CPUECTLR_EL1[46]=1 to disable hardware page aggregation on every affected core. That means an updated TF-A / BL31 build from the platform OEM, flashed to each node, with a reboot and therefore a job drain. Ampere, NVIDIA and cloud OEMs each ship their own TF-A fork, so you wait on your board vendor, not on Arm. Budget a rolling reboot of the whole Arm fleet. There is a small performance cost to losing HPA on large-page workloads; measure before assuming it is free.","references":["https://developer.arm.com/Arm%20Security%20Center/Arm%20CPU%20Vulnerability%20CVE-2024-5660","https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-12.html","https://nvd.nist.gov/vuln/detail/CVE-2024-5660"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56656","cve":"CVE-2024-56656","aliases":[],"title":"Linux bnxt_en driver (5760X / P7 aggregation ID mask): The bnxt_en driver mishandles the aggregation ID mask on 5760X (P7) chips, oopsing the kernel. P7 is the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (5760X / P7 aggregation ID mask)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The bnxt_en driver mishandles the aggregation ID mask on 5760X (P7) chips, oopsing the kernel. P7 is the Thor2 generation — the 400G-class Broadcom NIC going into current AI server designs — so this specifically hits new GPU-node builds rather than legacy hardware. A kernel oops on a training node kills the job for every rank in the collective, not just that node.","attack_vector":"Triggered by received traffic patterns hitting the hardware GRO/LRO path. Effectively remote from anything that can send traffic to the node.","remediation":"Kernel or driver package upgrade plus a host reboot. Standard rolling-reboot across the fleet, but coordinate with running jobs — a mid-training reboot is more expensive than the patch. If you build your own kernels, this is a backport-and-rebuild.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56656"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-6800","cve":"CVE-2024-6800","aliases":[],"title":"GitHub Enterprise Server: XML signature wrapping with publicly exposed federation metadata","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitHub Enterprise Server","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"XML signature wrapping with publicly exposed federation metadata -> forge a site-admin session","attack_vector":"Network (remote)","remediation":"Control-plane: GHES upgrade; review the admin list","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6800"],"status":"curated"},{"id":"CVE-2024-9053","cve":"CVE-2024-9053","aliases":[],"title":"vLLM (`AsyncEngineRPCServer`): Unsafe deserialization on RPC entrypoints","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`AsyncEngineRPCServer`)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unsafe deserialization on RPC entrypoints → RCE","attack_vector":"Unauthenticated network to the RPC server port","remediation":"Upgrade past 0.6.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9053"],"status":"curated"},{"id":"CVE-2024-9070","cve":"CVE-2024-9070","aliases":[],"title":"BentoML (runner server): Deserialization RCE on the internal runner server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (runner server)","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Deserialization RCE on the internal runner server","attack_vector":"Network to the runner port — reachable by a co-tenant in a flat cluster network","remediation":"Upgrade past 1.3.4.post1; runner ports must be namespace-local","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9070"],"status":"curated"},{"id":"CVE-2024-9486","cve":"CVE-2024-9486","aliases":[],"title":"Kubernetes Image Builder: VM images built with the Proxmox provider ship default credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes Image Builder","year":"2024","cvss_score":9.8,"severity":"critical","kev":false,"impact":"VM images built with the Proxmox provider ship default credentials; node takeover from the network","attack_vector":"Unauthenticated network reaching an affected node","remediation":"Rebuild and redeploy every affected node image; this is a full fleet reimage, not a package update","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-11200","cve":"CVE-2025-11200","aliases":[],"title":"MLflow (auth): Weak password requirements","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (auth)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Weak password requirements → authentication bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade; do not use MLflow's own auth as the tenant boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-11200"],"status":"curated"},{"id":"CVE-2025-11201","cve":"CVE-2025-11201","aliases":[],"title":"MLflow (model creation): Directory traversal on model creation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (model creation)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Directory traversal on model creation → RCE","attack_vector":"Network to the tracking server","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-11201"],"status":"curated"},{"id":"CVE-2025-15379","cve":"CVE-2025-15379","aliases":[],"title":"MLflow (serving container init): Command injection in `_install_model_dependencies`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (serving container init)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Command injection in `_install_model_dependencies`","attack_vector":"Customer-supplied model requirements consumed when the provider builds a serving container","remediation":"Provider-owned in managed serving: the tenant's requirements file executes shell in the build","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-15379"],"status":"curated"},{"id":"CVE-2025-1550","cve":"CVE-2025-1550","aliases":[],"title":"Keras (`Model.load_model`): Arbitrary code execution from a crafted `.keras` archive even with `safe_mode=True`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`Model.load_model`)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution from a crafted `.keras` archive even with `safe_mode=True`","attack_vector":"Customer-supplied Keras model file","remediation":"Upgrade Keras in base images; no host-level patch. The safe-mode flag is not a security boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1550"],"status":"curated","fleet":{"ubiquity":"Very common - Keras 3.0.0-3.8.0 ships in most TF/Keras training images","remediation_pain":"**Image rebuild** across the fleet to Keras 3.9.0+; no host restart, but every image and every cached model must be re-vetted","pain_class":"other","why_fleet_wide":"`Model.load_model` executes arbitrary Python from a crafted `.keras` config *even with `safe_mode=True`*, so any model artifact pulled from a hub or a customer bucket is RCE on the loading GPU node"}},{"id":"CVE-2025-1793","cve":"CVE-2025-1793","aliases":[],"title":"LlamaIndex (vector store integrations): SQL injection across multiple vector store integrations","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LlamaIndex (vector store integrations)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"SQL injection across multiple vector store integrations","attack_vector":"Attacker-influenced query/metadata reaching the store","remediation":"Upgrade past 0.12.21","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1793"],"status":"curated"},{"id":"CVE-2025-1945","cve":"CVE-2025-1945","aliases":[],"title":"picklescan (model scanner): Scanner fails to detect malicious pickles when ZIP flag bits are flipped","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan (model scanner)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Scanner fails to detect malicious pickles when ZIP flag bits are flipped","attack_vector":"Customer-supplied PyTorch archive submitted to a provider's \"we scan your models\" control","remediation":"Direct hit on the provider's own control plane. Upgrade picklescan and stop treating a clean scan as proof of safety","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1945"],"status":"curated"},{"id":"CVE-2025-1974","cve":"CVE-2025-1974","aliases":[],"title":"ingress-nginx: \"IngressNightmare\": unauthenticated RCE in the admission controller, reachable from any pod, giving…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"\"IngressNightmare\": unauthenticated RCE in the admission controller, reachable from any pod, giving cluster-wide secret access and full cluster takeover","attack_vector":"Any pod on the cluster network, no credentials needed","remediation":"Emergency controller upgrade, or delete the ValidatingWebhookConfiguration and network-restrict the admission port. Rolling controller upgrade, no GPU drain. The highest-priority item in this table for a multi-tenant neocloud","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"ubiquity":"Very common - ingress-nginx is the default ingress on most managed and self-run K8s clusters, including the control planes neoclouds put in front of GPU tenants","remediation_pain":"`daemon-restart` (rolling deployment upgrade of the controller) - cheap on the ingress pods themselves, but the *cleanup* is the pain: every secret in every namespace must be assumed stolen and rotated","pain_class":"daemon-restart","why_fleet_wide":"Anything on the pod network - i.e. any tenant workload - can hit the validating admission controller, load a shared library into the controller pod and read all secrets in all namespaces, which is cluster takeover"}},{"id":"CVE-2025-21927","cve":"CVE-2025-21927","aliases":["nvme-tcp header digest memory corruption","malicious NVMe/TCP target attacks initiator"],"title":"Linux kernel - NVMe/TCP host (initiator), drivers/nvme/host/tcp.c: TENANT ISOLATION: nvme_tcp_recv_pdu() did not validate the PDU header length, so with header digests enabled…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe/TCP host (initiator), drivers/nvme/host/tcp.c","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: nvme_tcp_recv_pdu() did not validate the PDU header length, so with header digests enabled a target can send a packet declaring an invalid header length (for example 255) and make nvme_tcp_verify_hdgst() access memory outside the allocation and overwrite it with the calculated digest. The attack direction is target-to-initiator: a compromised or rogue storage target corrupts kernel memory on every GPU compute node that mounts from it. One storage compromise becomes fleet-wide kernel compromise, and header digests - a data-integrity feature operators turn on deliberately - are the precondition.","attack_vector":"The attacker controls an NVMe/TCP target the victim connects to, or can spoof/inject into that TCP connection, and returns a PDU with an out-of-range header length. Because NVMe/TCP has no transport authentication by default, an on-path attacker or anyone who can win a race to the discovery address can pose as the target.","remediation":"Host reboot / kernel upgrade on all NVMe/TCP initiator nodes - that is the GPU compute fleet, not just storage, so plan a full rolling drain. Immediate mitigations: disable header digests on affected initiators (nvme connect option, applied on reconnect, no reboot) to remove the precondition, and enable TLS for NVMe/TCP where supported so the target's identity is proven. Also verify that discovery addresses cannot be hijacked on the storage network.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2025/CVE-2025-21927.json","https://nvd.nist.gov/vuln/detail/CVE-2025-21927"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23310","cve":"CVE-2025-23310","aliases":[],"title":"NVIDIA Triton: Stack buffer overflow","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack buffer overflow → code execution","attack_vector":"Unauthenticated network to the inference port","remediation":"Patch; part of the Aug-2025 Triton cluster disclosed by Wiz","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23310"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2025-23311","cve":"CVE-2025-23311","aliases":[],"title":"NVIDIA Triton: Stack overflow via crafted request","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Stack overflow via crafted request → code execution","attack_vector":"Unauthenticated network","remediation":"Patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23311"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2025-23316","cve":"CVE-2025-23316","aliases":[],"title":"NVIDIA Triton Inference Server (Python backend): Attacker-controlled input in the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton Inference Server (Python backend)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Attacker-controlled input in the Python backend → remote code execution","attack_vector":"Unauthenticated network to an exposed Triton HTTP/gRPC port","remediation":"Patch to the fixed Triton release and rebuild NGC-derived images. Provider-published Triton images must be rebuilt, not just re-tagged","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23316"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2025-25257","cve":"CVE-2025-25257","aliases":[],"title":"Fortinet FortiWeb: Unauthenticated SQL injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiWeb","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Unauthenticated SQL injection -> pre-auth code execution on the WAF","attack_vector":"Network (remote)","remediation":"Control-plane: firmware upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25257"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-25291","cve":"CVE-2025-25291","aliases":[],"title":"GitLab (ruby-saml): ReXML/Nokogiri parser differential","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab (ruby-saml)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"ReXML/Nokogiri parser differential -> SAML authentication bypass and account takeover","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade GitLab/ruby-saml; rotate the IdP signing certificate","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25291"],"status":"curated"},{"id":"CVE-2025-27520","cve":"CVE-2025-27520","aliases":[],"title":"BentoML: RCE via insecure deserialization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Network to the serving endpoint","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-27520"],"status":"curated"},{"id":"CVE-2025-32375","cve":"CVE-2025-32375","aliases":[],"title":"BentoML: Insecure deserialization RCE prior to 1.4.8","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Insecure deserialization RCE prior to 1.4.8","attack_vector":"Network to the serving endpoint","remediation":"Upgrade to 1.4.8+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32375"],"status":"curated"},{"id":"CVE-2025-32434","cve":"CVE-2025-32434","aliases":[],"title":"PyTorch (`torch.load`): RCE via unsafe deserialization even with `weights_only=True`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.load`)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via unsafe deserialization even with `weights_only=True`","attack_vector":"Customer-supplied `.pt`/`.pth` checkpoint file loaded by any tenant job or by a provider-run model-import service","remediation":"No host patch. Bump `torch>=2.6.0` in every base image the provider ships; if the tenant pins an old torch in their own image, the provider cannot remediate — advise migration to safetensors. Shared responsibility: provider owns base images, tenant owns pinned envs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32434"],"status":"curated","fleet":{"ubiquity":"Universal - PyTorch ≤2.5.1 is in essentially every AI container image in every GPU fleet","remediation_pain":"**Image rebuild across the fleet** (`node-drain` for anything long-running) - the fix is `torch>=2.6.0` inside every tenant and platform image; you cannot hot-patch a library already imported into a running training job","pain_class":"node-drain","why_fleet_wide":"`torch.load(weights_only=True)` - the setting everyone was told was the safe one - is bypassable, so *every* checkpoint-loading path in the fleet is an RCE sink; 130+ public PoCs"}},{"id":"CVE-2025-33222","cve":"CVE-2025-33222","aliases":[],"title":"NVIDIA Isaac Launchable: Hard-coded credentials in Isaac Launchable give an unauthenticated network attacker code execution and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Launchable","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Hard-coded credentials in Isaac Launchable give an unauthenticated network attacker code execution and privilege escalation - scored 9.8. Hard-coded credentials are unpatchable by configuration: the fix has to be a new build, and any deployment still running the old artifact stays exploitable regardless of what you change around it.","attack_vector":"Network, unauthenticated, no user interaction. The credentials are in the artifact, so anyone with the artifact has them.","remediation":"Update to the fixed Isaac Launchable release per bulletin 5749 and rotate anything the embedded credential could reach. Cost: redeploy. Critically, patching alone is insufficient - assume the credential is public and revoke it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33222","https://github.com/NVIDIA/product-security/tree/main/2025/5749"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-798"]},{"id":"CVE-2025-33223","cve":"CVE-2025-33223","aliases":[],"title":"NVIDIA Isaac Launchable: Execution with unnecessary privileges lets an unauthenticated network attacker reach code execution…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Launchable","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Execution with unnecessary privileges lets an unauthenticated network attacker reach code execution, privilege escalation, information disclosure and data tampering. Scored 9.8.","attack_vector":"Network, unauthenticated, no user interaction.","remediation":"Update to the fixed release per bulletin 5749 and redeploy. Cost: redeploy only, but review what privileges the workload actually needs - the underlying pattern is over-privileged execution, which recurs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33223","https://github.com/NVIDIA/product-security/tree/main/2025/5749"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"]},{"id":"CVE-2025-33224","cve":"CVE-2025-33224","aliases":[],"title":"NVIDIA Isaac Launchable: A second over-privileged execution path with the same unauthenticated network reach and 9.8 score","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Launchable","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A second over-privileged execution path with the same unauthenticated network reach and 9.8 score.","attack_vector":"Network, unauthenticated, no user interaction.","remediation":"Update to the fixed release per bulletin 5749 and redeploy. Cost: redeploy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33224","https://github.com/NVIDIA/product-security/tree/main/2025/5749"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"]},{"id":"CVE-2025-38246","cve":"CVE-2025-38246","aliases":[],"title":"Linux bnxt_en driver (XDP redirect list flush): List corruption in the XDP redirect path, found crashing production systems. Same family as the other bnxt…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (XDP redirect list flush)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"List corruption in the XDP redirect path, found crashing production systems. Same family as the other bnxt XDP defects: if you run XDP for packet steering in front of inference endpoints, the Broadcom driver's redirect path has repeatedly been the weak spot, and the failure mode is a host crash rather than a graceful drop.","attack_vector":"Traffic through an attached XDP program using redirect on a Broadcom NIC.","remediation":"Kernel/driver upgrade plus host reboot. Companion fix CVE-2025-38439 (wrong DMA unmap length on XDP_REDIRECT) is in the same area — take both. Interim: detach XDP from Broadcom interfaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38246","https://nvd.nist.gov/vuln/detail/CVE-2025-38439"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40350","cve":"CVE-2025-40350","aliases":[],"title":"Linux kernel mlx5_core RX datapath (striding RQ + XDP multi-buffer): The mlx5 driver assumed an XDP program could not change the xdp_buff layout. It can, via…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (striding RQ + XDP multi-buffer)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The mlx5 driver assumed an XDP program could not change the xdp_buff layout. It can, via bpf_xdp_adjust_head/tail - and when it shrinks non-linear data the driver hits a BUG_ON, or builds a malformed skb. Any node running an XDP program on mlx5 (load balancers, DDoS scrubbers, CNI dataplanes like Cilium) can be crashed or memory-corrupted by remote packets. Jumbo-MTU routed fabrics are standard in datacenters, so the attacker does not need to be L2-adjacent.","attack_vector":"Unauthenticated remote sender, provided the target node runs an XDP program on an mlx5 interface with striding RQ.","remediation":"Upgrade the host kernel to 6.18 or a stable backport (6.6.115, 6.12.56, 6.17.6). Rolling reboot. Interim: unload the XDP program from mlx5 interfaces if you can accept the performance/feature loss - a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40350","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-40350.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-46412","cve":"CVE-2025-46412","aliases":["CVE-2025-41426","ICSA-25-140-10"],"title":"Vertiv Liebert RDU101 (<=1.9.0.0) and Liebert IS-UNITY (<=8.4.1.0) communication cards: Authentication bypass plus a stack-based buffer overflow giving code execution on the card that sits inside…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Vertiv Liebert RDU101 (<=1.9.0.0) and Liebert IS-UNITY (<=8.4.1.0) communication cards","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Authentication bypass plus a stack-based buffer overflow giving code execution on the card that sits inside Liebert cooling and power equipment across most of the datacenter industry. This is the worst shape a facility bug can take: an unauthenticated attacker on the facility network gets a persistent foothold on a device physically wired into thermal control, and from there can command the attached unit, spoof telemetry upward to the DCIM/BMS so dashboards stay green, and pivot laterally to every other card on the same VLAN. Applied against the cooling serving a GPU hall this is a fleet-wide availability kill and a hardware-damage risk, not an IT incident. Because the card also relays alarms, the same access lets an attacker blind the operator during the event. For a multi-tenant operator it is also a tenant-handoff problem: a card compromised under one tenant's occupancy stays compromised across the next tenant, because nobody re-flashes communication cards between customers.","attack_vector":"Unauthenticated, network-reachable, low complexity - CISA rates it exploitable remotely. The realistic exposure is the facility/BMS VLAN. RDU101 and IS-UNITY cards are also routinely reachable from the DCIM collector and, in far too many sites, from a vendor remote-support VPN concentrator that the mechanical contractor owns. Shodan-class internet exposure of Liebert cards is a recurring finding, so check whether yours are NAT'd out for 'remote monitoring'.","remediation":"Patchable, and this one is worth the window: update RDU101 to v1.9.1.2_0000001 and IS-UNITY to v8.4.3.1_00160. The card firmware update does not require draining the cooling unit, but it does require the card to be reachable and briefly offline, so schedule it as a monitoring outage rather than a cooling outage - budget a per-card touch across the whole fleet, which for a large hall is dozens of cards and a real technician-day cost. Alongside the patch, segment: cards on a dedicated VLAN, no route to tenant or corporate networks, no inbound from the internet, and inventory every Liebert card by firmware version so you can prove the fleet is clean. Leased colo: this is the landlord's card in the landlord's unit - send them the CISA advisory ID and require written confirmation of the firmware level.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-25-140-10","https://nvd.nist.gov/vuln/detail/CVE-2025-46412","https://nvd.nist.gov/vuln/detail/CVE-2025-41426","https://www.vertiv.com/en-us/support/security-support-center/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-47277","cve":"CVE-2025-47277","aliases":[],"title":"vLLM (`PyNcclPipe` KV transfer): RCE via the KV cache transfer integration","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`PyNcclPipe` KV transfer)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via the KV cache transfer integration","attack_vector":"Unauthenticated network to the KV pipe port between prefill/decode nodes","remediation":"Upgrade past 0.8.4. Disaggregated-prefill deployments expose a new unauthenticated plane on the tenant fabric","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-47277"],"status":"curated"},{"id":"CVE-2025-49655","cve":"CVE-2025-49655","aliases":[],"title":"Keras: Deserialization of untrusted data in 3.11.0–3.11.2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Deserialization of untrusted data in 3.11.0–3.11.2 → malicious model runs arbitrary code","attack_vector":"Customer-supplied model file","remediation":"Upgrade to 3.11.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49655"],"status":"curated"},{"id":"CVE-2025-49825","cve":"CVE-2025-49825","aliases":[],"title":"Teleport: Remote authentication bypass in Teleport Community Edition (<=17.5.1)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Teleport","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Remote authentication bypass in Teleport Community Edition (<=17.5.1)","attack_vector":"Network (remote)","remediation":"Control-plane + data-plane: upgrade the proxy AND every node agent; rotate CAs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49825"],"status":"curated"},{"id":"CVE-2025-53521","cve":"CVE-2025-53521","aliases":["K000156741"],"title":"F5 BIG-IP (APM access policy): Specific malicious traffic against a virtual server with a BIG-IP APM access policy configured leads to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"F5 BIG-IP (APM access policy)","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Specific malicious traffic against a virtual server with a BIG-IP APM access policy configured leads to remote code execution. This is part of F5's October 2025 mass-remediation advisory issued after the nation-state breach of F5's internal systems and BIG-IP source code, and it's confirmed in CISA's KEV catalog — treat any internet- or cluster-VPN-facing BIG-IP APM instance as a priority target.","attack_vector":"Remote, against a virtual server that has an APM access policy applied — F5 has not disclosed full exploitation prerequisites, consistent with a critical RCE that's actively tracked as exploited.","remediation":"Software upgrade to the fixed BIG-IP version per F5 K000156741, then reboot/failover. Because this is part of the breach-remediation batch, treat the whole BIG-IP estate as needing review — not just this one CVE — and prioritize any unit that's End of Technical Support, since F5 does not evaluate or patch those.","references":["https://my.f5.com/manage/s/article/K000156741","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2025-53521"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-62863","cve":"CVE-2025-62863","aliases":["AMP-SB-0007","AmpereOne UEFI-MM PCIe driver OOB write"],"title":"Ampere AmpereOne AC03 before 3.5.9.3, AC04 before 4.4.5.2, AmpereOne M before 5.4.5.1 - UEFI Management Mode PCIe driver reached over SMC (sibling issues CVE-2025-62862 and CVE-2025-62864 in the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere AmpereOne AC03 before 3.5.9.3, AC04 before 4.4.5.2, AmpereOne M before 5.4.5.1 - UEFI Management Mode PCIe…","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A malformed SMC from the normal world produces an out-of-bounds write inside the S-EL0 UEFI-MM secure partition. That is code execution in the secure world reached from the OS - the boundary AmpereOne uses to protect runtime firmware services. Once there, the attacker can rewrite firmware state, forge attestation, or persist across a tenant handoff. The bulletin's siblings give a secure-partition-context OOB write (CVE-2025-62864) and an S-EL0 info leak (CVE-2025-62862), so the whole SMC surface should be treated as compromised until patched.","attack_vector":"Host kernel or hypervisor issuing SMC calls on an AmpereOne node. On bare-metal AmpereOne rental this is the tenant. No physical or network access needed.","remediation":"Update to the fixed firmware for your part - AC03 3.5.9.3, AC04 4.4.5.2, AmpereOne M 5.4.5.1 - from the board OEM. Flash + reboot + drain per node; because this is in UEFI-MM, it ships as part of the platform firmware bundle and the ODM must integrate Ampere's release before you can install it. Verify the running version after reboot; AmpereOne firmware version strings are per-SKU and easy to get wrong on a mixed fleet.","references":["https://amperecomputing.com/products/security-bulletins/amp-sb-0007","https://nvd.nist.gov/vuln/detail/CVE-2025-62863","https://amperecomputing.com/products/product-security"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-63389","cve":"CVE-2025-63389","aliases":[],"title":"Ollama (API auth): Critical authentication bypass on API endpoints through v0.12.3","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (API auth)","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Critical authentication bypass on API endpoints through v0.12.3","attack_vector":"Unauthenticated network to the Ollama API","remediation":"Upgrade; assume any tenant-reachable Ollama instance is fully controllable","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-63389"],"status":"curated"},{"id":"CVE-2025-6543","cve":"CVE-2025-6543","aliases":["CTX694788"],"title":"Citrix NetScaler ADC / Gateway (configured as VPN Gateway, ICA Proxy, CVPN, RDP Proxy, or AAA virtual server): A memory overflow in the Gateway/AAA code path corrupts control flow, letting an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Citrix NetScaler ADC / Gateway (configured as VPN Gateway, ICA Proxy, CVPN, RDP Proxy, or AAA virtual server)","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"A memory overflow in the Gateway/AAA code path corrupts control flow, letting an attacker crash the appliance (denial of service) and, per CISA's KEV listing, this has been exploited in the wild — meaning it's already a live threat to any NetScaler used as the VPN entry point into a cluster's management network.","attack_vector":"Remote, targeting a NetScaler configured as a Gateway/AAA virtual server (i.e. the VPN/ICA-proxy entry point) — exact prerequisites are undisclosed by Citrix, consistent with an actively-exploited memory-corruption bug.","remediation":"Software upgrade to the fixed NetScaler ADC/Gateway build per Citrix advisory CTX694788, then reboot. Given confirmed in-the-wild exploitation, patch ahead of the normal cycle — this is the box guarding VPN access into the cluster's OOB/management network, not just a data-plane load balancer.","references":["https://support.citrix.com/support-home/kbsearch/article?articleNumber=CTX694788","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2025-6543"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-66405","cve":"CVE-2025-66405","aliases":[],"title":"Portkey AI Gateway: Gateway resolves the destination baseURL from attacker-controlled precedence","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Portkey AI Gateway","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Gateway resolves the destination baseURL from attacker-controlled precedence → request hijack","attack_vector":"Tenant request to the shared gateway","remediation":"Upgrade to 1.14.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-66405"],"status":"curated"},{"id":"CVE-2025-67038","cve":"CVE-2025-67038","aliases":["ICSA-26-069-02"],"title":"Lantronix EDS5000 serial-to-Ethernet device server: Root command execution on the device server. The HTTP RPC module writes a log line whenever a login fails…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix EDS5000 serial-to-Ethernet device server","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"Root command execution on the device server. The HTTP RPC module writes a log line whenever a login fails, and it builds that log line by shell-concatenating the attempted username — so a failed login with a malicious username is enough to run arbitrary commands as root. This CVE is in CISA's Known Exploited Vulnerabilities catalog, meaning it has been used in real attacks.","attack_vector":"No valid credentials needed — the trigger is a failed authentication attempt, so the attacker just needs network reachability to the device's HTTP interface and controls the username field.","remediation":"Firmware flash is required; there's no config toggle to disable the vulnerable logging path. Given the confirmed exploitation, treat this as urgent — patch every EDS5000 in the fleet ahead of routine maintenance windows, one device at a time (each flash drops the serial sessions it's bridging).","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-069-02","https://www.cisa.gov/known-exploited-vulnerabilities-catalog?field_cve=CVE-2025-67038"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-6802","cve":"CVE-2025-6802","aliases":["CVE-2025-6793","CVE-2025-6794","CVE-2025-6795","CVE-2025-6796","CVE-2025-6797","CVE-2025-6798","CVE-2025-6799","CVE-2025-6800","CVE-2025-6801","CVE-2025-6803","CVE-2025-6804","CVE-2025-6805","CVE-2025-6806","CVE-2025-6807","CVE-2025-8426"],"title":"Marvell QConvergeConsole (QLogic Fibre Channel / FC-NVMe / CNA HBA management web console), 5.5.0.78 and earlier: A 2025 batch of sixteen ZDI advisories against the same Tomcat/GWT console, most…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Marvell QConvergeConsole (QLogic Fibre Channel / FC-NVMe / CNA HBA management web console), 5.5.0.78 and earlier","year":"2025","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A 2025 batch of sixteen ZDI advisories against the same Tomcat/GWT console, most of them reachable with no authentication at all. getFileFromURL gives unrestricted file upload leading to code execution as the console's service account (SYSTEM or root); saveAsText gives directory-traversal RCE; the rest are arbitrary file write, arbitrary file deletion and arbitrary file read. QConvergeConsole is the box that flashes QLogic HBA firmware and boot-from-SAN parameters across the fleet, so owning it means owning the HBA firmware and the SAN login identity of every server it manages - a persistence layer beneath the tenant OS, and a way to re-point a node's boot LUN. The file-deletion variants alone are a fleet availability event.","attack_vector":"Any host that can reach the QConvergeConsole web listener (typically TCP 8080/8443 on a management server or on individual hosts running the agent). No credentials for the unauthenticated set; the older 2020 cluster's 'authenticated' variants were shown to be reachable by bypassing the console's own auth.","remediation":"Upgrade to the fixed Marvell QConvergeConsole release, but treat that as a stopgap: this codebase has now produced two large unauthenticated-RCE clusters (2020 and 2025) and has no meaningful hardening. The durable fix for a GPU fleet is to remove QConvergeConsole from production hosts entirely and manage QLogic HBAs with the qaucli/qlogic CLI driven from configuration management, then firewall the console port. Software-only change, no HBA flash and no array downtime. Any host whose console was internet- or tenant-reachable should have its HBA firmware reflashed from a known-good vendor image before reuse.","references":["https://www.zerodayinitiative.com/advisories/ZDI-25-464/","https://www.zerodayinitiative.com/advisories/ZDI-25-450/","https://nvd.nist.gov/vuln/detail/CVE-2025-6802"],"status":"curated"},{"id":"CVE-2025-7775","cve":"CVE-2025-7775","aliases":[],"title":"Citrix NetScaler: Memory overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler","year":"2025","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Memory overflow -> pre-authentication remote code execution and/or DoS","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7775"],"status":"curated"},{"id":"CVE-2026-0300","cve":"CVE-2026-0300","aliases":[],"title":"Palo Alto PAN-OS: Buffer overflow in the User-ID Captive Portal","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Buffer overflow in the User-ID Captive Portal -> unauthenticated arbitrary code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency patch; disable the Captive Portal if unused","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-0300"],"status":"curated"},{"id":"CVE-2026-0545","cve":"CVE-2026-0545","aliases":[],"title":"MLflow (jobs API): `/ajax-api/3.0/jobs/*` unauthenticated even with basic-auth enabled","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (jobs API)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`/ajax-api/3.0/jobs/*` unauthenticated even with basic-auth enabled","attack_vector":"Unauthenticated network","remediation":"Upgrade; auth coverage gaps recur across MLflow releases","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-0545"],"status":"curated"},{"id":"CVE-2026-12481","cve":"CVE-2026-12481","aliases":[],"title":"Keras (Lambda layer): Arbitrary code execution via Lambda-layer deserialization in 3.14.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (Lambda layer)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Arbitrary code execution via Lambda-layer deserialization in 3.14.0","attack_vector":"Customer-supplied model containing a Lambda layer","remediation":"No safe fix short of refusing Lambda layers. Provider should reject models declaring Lambda layers at ingest","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-12481"],"status":"curated"},{"id":"CVE-2026-1281","cve":"CVE-2026-1281","aliases":[],"title":"Ivanti Endpoint Manager Mobile: Code injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager Mobile","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Code injection -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately; the MDM plane reaches operator devices","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1281"],"status":"curated"},{"id":"CVE-2026-1340","cve":"CVE-2026-1340","aliases":[],"title":"Ivanti Endpoint Manager Mobile: Code injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager Mobile","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Code injection -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch immediately; assume compromise if internet-facing","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1340"],"status":"curated"},{"id":"CVE-2026-15969","cve":"CVE-2026-15969","aliases":[],"title":"SGLang (`/load_lora_adapter_from_tensors`): Unauthenticated RCE bypassing `SafeUnpickler`'s incomplete denylist","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/load_lora_adapter_from_tensors`)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE bypassing `SafeUnpickler`'s incomplete denylist","attack_vector":"Customer-supplied LoRA adapter posted to the serving API","remediation":"Upgrade. LoRA adapters are executable content — a denylist unpickler is not a boundary","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15969"],"status":"curated"},{"id":"CVE-2026-21643","cve":"CVE-2026-21643","aliases":[],"title":"Fortinet FortiClient EMS: Unauthenticated SQL injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiClient EMS","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Unauthenticated SQL injection -> code execution via crafted HTTP requests","attack_vector":"Network (remote)","remediation":"Control-plane: patch the endpoint-management server","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-21643"],"status":"curated"},{"id":"CVE-2026-22778","cve":"CVE-2026-22778","aliases":[],"title":"vLLM (image error echo): Error path returns sensitive content on invalid image input","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (image error echo)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Error path returns sensitive content on invalid image input","attack_vector":"Unauthenticated network to the multimodal endpoint","remediation":"Upgrade past 0.14.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22778"],"status":"curated"},{"id":"CVE-2026-23112","cve":"CVE-2026-23112","aliases":["CVE-2026-52989","CVE-2022-50717"],"title":"Linux kernel nvmet-tcp - PDU iovec construction and H2C Transfer Tag handling: nvmet_tcp_build_pdu_iovec() walks past the command's scatterlist when a PDU length or offset exceeds sg_cnt…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel nvmet-tcp - PDU iovec construction and H2C Transfer Tag handling","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"nvmet_tcp_build_pdu_iovec() walks past the command's scatterlist when a PDU length or offset exceeds sg_cnt and then builds a bvec from garbage sg->length/offset values, and in the related defect the error is not propagated so the socket receive loop reads network data into an uninitialised iterator. The older sibling is the same shape: the Transfer Tag in an H2C_DATA PDU is used as an array index with no bounds check. All three give an initiator on the storage network control over where the target kernel copies received data - memory corruption in the process that owns every tenant's namespace mappings, and a straightforward crash of the entire target if you only want the outage.","attack_vector":"Any host that can complete an NVMe/TCP connection to the target on TCP 4420 and send a malformed PDU. A provisioned namespace is not required for the connection-level PDU handling; on a fabric with no in-band auth, any host on the storage VLAN qualifies.","remediation":"Kernel update on target nodes, reboot, arrays offline for the duration unless you can fail initiators to a peer target first. There is no runtime workaround - the parsing happens before any policy check. Between now and the maintenance window, the storage VLAN ACL is the control: only known initiator IPs reach 4420. Note the 2022 Transfer Tag bounds check applies to long-lived kernels many operators are still running under distro LTS, so check your actual kernel rather than assuming a recent distro release covers it.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2026-23112","https://nvd.nist.gov/vuln/detail/CVE-2022-50717"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24178","cve":"CVE-2026-24178","aliases":[],"title":"NVIDIA FLARE SDK: Unauthenticated remote code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA FLARE SDK","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution","attack_vector":"Network-adjacent unauthenticated","remediation":"Emergency upgrade of FLARE; redeploy federated-learning services","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24178","https://github.com/NVIDIA/product-security/tree/main/2026/5819"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-639"]},{"id":"CVE-2026-24207","cve":"CVE-2026-24207","aliases":[],"title":"Triton Inference Server: Missing authentication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Missing authentication -> full unauthorized control of the server","attack_vector":"Network-adjacent unauthenticated","remediation":"Emergency Triton upgrade + redeploy all serving images; gate the endpoint behind authn immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24207","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-288"]},{"id":"CVE-2026-24254","cve":"CVE-2026-24254","aliases":[],"title":"NVIDIA Dynamo: Unauthenticated remote code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated remote code execution","attack_vector":"Network-adjacent unauthenticated","remediation":"Emergency Dynamo upgrade; redeploy; gate endpoints behind authn and network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24254","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-288"]},{"id":"CVE-2026-24270","cve":"CVE-2026-24270","aliases":[],"title":"NVIDIA AIStore: Missing authentication on API endpoints","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA AIStore","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Missing authentication on API endpoints -> full object-store access","attack_vector":"Network-adjacent unauthenticated inside the cluster","remediation":"Emergency upgrade of AIStore; enforce authn + network policy; audit stored tenant data","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24270","https://github.com/NVIDIA/product-security/tree/main/2026/5849"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-290"]},{"id":"CVE-2026-24858","cve":"CVE-2026-24858","aliases":[],"title":"Fortinet (FortiOS/FortiManager/FortiProxy): Auth bypass via alternate path using a FortiCloud account and a registered device","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet (FortiOS/FortiManager/FortiProxy)","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Auth bypass via alternate path using a FortiCloud account and a registered device","attack_vector":"Network (remote)","remediation":"Control-plane: fleet-wide firmware; audit FortiCloud account linkage","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24858"],"status":"curated"},{"id":"CVE-2026-25089","cve":"CVE-2026-25089","aliases":[],"title":"Fortinet FortiSandbox: OS command injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiSandbox","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] OS command injection -> unauthenticated command execution via crafted requests","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-25089"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-26190","cve":"CVE-2026-26190","aliases":[],"title":"Milvus (port 9091): Management port 9091 exposed by default enabling compromise","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Milvus (port 9091)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Management port 9091 exposed by default enabling compromise","attack_vector":"Unauthenticated network — default config","remediation":"Upgrade past 2.5.27/2.6.10 and firewall 9091. Insecure default, so an unmodified deployment is exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-26190"],"status":"curated"},{"id":"CVE-2026-3055","cve":"CVE-2026-3055","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Insufficient input validation as SAML IdP","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Insufficient input validation as SAML IdP -> out-of-bounds read / memory disclosure","attack_vector":"Network (remote)","remediation":"Control-plane: patch; rotate SAML signing material and kill sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3055"],"status":"curated"},{"id":"CVE-2026-3059","cve":"CVE-2026-3059","aliases":[],"title":"SGLang (multimodal ZMQ broker): Unauthenticated RCE via `pickle.loads()` on the ZMQ broker","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (multimodal ZMQ broker)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE via `pickle.loads()` on the ZMQ broker","attack_vector":"Unauthenticated network from any host that can reach the broker socket","remediation":"Upgrade. Providers running SGLang as a managed endpoint own this; tenants running their own own the patch but the provider owns fabric isolation","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3059"],"status":"curated","fleet":{"ubiquity":"Very common - affects the disaggregated/multimodal serving paths that large GPU deployments specifically use","remediation_pain":"`daemon-restart` plus network re-segmentation - the ZMQ broker binds 0.0.0.0 with zero auth, so patching alone is insufficient without isolating the internal serving plane","pain_class":"daemon-restart","why_fleet_wide":"`pickle.loads()` runs immediately on any payload received by an unauthenticated all-interfaces ZMQ broker, so anything with pod-network reach owns every disaggregated serving node - the multi-node serving fabric is the blast radius"}},{"id":"CVE-2026-3060","cve":"CVE-2026-3060","aliases":[],"title":"SGLang (encoder parallel disaggregation): Unauthenticated RCE via `pickle.loads()` in the disaggregation module","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (encoder parallel disaggregation)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Unauthenticated RCE via `pickle.loads()` in the disaggregation module","attack_vector":"Unauthenticated network on the intra-cluster fabric","remediation":"Upgrade; disaggregated serving multiplies unauthenticated internal planes","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3060"],"status":"curated"},{"id":"CVE-2026-30623","cve":"CVE-2026-30623","aliases":[],"title":"LiteLLM (MCP server creation): RCE via MCP server registration","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (MCP server creation)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via MCP server registration","attack_vector":"Network user able to add an MCP server","remediation":"Upgrade; MCP registration is a code-execution primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-30623"],"status":"curated"},{"id":"CVE-2026-31228","cve":"CVE-2026-31228","aliases":[],"title":"Kubeflow (ART component): RCE in the robustness evaluation function","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow (ART component)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE in the robustness evaluation function","attack_vector":"Customer-supplied evaluation config","remediation":"Upgrade ART","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31228"],"status":"curated"},{"id":"CVE-2026-31229","cve":"CVE-2026-31229","aliases":[],"title":"Kubeflow (Adversarial Robustness Toolbox component): Insecure deserialization in the Kubeflow model-loading component","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow (Adversarial Robustness Toolbox component)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Insecure deserialization in the Kubeflow model-loading component","attack_vector":"Customer-supplied model evaluated by a robustness pipeline","remediation":"Upgrade ART past 1.20.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31229"],"status":"curated"},{"id":"CVE-2026-34159","cve":"CVE-2026-34159","aliases":[],"title":"llama.cpp (RPC `deserialize_tensor`): RPC backend skips all bounds validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (RPC `deserialize_tensor`)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RPC backend skips all bounds validation","attack_vector":"Unauthenticated network to the RPC port","remediation":"Rebuild past b8492. Same unauthenticated-RPC class as 2024 — the design has not changed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34159"],"status":"curated"},{"id":"CVE-2026-35616","cve":"CVE-2026-35616","aliases":[],"title":"Fortinet FortiClient EMS: Improper access control","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiClient EMS","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Improper access control -> unauthenticated code execution via crafted requests","attack_vector":"Network (remote)","remediation":"Control-plane: patch FortiClientEMS 7.4.5-7.4.6","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35616"],"status":"curated"},{"id":"CVE-2026-39808","cve":"CVE-2026-39808","aliases":[],"title":"Fortinet FortiSandbox: OS command injection via crafted HTTP requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fortinet FortiSandbox","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] OS command injection via crafted HTTP requests -> unauthenticated code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency firmware; same window as CVE-2026-25089","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39808"],"status":"curated"},{"id":"CVE-2026-42208","cve":"CVE-2026-42208","aliases":[],"title":"LiteLLM proxy: SQL injection in a database query path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM proxy","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"SQL injection in a database query path","attack_vector":"Network to the LiteLLM proxy","remediation":"**[KEV]** Patch to 1.83.7+ immediately — known exploited. Any AI gateway a neocloud runs for tenants is in scope","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-42208"],"status":"curated"},{"id":"CVE-2026-43465","cve":"CVE-2026-43465","aliases":["net/mlx5e RX XDP multi-buf frag counting for striding RQ"],"title":"Linux kernel mlx5_core RX datapath (striding RQ, page_pool): A regression introduced by the fix for CVE-2025-40350: dropped XDP fragments stopped being counted…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core RX datapath (striding RQ, page_pool)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A regression introduced by the fix for CVE-2025-40350: dropped XDP fragments stopped being counted driver-side, so page_pool reference counts go negative and mlx5 releases 64 fragments against a refcount of 63. Remote packets corrupt page-pool refcounting in the host kernel. Worth flagging to operators as a pattern - patching the earlier RX bug without moving to a current stable re-exposes you.","attack_vector":"Unauthenticated remote sender to a node running XDP multi-buffer on mlx5 striding RQ, on a kernel that carries the CVE-2025-40350 fix but not this one.","remediation":"Upgrade the host kernel to 7.0 or a stable backport (6.18.19, 6.19.9). Rolling reboot. Do not stop at the CVE-2025-40350 fix level - verify your running kernel carries both.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43465","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-43465.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-44484","cve":"CVE-2026-44484","aliases":[],"title":"PyTorch Lightning: Reintroduced unsafe deserialization in 2.6.2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch Lightning","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Reintroduced unsafe deserialization in 2.6.2","attack_vector":"Customer-supplied checkpoint","remediation":"Tenant-owned; provider blocks pickle-format checkpoints at ingest","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44484"],"status":"curated"},{"id":"CVE-2026-46135","cve":"CVE-2026-46135","aliases":["CVE-2026-64534","CVE-2026-64535"],"title":"Linux kernel nvmet-tcp - ICReq/teardown race and data-digest error paths: Three lifecycle bugs an initiator can drive on purpose. Send an ICReq and immediately close the socket and…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel nvmet-tcp - ICReq/teardown race and data-digest error paths","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Three lifecycle bugs an initiator can drive on purpose. Send an ICReq and immediately close the socket and the target's queue state gets flipped back to LIVE after teardown already started, defeating the DISCONNECTING guard and leading to use-after-free. Negotiate data digests and deliberately mismatch one on a non-final H2C_DATA PDU and the target either underflows a refcount into a permanent workqueue deadlock or leaves a half-uninitialised command that teardown uninits a second time. The workqueue deadlock variant is the one that hurts operationally: the target stops serving I/O for every tenant and only a reboot clears it, so a single initiator - including a tenant VM with a legitimately provisioned namespace - can take the whole storage node down on demand and repeat it after every restart.","attack_vector":"Any host that can open an NVMe/TCP connection to the target. For the digest variants the attacker needs a connection that negotiates data digests, which any initiator can request; no valid namespace or credential is required for the ICReq race.","remediation":"Kernel update on the target nodes and a reboot. Disabling data digests removes two of the three but not the ICReq race, and costs you the end-to-end integrity check, so it is a poor trade. Because these are cheap remote denial-of-service primitives rather than deep exploitation, prioritise them by blast radius: any target serving more than one tenant should be patched before targets serving one. Rate-limiting or per-initiator connection caps on the storage VLAN reduce the repeat rate but do not fix it.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2026-46135","https://nvd.nist.gov/vuln/detail/CVE-2026-64535"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46325","cve":"CVE-2026-46325","aliases":["RDMA/rxe iova-to-va conversion error","Soft-RoCE MR page size mismatch"],"title":"Linux kernel - RDMA/rxe memory region translation, drivers/infiniband/sw/rxe/rxe_mr.c: TENANT ISOLATION: rxe mishandles memory regions whose page size differs from the system PAGE_SIZE - it steps…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe memory region translation, drivers/infiniband/sw/rxe/rxe_mr.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: rxe mishandles memory regions whose page size differs from the system PAGE_SIZE - it steps the page list by mr->page_size while each stored entry actually represents PAGE_SIZE of memory. The result is that an IO virtual address resolves to the wrong physical page, so a legitimate-looking remote access lands on memory outside the intended region. This is the most dangerous shape a memory-registration bug can take: the rkey check passes, the access is authorised, and the data returned or overwritten belongs to something else entirely. It matters concretely on ARM64 GPU hosts (64K pages) and anywhere hugepages are used for RDMA buffers - which is standard practice for training-job memory. A reported kernel panic is the visible symptom; silent cross-region reads and writes are the security consequence.","attack_vector":"A remote peer issues ordinary RDMA reads or writes against a memory region registered with a page size differing from the host PAGE_SIZE. No malformed packets needed - the mistranslation happens in the victim's own code path. Reachability is whatever the rxe endpoint's reachability is; combined with the rkey weaknesses described in the ReDMArk entry, an attacker who guesses into a region gets misdirected access on top of unauthorised access.","remediation":"Host reboot / kernel upgrade. Interim: unload and blacklist rdma_rxe if Soft-RoCE is not in deliberate use (config change, no downtime). If rxe is required, avoid registering regions whose page size differs from PAGE_SIZE until patched - in practice that means not backing RDMA buffers with hugepages on affected kernels, which costs performance but is a same-day application/config change. Prioritise ARM64 GPU hosts and any node with 64K pages.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46325.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46325"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-47627","cve":"CVE-2026-47627","aliases":[],"title":"NVIDIA Triton Inference Server: A path traversal scored 9.8 (network, no privileges, full confidentiality/integrity/availability impact).…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"A path traversal scored 9.8 (network, no privileges, full confidentiality/integrity/availability impact). NVIDIA's summary understates it as denial of service while the vector says complete compromise; treat the vector as authoritative and patch this first among the Triton set. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47627","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-22"]},{"id":"CVE-2026-49468","cve":"CVE-2026-49468","aliases":[],"title":"LiteLLM proxy: Host-header parsing flaw in the proxy","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM proxy","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Host-header parsing flaw in the proxy","attack_vector":"Unauthenticated network","remediation":"Upgrade to 1.84.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49468"],"status":"curated"},{"id":"CVE-2026-53176","cve":"CVE-2026-53176","aliases":["IB/isert short login PDU","iSER target pre-authentication crash"],"title":"Linux kernel - iSER (iSCSI Extensions for RDMA) target, drivers/infiniband/ulp/isert/ib_isert.c: TENANT ISOLATION: The iSER target accepted login PDUs shorter than ISER_HEADERS_LEN and parsed them…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - iSER (iSCSI Extensions for RDMA) target, drivers/infiniband/ulp/isert/ib_isert.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: The iSER target accepted login PDUs shorter than ISER_HEADERS_LEN and parsed them anyway, giving out-of-bounds access in the kernel before any authentication has occurred. Login is by definition the pre-auth surface, so anyone who can reach the iSER target port gets remote kernel compromise on the storage node with no credentials at all. A storage node in a GPU cluster typically mounts and serves many tenants' volumes, so compromising it is equivalent to compromising every dataset it fronts.","attack_vector":"Connect to the iSER target and send a truncated login PDU. Fully pre-authentication, no valid initiator identity required, reachable from anywhere on the storage fabric. If the storage fabric is not separated from the tenant fabric - a common shortcut - this is reachable from tenant workloads directly.","remediation":"Host reboot / kernel upgrade on iSER target nodes, treated as urgent given it is pre-auth and scored 9.8. Immediate config controls while you schedule the reboot: restrict the iSER/iSCSI target port to known initiator addresses at the switch and host firewall, and place storage targets on a fabric partition tenants cannot reach. If iSER is unused, unload ib_isert and disable the target configuration.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-53176.json","https://nvd.nist.gov/vuln/detail/CVE-2026-53176"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-5760","cve":"CVE-2026-5760","aliases":[],"title":"SGLang (`/v1/rerank`): RCE via a malicious `tokenizer.chat_template` rendered as Jinja2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/v1/rerank`)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"RCE via a malicious `tokenizer.chat_template` rendered as Jinja2","attack_vector":"Customer-supplied model file — the chat template inside the model repo is the payload","remediation":"Upgrade. Jinja chat templates are code; scanning the weights does not cover the tokenizer config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5760"],"status":"curated"},{"id":"CVE-2026-63077","cve":"CVE-2026-63077","aliases":[],"title":"JetBrains TeamCity: Deserialization in the agent polling protocol","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"JetBrains TeamCity","year":"2026","cvss_score":9.8,"severity":"critical","kev":true,"impact":"[KEV] Deserialization in the agent polling protocol -> unauthenticated remote code execution on the CI server","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT upgrade to 2026.1.3/2025.11.7; rotate all build secrets and signing keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63077"],"status":"curated"},{"id":"CVE-2026-63886","cve":"CVE-2026-63886","aliases":["LIO iSCSI CHAP_R base64 overflow","iscsi_target_auth pre-auth heap overflow"],"title":"Linux kernel - LIO iSCSI target CHAP authentication, drivers/target/iscsi/iscsi_target_auth.c: TENANT ISOLATION: chap_server_compute_hash() allocated the client digest buffer at the hash's digest…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - LIO iSCSI target CHAP authentication, drivers/target/iscsi/iscsi_target_auth.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: chap_server_compute_hash() allocated the client digest buffer at the hash's digest size, then passed a base64 CHAP_R response to the decoder without checking whether the input could produce more output than that. Up to 127 base64 characters decode to 95 bytes - a 63-byte overflow for SHA-256, 79 bytes for MD5 - and the existing length check fired only after the write had already happened. This is a heap overflow reached during CHAP authentication, meaning pre-authentication by definition: anyone who can open an iSCSI session to the target gets a controlled kernel heap write on the storage node. The HEX branch of the same switch already validated its length, so the bug was a straightforward omission on the base64 path.","attack_vector":"Open an iSCSI session to the LIO target and send a CHAP_R value with the '0b' base64 prefix and a long payload. No valid credentials required - the overflow happens while the target is still working out whether your credentials are valid. Reachable from any host that can reach the iSCSI port.","remediation":"Host reboot / kernel upgrade on all LIO iSCSI target nodes - urgent, pre-auth, 9.8. Interim controls: restrict the iSCSI target port to known initiator addresses at the firewall and switch, and if CHAP is not required, note that disabling it does not help since the parser is reached during login negotiation. Move iSCSI targets off tenant-reachable networks. Where the iSCSI target is legacy, decommissioning it is the cleanest fix.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-63886.json","https://nvd.nist.gov/vuln/detail/CVE-2026-63886"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64102","cve":"CVE-2026-64102","aliases":["RDMA/siw MPA FPDU length underflow"],"title":"Linux kernel - RDMA/siw (soft-iWARP) MPA framing, drivers/infiniband/sw/siw/siw_qp_rx.c: TENANT ISOLATION: The siw receive path decodes the MPA framed-PDU length from the wire and feeds it into…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/siw (soft-iWARP) MPA framing, drivers/infiniband/sw/siw/siw_qp_rx.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: The siw receive path decodes the MPA framed-PDU length from the wire and feeds it into signed arithmetic against the per-fragment remainder without rejecting an underflow first. A remote peer that sends a crafted FPDU length drives the subsequent length math negative, which downstream is used to size receive operations - the classic path to out-of-bounds kernel access from a network packet. Same reachability profile as the Read Response bug in the same file: remote, unauthenticated relative to the connection, over plain TCP, giving full compromise of the victim kernel.","attack_vector":"Any peer on an established siw connection sends an MPA FPDU whose declared length underflows the remaining fragment length. Because siw rides ordinary TCP, this is reachable from anything the node peers with, including across routed networks if the RDMA port is not firewalled. No local access and no credentials on the victim.","remediation":"Host reboot / kernel upgrade. Same practical advice as the sibling siw bug: unload and blacklist the siw module wherever soft-iWARP is not deliberately in use - a config change with zero downtime that eliminates both findings at once. Where siw is required, upgrade the kernel and reboot on a rolling drain; also firewall the iWARP TCP ports to known peers as an interim control.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64102.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64102"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64122","cve":"CVE-2026-64122","aliases":[],"title":"Linux kernel mlx5_core TX timeout devlink health reporter: The TX timeout recovery handler accesses the netdev pointer after the channel and its send queues have…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core TX timeout devlink health reporter","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"The TX timeout recovery handler accesses the netdev pointer after the channel and its send queues have already been torn down and freed - a use-after-free in the code path that is supposed to recover the NIC from a stall. So the failure mode is: the NIC wedges under load, the recovery machinery fires, and instead of recovering it corrupts host kernel memory.","attack_vector":"Remote and unauthenticated in effect - an attacker who can drive enough traffic to induce a TX timeout on the mlx5 interface reaches the recovery path. No credentials on the host.","remediation":"Upgrade the host kernel to a build carrying the fix (mainline 7.x and current stable series). Rolling reboot of the fleet. There is no useful config mitigation - you cannot safely turn off TX timeout recovery.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64122","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64122.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64268","cve":"CVE-2026-64268","aliases":["RDMA/siw Read Response overrun","soft-iWARP RREAD sink overflow"],"title":"Linux kernel - RDMA/siw (soft-iWARP), drivers/infiniband/sw/siw/siw_qp_rx.c: TENANT ISOLATION: siw places inbound Read Response segments at the sink buffer without checking the running…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/siw (soft-iWARP), drivers/infiniband/sw/siw/siw_qp_rx.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: siw places inbound Read Response segments at the sink buffer without checking the running total against the sink length on continuation segments - the sink is validated only on the first fragment and the cumulative length only on the last. A connected peer that answers an outstanding RDMA READ with segments that never set the DDP Last flag, carrying more payload than was requested, walks the write pointer past the validated buffer and writes attacker-controlled data out of bounds in the kernel. siw runs iWARP over ordinary routable TCP, so the attacker is simply the far end of an established connection and needs no local privilege on the victim. Remote kernel memory corruption with full confidentiality, integrity and availability loss - the worst-scored item in this slice.","attack_vector":"The attacker is the remote peer of an siw connection, reachable over normal TCP - which means anything the victim connects out to, including a storage target or a peer node in a job, can attack back. They respond to any RDMA READ the victim issues with a stream of Read Response segments with DDP Last clear and total length exceeding the RREAD length. No authentication is involved; the connection itself is the only prerequisite. Note siw is a software iWARP provider commonly loaded for testing, for nodes without an RNIC, and by blktests/nvme-over-fabrics setups - it is often present on GPU nodes that do not intend to use it.","remediation":"Host reboot / kernel upgrade to a version carrying the fix (backported across stable trees) - no firmware, no switch work. Immediate zero-downtime mitigation: if you are not deliberately using soft-iWARP, blacklist and unload the module (rmmod siw; add to modprobe blacklist), which removes the attack surface entirely and costs nothing on the vast majority of GPU nodes. Audit with lsmod across the fleet - siw is frequently loaded by dependency rather than by intent. If siw is in use, treat the kernel upgrade as urgent and drain nodes in waves.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64268.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64268"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-65791","cve":"CVE-2026-65791","aliases":["CVE-2026-65679","CVE-2026-65796","CVE-2026-65681"],"title":"Windows iSCSI Target Service (Windows Server 2012 through Windows Server 2025 / Windows 10 1607+): Three heap-based buffer overflows allowing an unauthorized attacker to execute code over the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Windows iSCSI Target Service (Windows Server 2012 through Windows Server 2025 / Windows 10 1607+)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Three heap-based buffer overflows allowing an unauthorized attacker to execute code over the network against the Windows iSCSI Target Service, plus a null-dereference denial of service in the same August 2026 batch. Unauthenticated network RCE against the service that owns every exported virtual disk on the box - the attacker gets SYSTEM on the storage server and, with it, read/write to every tenant VHD it serves. Relevant to GPU operators running Windows-based storage nodes or Hyper-V clusters backing GPU VMs, and to the long tail of Server 2012-era boxes still exporting iSCSI for management infrastructure.","attack_vector":"Any host that can reach TCP 3260 on the Windows storage server. No credentials.","remediation":"Apply the August 2026 Windows cumulative update on every server running the iSCSI Target Service role and reboot - which drops all iSCSI sessions and any VM or host booting from those LUNs, so drain first. If the role is enabled but unused (a common leftover on general-purpose Windows servers), remove the role instead of patching it; that permanently deletes the exposure. Restrict 3260 to known initiator addresses with Windows Firewall regardless of patch state. Note Server 2012 is out of mainstream support - confirm your ESU covers this batch or plan the migration.","references":["https://msrc.microsoft.com/update-guide/vulnerability/CVE-2026-65791","https://nvd.nist.gov/vuln/detail/CVE-2026-65791"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72129","cve":"CVE-2026-72129","aliases":["nvmet-rdma inline data nonzero offset","NVMe-oF RDMA inline scatterlist underflow"],"title":"Linux kernel - NVMe-oF RDMA target, drivers/nvme/target/rdma.c: TENANT ISOLATION: nvmet_rdma_use_inline_sg() accepted any host-controlled inline data offset that satisfied…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - NVMe-oF RDMA target, drivers/nvme/target/rdma.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: nvmet_rdma_use_inline_sg() accepted any host-controlled inline data offset that satisfied off + len <= inline_data_size, but the mapping still assumed the data started in the first inline page. When a port is configured with inline_data_size larger than PAGE_SIZE - a legitimate, performance-motivated setting, allowed up to 16K - an offset past PAGE_SIZE makes the length calculation underflow to roughly 4 GiB, and the block backend reads far past the intended page. There is also a double-free path where a persistent inline page is mistaken for an allocated scatterlist and freed. Remote, host-controlled, on the NVMe-oF-over-RDMA target that fronts tenant storage: this is the disaggregated-storage isolation break in its most direct form.","attack_vector":"A connected NVMe-oF initiator submits a command with an inline data offset in the range between PAGE_SIZE and the port's configured inline_data_size. Requires the target port to be configured with inline_data_size > PAGE_SIZE, which operators do deliberately to cut latency on small writes - so the hardened, performance-tuned configuration is the vulnerable one.","remediation":"Host reboot / kernel upgrade on NVMe-oF RDMA targets. Immediate config-only mitigation with a measurable but acceptable cost: set the port's inline_data_size to PAGE_SIZE or less (nvmet configfs, applied on port re-enable, brief reconnect for attached initiators) - this removes the precondition entirely without a reboot. Then schedule the kernel upgrade and restore the larger inline size afterwards if the latency mattered.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-72129.json","https://nvd.nist.gov/vuln/detail/CVE-2026-72129"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72130","cve":"CVE-2026-72130","aliases":["nvmet-auth short AUTH_RECEIVE heap write","DH-HMAC-CHAP SUCCESS1 overflow"],"title":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c: TENANT ISOLATION: The sibling of the previous finding and worse - this one writes.…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: The sibling of the previous finding and worse - this one writes. nvmet_execute_auth_receive() trusted the AUTH_RECEIVE allocation length after checking only that it was non-zero, so a remote initiator supplying a one-byte allocation length reaches the fixed-size response builders with an undersized buffer and triggers a 16-byte heap out-of-bounds write. The code merely warned about the short length and then formatted the response anyway. A controlled heap write from a remote peer on a storage node is a direct path to kernel code execution and thus to every tenant volume the target serves.","attack_vector":"A remote NVMe-oF initiator with access to an auth-enabled target sends an AUTH_RECEIVE with a one-byte allocation length while the exchange is in the SUCCESS1 or FAILURE1 state. Reachable only when in-band DH-HMAC-CHAP is configured - so again, this is a risk you take on by following the standard hardening advice on an unpatched kernel.","remediation":"Host reboot / kernel upgrade on NVMe-oF targets. Patch before enabling DH-HMAC-CHAP; if auth is already on and the kernel is unpatched, treat the upgrade as urgent and restrict target reachability to known initiators at the firewall in the meantime. Roll targets in waves with multipath initiators so tenants see no I/O interruption.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-72130.json","https://nvd.nist.gov/vuln/detail/CVE-2026-72130"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-7301","cve":"CVE-2026-7301","aliases":[],"title":"SGLang (scheduler ROUTER socket): ROUTER socket binds `0.0.0.0` by default and `pickle.loads()` incoming messages","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (scheduler ROUTER socket)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"ROUTER socket binds `0.0.0.0` by default and `pickle.loads()` incoming messages","attack_vector":"Unauthenticated network — default configuration is exploitable","remediation":"Upgrade and override the bind address. The insecure default means an unmodified deployment is exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7301"],"status":"curated"},{"id":"CVE-2026-7304","cve":"CVE-2026-7304","aliases":[],"title":"SGLang (custom logit processor): `dill.loads` on user objects when `--enable-custom-logit-processor` is set","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (custom logit processor)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"`dill.loads` on user objects when `--enable-custom-logit-processor` is set → unauthenticated RCE","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade; never enable custom logit processors on a tenant-facing endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7304"],"status":"curated"},{"id":"CVE-2026-74269","cve":"CVE-2026-74269","aliases":[],"title":"Linux bnxt_en driver (XDP head-grow underflow): Head underflow when an XDP program grows the packet head on a Broadcom NIC, crashing the host. Found by the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (XDP head-grow underflow)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"Head underflow when an XDP program grows the packet head on a Broadcom NIC, crashing the host. Found by the kernel's own XDP test suite, which is a reminder that the XDP fast paths on vendor NIC drivers get materially less real-world coverage than the normal receive path — and AI-serving front-ends are one of the few places they run at scale.","attack_vector":"An attached XDP program that grows the packet head, processing received traffic.","remediation":"Kernel/driver upgrade plus host reboot. Interim: detach head-growing XDP programs from Broadcom interfaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74269"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74394","cve":"CVE-2026-74394","aliases":["RDMA/srpt immediate data integer overflow","SRP target"],"title":"Linux kernel - SRP target (srpt), drivers/infiniband/ulp/srpt/ib_srpt.c: TENANT ISOLATION: An integer overflow in the immediate-data length check on the SRP target lets a remote…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - SRP target (srpt), drivers/infiniband/ulp/srpt/ib_srpt.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: An integer overflow in the immediate-data length check on the SRP target lets a remote initiator bypass the bound and reach kernel memory it should not. The srpt target serves block storage over InfiniBand to compute clients, so a single tenant that can connect as an initiator gets kernel-level compromise of the storage node - and through it, access to every other tenant's volumes exported from the same target.","attack_vector":"A remote SRP initiator submits a command whose immediate data length is chosen so the length arithmetic wraps, defeating the check. Any host allowed to connect to the target can do it; on fabrics without per-tenant partitioning that is any host on the subnet.","remediation":"Host reboot / kernel upgrade on SRP target nodes. Interim: enforce InfiniBand partitioning so only authorised initiators can reach the target port (SM config change, effective on the next sweep, no reload), and restrict target ACLs to known initiator GUIDs. Where SRP has been superseded by NVMe-oF, decommission the srpt target and unload the module - the cleanest fix.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-74394.json","https://nvd.nist.gov/vuln/detail/CVE-2026-74394"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74474","cve":"CVE-2026-74474","aliases":[],"title":"Linux VXLAN driver (transmit-path header pulls): TENANT ISOLATION: `vxlan_xmit()`, `arp_reduce()` and `vxlan_mdb_entry_skb_get()` validate header availability…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux VXLAN driver (transmit-path header pulls)","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: `vxlan_xmit()`, `arp_reduce()` and `vxlan_mdb_entry_skb_get()` validate header availability with `pskb_may_pull()`, but on the transmit path `skb->data` points at the MAC header, so the offset accounting is wrong by the Ethernet header length and the driver reads past the validated region. This is the Linux kernel VXLAN data path — the software VTEP used by Linux-based switch OSes (SONiC, Cumulus), by container overlay networks, and by any host doing VXLAN encapsulation itself. Memory corruption in the encapsulation path of a multi-tenant overlay is as close to the centre of the tenant-isolation boundary as this database gets.","attack_vector":"Traffic traversing the VXLAN transmit path on an affected host or switch — reachable from inside a tenant overlay, since tenants generate the frames being encapsulated.","remediation":"Kernel upgrade plus host reboot; on Linux-based switch NOSes it arrives as a NOS image upgrade plus switch reload, so stage across MLAG/ECMP pairs. No config workaround short of not using VXLAN. Part of a 2026 cluster of VXLAN data-path fixes — CVE-2026-74475, CVE-2026-74473, CVE-2026-74406, CVE-2026-63993 — so take the whole batch in one upgrade rather than chasing them individually.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74474","https://nvd.nist.gov/vuln/detail/CVE-2026-74473","https://nvd.nist.gov/vuln/detail/CVE-2026-63993"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74556","cve":"CVE-2026-74556","aliases":["libiscsi_tcp SCSI Response data segment overflow","open-iscsi initiator conn->data overflow"],"title":"Linux kernel - iSCSI TCP initiator, drivers/scsi/libiscsi_tcp.c: TENANT ISOLATION: The iSCSI initiator receives PDU data segments into a fixed 8192-byte conn","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - iSCSI TCP initiator, drivers/scsi/libiscsi_tcp.c","year":"2026","cvss_score":9.8,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: The iSCSI initiator receives PDU data segments into a fixed 8192-byte conn->data buffer. The login, text, reject and async-event paths all check the DataSegmentLength against that size, but the SCSI Command Response path copies sense and response data in without the same check. The only remaining bound is the initiator's advertised MaxRecvDataSegmentLength, which open-iscsi defaults to 262144 - so a target returning a SCSI Response with a data segment between 8193 and 262144 bytes overflows the buffer by up to 250 KB. Target-to-initiator again: one compromised or rogue iSCSI target achieves kernel memory corruption on every compute node that mounts from it.","attack_vector":"The attacker controls an iSCSI target the victim connects to, or can inject into the TCP session, and returns a SCSI Command Response with an oversized data segment. Because open-iscsi's default MaxRecvDataSegmentLength is far above the buffer size, the vulnerable window is wide in stock configurations rather than requiring unusual tuning.","remediation":"Host reboot / kernel upgrade on all iSCSI initiator nodes - the compute fleet, so plan a rolling drain. Immediate config mitigation that closes the window without a reboot: set MaxRecvDataSegmentLength to 8192 or below in /etc/iscsi/iscsid.conf and reconnect sessions, which caps the receive length at the buffer size. Also verify targets are authenticated (mutual CHAP) so a rogue target cannot attract sessions.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-74556.json","https://nvd.nist.gov/vuln/detail/CVE-2026-74556"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-3679","cve":"CVE-2018-3679","aliases":[],"title":"Intel Data Center Manager SDK (reference UI): The DCM SDK's reference UI allows an unauthenticated remote attacker to execute code with administrator…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center Manager SDK (reference UI)","year":"2018","cvss_score":9.6,"severity":"critical","kev":false,"impact":"The DCM SDK's reference UI allows an unauthenticated remote attacker to execute code with administrator privileges. Reference UIs get shipped into production more often than vendors expect - if you built anything on the DCM SDK, check whether the sample UI went with it.","attack_vector":"Unauthenticated remote attacker able to reach the reference UI.","remediation":"Upgrade the DCM SDK past 5.0 and remove the reference UI from any production deployment. Application-level change; no host reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3679","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00143.html"],"status":"curated"},{"id":"CVE-2019-15897","cve":"CVE-2019-15897","aliases":[],"title":"BeeGFS (beegfs-ctl / metadata server): TENANT ISOLATION: authentication bypass by talking directly to a BeeGFS metadata server. BeeGFS is a common…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"BeeGFS (beegfs-ctl / metadata server)","year":"2019","cvss_score":9.6,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: authentication bypass by talking directly to a BeeGFS metadata server. BeeGFS is a common choice for AI-training scratch storage because it is fast and easy to stand up, and its threat model assumes the storage network is private. Anyone who reaches the metadata server can act against the filesystem's metadata — which means other tenants' namespaces on a shared BeeGFS deployment.","attack_vector":"Network access to a BeeGFS metadata server. The advisory notes such servers are typically not exposed externally — but inside a GPU cluster, 'not externally exposed' still means reachable by every tenant workload on the storage VLAN.","remediation":"Upgrade BeeGFS past 7.1.3 and enable connection authentication (the shared-secret `connAuthFile`), then restart the metadata, storage and client services — a coordinated restart that stalls I/O, so drain jobs first. Enabling connAuth is the load-bearing step; the version upgrade alone does not help if authentication stays off.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-15897"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-21538","cve":"CVE-2021-21538","aliases":["DSA-2021-082"],"title":"Dell iDRAC9 (Virtual Console / authentication): An attacker with no credentials lands directly inside the server's Virtual Console - the same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (Virtual Console / authentication)","year":"2021","cvss_score":9.6,"severity":"critical","kev":false,"impact":"An attacker with no credentials lands directly inside the server's Virtual Console - the same screen-and-keyboard the operator uses. From there they see whatever the tenant is running, drive the BIOS/boot menu, and pair it with Virtual Media to boot a node off an image they supply. On a bare-metal cloud that is a silent cross-tenant handoff failure: the previous or a neighbouring tenant's session is visible and controllable without ever touching the production network. Only iDRAC9 firmware in the 4.40.00.00-4.40.09.99 band is affected, so this is a narrow-window regression that is easy to miss in a mixed-vintage fleet.","attack_vector":"Anything routable to the iDRAC address on the out-of-band management VLAN, no account and no prior foothold. If the OOB network is flat across racks, a single compromised jump host or a mis-scoped VPN split-tunnel reaches every node in the affected firmware band.","remediation":"Flash iDRAC9 to 4.40.10.00 or later. This is per-node but fully out-of-band (iDRAC web UI, racadm, Redfish SimpleUpdate, or OME) and does NOT require a host reboot - the iDRAC resets itself and you lose OOB reachability for roughly two to five minutes while the host keeps running its jobs. No drain needed. Immediate config-only mitigation while you roll: disable Virtual Console and Virtual Media under iDRAC Settings, and ACL the management VLAN down to the jump hosts.","references":["https://www.dell.com/support/kbdoc/000186420","https://nvd.nist.gov/vuln/detail/CVE-2021-21538"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2021-21596","cve":"CVE-2021-21596","aliases":[],"title":"Dell OpenManage Enterprise (remote code execution): Remote code execution on the OpenManage Enterprise console. OME is the fleet-wide control plane that already…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell OpenManage Enterprise (remote code execution)","year":"2021","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Remote code execution on the OpenManage Enterprise console. OME is the fleet-wide control plane that already holds credentials for, and can push firmware to, every iDRAC it manages - so compromising it is not compromising one node, it is compromising the mechanism that drives all of them. An attacker in OME can trigger firmware deployment, mount Virtual Media, and power-cycle at fleet scale from a single box. Affects OME versions before 3.6.2 and the corresponding OME-Modular builds.","attack_vector":"An attacker with access to the immediate subnet the OME appliance sits on. That is usually the management network segment, so the practical question is who else lives on the same VLAN as your OME appliance - jump hosts, monitoring, DCIM, and often a broader IT segment than anyone intends.","remediation":"Upgrade the OME appliance to 3.6.2 or later. This is a single appliance upgrade, not a per-node campaign, so it is cheap in rollout terms - no node reboots, no job drain, only the OME console's own downtime. The structural fix is network placement: put OME on its own segment with an explicit allowlist rather than sharing the general management VLAN, and treat it as tier-0 infrastructure because it holds BMC credentials for the whole fleet.","references":["https://www.dell.com/support/kbdoc/000189673","https://nvd.nist.gov/vuln/detail/CVE-2021-21596"],"status":"curated"},{"id":"CVE-2022-24422","cve":"CVE-2022-24422","aliases":["DSA-2022-068"],"title":"Dell iDRAC9 (VNC server): Unauthenticated access to the iDRAC VNC console. Same practical outcome as an unauthenticated KVM: the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (VNC server)","year":"2022","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Unauthenticated access to the iDRAC VNC console. Same practical outcome as an unauthenticated KVM: the attacker watches and drives the tenant's console, can reach the boot menu and firmware setup, and can chain into Virtual Media to boot an attacker image below the hypervisor. Affects iDRAC9 5.00.00.00 up to 5.10.10.00, so it hits the 15G PowerEdge generation that a lot of first-wave GPU fleets standardised on.","attack_vector":"Anything routable to the iDRAC VNC port on the management VLAN, unauthenticated - but only where the iDRAC VNC server has actually been enabled. It is off by default, so the real exposure is the subset of nodes where someone turned it on for remote hands.","remediation":"Flash iDRAC9 to 5.10.10.00 or later - out-of-band, per-node, no host reboot and no job drain (iDRAC self-resets, brief OOB blackout only). Because VNC is opt-in, the fastest mitigation is config-only and costs nothing: audit which nodes have the iDRAC VNC server enabled and turn it off. That closes the hole fleet-wide in minutes while the firmware campaign runs on its own schedule.","references":["https://www.dell.com/support/kbdoc/en-us/000199267/dsa-2022-068-dell-idrac9-security-update-for-an-improper-authentication-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2022-24422"],"status":"curated"},{"id":"CVE-2022-41924","cve":"CVE-2022-41924","aliases":[],"title":"Tailscale (Windows client): Local API bound to a TCP socket","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Tailscale (Windows client)","year":"2022","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Local API bound to a TCP socket -> a malicious website reconfigures tailscaled and achieves RCE","attack_vector":"Network (remote)","remediation":"Control-plane: force-update all operator clients","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41924"],"status":"curated"},{"id":"CVE-2023-3043","cve":"CVE-2023-3043","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): The twin of CVE-2023-37293: a stack smash in the BMC's network-facing path that hands an unauthenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":9.6,"severity":"critical","kev":false,"impact":"The twin of CVE-2023-37293: a stack smash in the BMC's network-facing path that hands an unauthenticated attacker execution on the management controller. Once there the attacker can hold power control over the node indefinitely, mount virtual media to boot an attacker image, read whatever the BMC can see, and write persistent firmware. The persistence is the point - this survives OS reinstall, disk replacement and node rebuild, so a fleet operator who rotates a suspect node back into the pool reintroduces the implant.","attack_vector":"Adjacent network, unauthenticated, low complexity. Reachable from anything sharing the BMC's broadcast domain. No tenant credentials, no BMC credentials, no clicking required.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6 - out-of-band, per node, gated on your ODM publishing a rebased image. Budget for a fleet-wide BMC flash campaign, not a patch window: BMC images are not managed by your OS config management and each SKU needs its own tested build. Config-only interim mitigation is VLAN isolation of the BMC plane plus removing any tenant-network route to it; disabling IPMI-over-LAN does not close this one because the exposure is the BMC's own network stack.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-3043"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-37293","cve":"CVE-2023-37293","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): Unauthenticated code execution inside the BMC, reached with a crafted packet and nothing else. The attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Unauthenticated code execution inside the BMC, reached with a crafted packet and nothing else. The attacker lands below the hypervisor on a processor that stays powered whether or not the node is booted, and from there owns the node's power state, its console, its virtual media and its BIOS flash. On a GPU fleet this is the classic rack-wide implant: reimaging the tenant OS, replacing the NVMe, or rebuilding the node from the provisioning system does not remove it. AMI rates the scope as changed, meaning the BMC compromise is expected to reach beyond the BMC into the host.","attack_vector":"Any host on the same L2 segment as the BMC management interface - no credentials, no user interaction, low attack complexity. In practice that means any compromised node, switch, PDU, out-of-band jump box or IPMI-speaking tool that shares the management VLAN. If BMCs are flat-VLANed across a hall (a common shortcut in leased colo and neocloud builds), one foothold reaches every BMC in the hall.","remediation":"BMC firmware flash to MegaRAC SPx_12.7 / SPx_13.6 or later - out-of-band, per node, with real bricking risk if power or the transfer is interrupted mid-write. You cannot take AMI's build directly: your ODM (Supermicro, Quanta, Gigabyte, Wiwynn, ASRock Rack, Inventec) must rebase and requalify, and that lag has historically run six months to over a year, longer on white-box SKUs whose vendor has stopped shipping BMC images. The host generally stays up through the flash but KVM/vMedia sessions drop and the BMC reboots. Until the image exists, the only real control is network: put BMCs on an isolated management VLAN with no route from tenant, storage or corporate networks, and reach them only through a bastion.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-37293","https://www.ami.com/security-advisories/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-24593","cve":"CVE-2024-24593","aliases":[],"title":"ClearML API server: CSRF against the API server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML API server","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"CSRF against the API server","attack_vector":"Logged-in operator visiting a malicious page","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24593"],"status":"curated"},{"id":"CVE-2024-27892","cve":"CVE-2024-27892","aliases":[],"title":"Arista EOS (OpenConfig gNMI Set authorization): A gNMI Set request that authorization should have rejected is executed, so a caller writes switch…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (OpenConfig gNMI Set authorization)","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"A gNMI Set request that authorization should have rejected is executed, so a caller writes switch configuration it has no right to write. Model-driven management is how large fabrics are actually operated now, and gNMI is the write path — if its authorization does not hold, your RBAC on the fabric does not exist. Pairs with CVE-2024-27890 (same defect, separate advisory) and CVE-2025-1260 on the gNOI side.","attack_vector":"A client able to reach the gNMI endpoint with credentials whose authorization should have been insufficient. Requires OpenConfig to be configured.","remediation":"EOS upgrade plus reload. Interim: restrict gNMI/gNOI reachability with a control-plane ACL to only the automation hosts that legitimately write config — live config change and worth doing permanently.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27892","https://nvd.nist.gov/vuln/detail/CVE-2024-27890"],"status":"curated"},{"id":"CVE-2024-34359","cve":"CVE-2024-34359","aliases":[],"title":"llama-cpp-python: RCE via Jinja2 template in a GGUF model's metadata (`Llama` class)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-cpp-python","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"RCE via Jinja2 template in a GGUF model's metadata (`Llama` class)","attack_vector":"Customer-supplied GGUF model file — the chat template inside it is the payload","remediation":"Upgrade; GGUF chat templates are executable and are not covered by pickle scanning","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-34359"],"status":"curated"},{"id":"CVE-2024-35225","cve":"CVE-2024-35225","aliases":[],"title":"Jupyter Server Proxy: Unauthenticated web access to a user's proxied processes","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server Proxy","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Unauthenticated web access to a user's proxied processes","attack_vector":"Network attacker reaching the hub","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35225"],"status":"curated"},{"id":"CVE-2024-6385","cve":"CVE-2024-6385","aliases":[],"title":"GitLab: Attacker can trigger a CI pipeline as another user","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitLab","year":"2024","cvss_score":9.6,"severity":"critical","kev":false,"impact":"Attacker can trigger a CI pipeline as another user -> jobs run with that user's token","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate CI job and runner registration tokens","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6385"],"status":"curated"},{"id":"CVE-2026-50540","cve":"CVE-2026-50540","aliases":[],"title":"Kata Containers: kata-runtime host code execution via an untrusted input path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":9.6,"severity":"critical","kev":false,"impact":"kata-runtime host code execution via an untrusted input path","attack_vector":"Any tenant workload / malicious image under Kata","remediation":"Emergency Kata upgrade; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-50540"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-8037","cve":"CVE-2026-8037","aliases":[],"title":"Progress LoadMaster (ADC): OS command injection in the API","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Progress LoadMaster (ADC)","year":"2026","cvss_score":9.6,"severity":"critical","kev":true,"impact":"[KEV] OS command injection in the API -> unauthenticated remote code execution on the load balancer","attack_vector":"Adjacent network","remediation":"Control-plane: patch; the ADC terminates tenant-facing TLS","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-8037"],"status":"curated"},{"id":"CVE-2025-26385","cve":"CVE-2025-26385","aliases":["ICSA-26-027-04"],"title":"Johnson Controls Metasys Application and Data Server (ADS) deployed with SQL Express: Command injection on the Metasys ADS that yields remote SQL execution. The ADS is the site's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Johnson Controls Metasys Application and Data Server (ADS) deployed with SQL Express","year":"2025","cvss_score":9.5,"severity":"critical","kev":false,"impact":"Command injection on the Metasys ADS that yields remote SQL execution. The ADS is the site's building-automation database and supervisory server: it holds the point database, the schedules, the trends and the operator accounts for every NAE/SNE/SNC engine driving air handling in the building. Getting arbitrary SQL - and, through it, the usual SQL-Server-to-OS escalation paths - means an attacker owns the system that both commands and reports on cooling. They can rewrite schedules and setpoints so the change persists, and rewrite the trend history so the postmortem shows nothing anomalous. For a hall of 40 kW+ racks that is fleet-wide availability risk with an intact-looking dashboard, which is the worst combination for an operator trying to diagnose why a training run just died.","attack_vector":"Network access to the Metasys ADS. Metasys servers are Windows machines that facilities teams routinely domain-join and expose to the corporate network for the Site Management Portal, so the realistic path is a corporate-network foothold rather than direct internet exposure - though internet-published SMP instances do exist. Anyone who compromises a facilities workstation is one hop away.","remediation":"Vendor software update to the ADS per the Johnson Controls advisory - a server-side patch and a normal Windows change window, no controller firmware and no cooling downtime, so schedule it. Additionally: get the ADS off the corporate network segment entirely, restrict SQL Express to localhost, run the Metasys service account with least privilege, and require a jump host for SMP access. Leased colo: the Metasys ADS is the landlord's building server and typically serves all tenants - you cannot patch it, so make its version and network placement a contractual disclosure and demand notification when JCI publishes an advisory.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-027-04","https://nvd.nist.gov/vuln/detail/CVE-2025-26385","https://www.johnsoncontrols.com/trust-center/cybersecurity/security-advisories"],"status":"curated"},{"id":"CVE-2026-44946","cve":"CVE-2026-44946","aliases":[],"title":"Rancher: SAML assertion replay: the ACS handler does not enforce one-time use, so a captured assertion logs an…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2026","cvss_score":9.5,"severity":"critical","kev":false,"impact":"SAML assertion replay: the ACS handler does not enforce one-time use, so a captured assertion logs an attacker in","attack_vector":"Unauthenticated network in a MITM or log-access position","remediation":"Emergency Rancher upgrade; invalidate all SAML sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44946"],"status":"curated"},{"id":"CVE-2023-25131","cve":"CVE-2023-25131","aliases":["JVN#95119483"],"title":"CyberPower PowerPanel Business Local/Remote/Management v4.8.6 and earlier (Windows and Linux): A default password that ships enabled. This is the most boring vulnerability in the facility layer and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Business Local/Remote/Management v4.8.6 and earlier (Windows and Linux)","year":"2023","cvss_score":9.4,"severity":"critical","kev":false,"impact":"A default password that ships enabled. This is the most boring vulnerability in the facility layer and probably the most exploited class of them in practice - power management software gets installed once by whoever racked the gear and never revisited.","attack_vector":"Unauthenticated remote access to the PowerPanel Business interface using published default credentials.","remediation":"Upgrade past v4.8.6 and change the credential. Then go and check every other piece of power-management software you inherited with a build: default credentials on facility gear are the single highest-yield internal audit an operator can run, and it costs a day.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25131"],"status":"curated"},{"id":"CVE-2023-3128","cve":"CVE-2023-3128","aliases":[],"title":"Grafana: Azure AD accounts validated on the mutable, non-unique email claim","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2023","cvss_score":9.4,"severity":"critical","kev":false,"impact":"Azure AD accounts validated on the mutable, non-unique email claim -> account takeover / auth bypass","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; pin the OAuth tenant and allowed groups","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3128"],"status":"curated"},{"id":"CVE-2023-4966","cve":"CVE-2023-4966","aliases":[],"title":"Citrix NetScaler ADC/Gateway: \"CitrixBleed\" - memory overread leaking valid session tokens","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":9.4,"severity":"critical","kev":true,"impact":"[KEV] \"CitrixBleed\" - memory overread leaking valid session tokens -> MFA bypass","attack_vector":"Network (remote)","remediation":"Control-plane: patch AND terminate every existing session; tokens survive patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4966"],"status":"curated"},{"id":"CVE-2024-0964","cve":"CVE-2024-0964","aliases":[],"title":"Gradio: Remotely triggerable local file include via a JSON value in an API request","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":9.4,"severity":"critical","kev":false,"impact":"Remotely triggerable local file include via a JSON value in an API request","attack_vector":"Unauthenticated network to the demo","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0964"],"status":"curated"},{"id":"CVE-2026-4404","cve":"CVE-2026-4404","aliases":[],"title":"Harbor: Hard-coded default credentials give web UI access to the whole registry","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2026","cvss_score":9.4,"severity":"critical","kev":false,"impact":"Hard-coded default credentials give web UI access to the whole registry","attack_vector":"Unauthenticated network","remediation":"Upgrade Harbor immediately and change the default password; treat every hosted image as potentially tampered with and re-verify signatures","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-4404"],"status":"curated"},{"id":"CVE-2026-53488","cve":"CVE-2026-53488","aliases":[],"title":"containerd: CRI plugin propagates unvalidated image LABEL values into container config","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":9.4,"severity":"critical","kev":false,"impact":"CRI plugin propagates unvalidated image LABEL values into container config; highest-severity containerd issue to date","attack_vector":"Malicious image","remediation":"Emergency rolling containerd upgrade with node drain; gate tenant image sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53488"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-37678","cve":"CVE-2021-37678","aliases":[],"title":"TensorFlow / Keras: Arbitrary code execution via unsafe YAML deserialization of model config","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TensorFlow / Keras","year":"2021","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Arbitrary code execution via unsafe YAML deserialization of model config","attack_vector":"Customer-supplied model YAML","remediation":"Historical but still present in pinned tenant envs; patch base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-37678"],"status":"curated"},{"id":"CVE-2023-1177","cve":"CVE-2023-1177","aliases":[],"title":"MLflow (tracking server): Path traversal (`\\..\\filename`)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (tracking server)","year":"2023","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Path traversal (`\\..\\filename`) → arbitrary file read","attack_vector":"Unauthenticated network to the MLflow tracking server","remediation":"Upgrade to 2.2.1+; MLflow has no auth by default","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-1177"],"status":"curated","fleet":{"ubiquity":"Common - MLflow is the default experiment/model registry alongside GPU training clusters","remediation_pain":"`daemon-restart` - upgrade the tracking server; the pain is credential rotation, since leaked SSH keys and cloud creds must be assumed compromised","pain_class":"daemon-restart","why_fleet_wide":"Unauthenticated LFI via the Model Versions API reads any file the server can read - SSH keys, cloud credentials - and one tracking server usually fronts every team's models and artifact buckets"}},{"id":"CVE-2023-24509","cve":"CVE-2023-24509","aliases":[],"title":"Arista EOS (redundant supervisor, RPR/SSO): On modular chassis with dual supervisors running RPR or SSO redundancy, an existing unprivileged user can log…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (redundant supervisor, RPR/SSO)","year":"2023","cvss_score":9.3,"severity":"critical","kev":false,"impact":"On modular chassis with dual supervisors running RPR or SSO redundancy, an existing unprivileged user can log in to the standby supervisor as root. The standby has full access to the chassis state and becomes the active supervisor on failover, so this is a straight unprivileged-to-root path on your largest, most central switches — typically the spines.","attack_vector":"Authenticated but unprivileged user with network access to the standby supervisor.","remediation":"EOS upgrade. Requires a reload of the supervisors — do the standby first, fail over, then the former active, which keeps the chassis forwarding throughout. Until patched, restrict management reachability to the standby supervisor's address.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-24509"],"status":"curated"},{"id":"CVE-2024-22252","cve":"CVE-2024-22252","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Use-after-free in the XHCI USB controller - guest-to-host code execution as the VMX process","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2024","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Use-after-free in the XHCI USB controller - guest-to-host code execution as the VMX process","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with evacuation. Interim mitigation: remove USB controllers from tenant VMs, which is cheap and usually harmless for GPU workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22252"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-22253","cve":"CVE-2024-22253","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Use-after-free in the UHCI USB controller - guest-to-host code execution as the VMX process","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2024","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Use-after-free in the UHCI USB controller - guest-to-host code execution as the VMX process","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation; same USB-controller removal mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22253"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-3573","cve":"CVE-2024-3573","aliases":[],"title":"MLflow (LFI via URI parsing): Local file inclusion — read arbitrary files","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (LFI via URI parsing)","year":"2024","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Local file inclusion — read arbitrary files","attack_vector":"Unauthenticated network to the tracking server","remediation":"Upgrade; one of ~15 traversal variants, each a bypass of the last","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3573"],"status":"curated"},{"id":"CVE-2025-22224","cve":"CVE-2025-22224","aliases":[],"title":"VMware ESXi / Workstation: TOCTOU race leading to an out-of-bounds write in VMX - full VM escape to host code execution","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation","year":"2025","cvss_score":9.3,"severity":"critical","kev":true,"impact":"TOCTOU race leading to an out-of-bounds write in VMX - full VM escape to host code execution; exploited in the wild as a zero-day [KEV]","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with vMotion evacuation. GPU-passthrough VMs cannot vMotion, so this is a tenant-visible drain. Top-priority escape of the 2025 set","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22224"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-33187","cve":"CVE-2025-33187","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An attacker with privileged access reaches SoC-protected areas through SROOT, ending in code execution below…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"An attacker with privileged access reaches SoC-protected areas through SROOT, ending in code execution below the OS. At 9.3 this is the most severe entry in the GB10 firmware set. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33187","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-41236","cve":"CVE-2025-41236","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Integer overflow in the VMXNET3 virtual NIC - VM escape to host code execution (Pwn2Own Berlin 2025)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Integer overflow in the VMXNET3 virtual NIC - VM escape to host code execution (Pwn2Own Berlin 2025)","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation. VMXNET3 is the default adapter, so there is no practical configuration mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41236"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-41237","cve":"CVE-2025-41237","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Integer underflow in VMCI leading to an out-of-bounds write - VM escape (Pwn2Own Berlin 2025)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Integer underflow in VMCI leading to an out-of-bounds write - VM escape (Pwn2Own Berlin 2025)","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41237"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-41238","cve":"CVE-2025-41238","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Heap overflow in the PVSCSI controller - out-of-bounds write and VM escape (Pwn2Own Berlin 2025)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Heap overflow in the PVSCSI controller - out-of-bounds write and VM escape (Pwn2Own Berlin 2025)","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41238"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-53696","cve":"CVE-2025-53696","aliases":["CVE-2025-53695"],"title":"Software House iSTAR Ultra firmware verification and web application (tested through 6.9.2): The controller verifies its firmware at boot, but the verification skips portions of the image - so an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Software House iSTAR Ultra firmware verification and web application (tested through 6.9.2)","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The controller verifies its firmware at boot, but the verification skips portions of the image - so an attacker who can push firmware gets persistent, boot-surviving code on the door controller that the device's own integrity check will happily bless. The companion issue is an authenticated OS command injection in the web application that escalates to root, which is a plausible way to reach the firmware write in the first place. This is the persistence problem in its purest form: you can re-image the panel, but if the attacker's implant is in the unverified region and you restore from a backup taken after compromise, it comes back. For a GPU operator the practical meaning is that door control - the boundary protecting drives, console ports and the OOB switch inside the cage - can be silently and durably owned, and no amount of log review will show it because the implant controls the logs. This is also a tenant-handoff failure: a panel compromised during one customer's tenancy stays compromised for the next one, since nobody re-flashes access-control hardware between tenants.","attack_vector":"The command-injection path requires an authenticated session to the iSTAR web application, so realistically credential theft, a default or shared integrator account, or chaining from one of the unauthenticated iSTAR issues. The firmware verification gap is then exercised by an attacker who already has that access. All of it lives on the physical-security VLAN.","remediation":"Move to firmware later than 6.9.2 per Johnson Controls' guidance and confirm with the vendor that the verification gap is closed in the version you land on - the disclosure notes later firmware may also be affected, so a version number alone is not evidence. Because the flaw defeats integrity verification, patching is not sufficient for a panel you believe was reachable while vulnerable: replace it or have the vendor perform a full low-level reflash, and rebuild its configuration from a known-good source rather than a device backup. Operationally, rotate every credential on the access-control system, audit the account list for integrator and service accounts, and require MFA on any path to the iSTAR web application.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53696","https://nvd.nist.gov/vuln/detail/CVE-2025-53695","https://www.johnsoncontrols.com/trust-center/cybersecurity/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-64513","cve":"CVE-2025-64513","aliases":[],"title":"Milvus: Unauthenticated attacker exploits the server directly","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Milvus","year":"2025","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Unauthenticated attacker exploits the server directly","attack_vector":"Unauthenticated network to Milvus","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64513"],"status":"curated"},{"id":"CVE-2026-22696","cve":"CVE-2026-22696","aliases":["dcap-qvl QE Identity verification bypass","DCAP quote verification flaw"],"title":"Phala dcap-qvl - the Rust/npm/Python DCAP quote verification library used to verify Intel SGX and TDX attestation quotes; Rust crate before 0.3.9, npm packages @phala/dcap-qvl <=0.3.0 and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phala dcap-qvl - the Rust/npm/Python DCAP quote verification library used to verify Intel SGX and TDX attestation…","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The library fetches Quoting Enclave identity collateral from the provisioning service but never verifies its signature against the certificate chain, and does not enforce MRSIGNER, ISVPRODID or ISVSVN policy on the QE report. An attacker forges QE Identity data to whitelist a Quoting Enclave that is not Intel's, then signs arbitrary quotes that the verifier accepts as genuine. That is total forgery of SGX and TDX remote attestation: a machine can claim to be running an attested confidential workload while running anything at all. For an operator or a customer relying on attestation to prove that model weights only decrypt inside a genuine TDX trust domain, the proof is worthless. This is the highest-scored item in this set and it is a pure software bug in the verifier, not a CPU flaw.","attack_vector":"Whoever controls the machine claiming to be attested, plus the ability to serve or influence the collateral the verifier fetches. No CPU vulnerability, no physical access, no privileged position on the verifier - the verifier simply accepts a forged identity. Anyone running a confidential-compute service whose verification path uses this library is exposed.","remediation":"Pure software update, no firmware and no reboot: upgrade to dcap-qvl 0.3.9 or later (npm @phala/dcap-qvl-node / -web 0.3.4+). Fast and cheap to deploy, which is the good news. The important operator action is inventory: DCAP quote verification is usually buried inside a confidential-computing framework or a TEE-attestation SaaS rather than being a dependency anyone declared deliberately, so audit which library your attestation path actually calls. Any attestation accepted by a vulnerable version before the upgrade should be treated as unverified and re-attested - the fix does not retroactively invalidate quotes you already trusted.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22696","https://github.com/Phala-Network/dcap-qvl/security/advisories/GHSA-796p-j2gh-9m2q"],"status":"curated"},{"id":"CVE-2026-24834","cve":"CVE-2026-24834","aliases":[],"title":"Kata Containers: Kata with Cloud Hypervisor allows a user to break the VM isolation boundary","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"Kata with Cloud Hypervisor allows a user to break the VM isolation boundary","attack_vector":"Any tenant workload running under Kata","remediation":"Emergency Kata upgrade; drain and recreate Kata pods. The whole point of Kata is this boundary, so treat as critical","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24834"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-48797","cve":"CVE-2026-48797","aliases":[],"title":"Backpropagate (single-GPU LLM fine-tuning library) - Reflex web UI: The optional web UI exposes a training control surface without authentication, scored 9.3. On a GPU box this…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Backpropagate (single-GPU LLM fine-tuning library) - Reflex web UI","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"The optional web UI exposes a training control surface without authentication, scored 9.3. On a GPU box this means an unauthenticated caller drives training jobs on your hardware - the practical outcomes are GPU theft for someone else's workload, poisoning of a fine-tune in progress, and read access to whatever the training process can see.","attack_vector":"Network, unauthenticated. Anyone who can reach the UI port. Fine-tuning UIs get bound to 0.0.0.0 during experimentation far more often than anyone admits.","remediation":"Upgrade past 1.1.1, or disable the Reflex web UI entirely if you do not use it. Cost: restart the service. The standing control is a network policy that stops research workloads binding tooling ports to anything but localhost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-48797"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-72329","cve":"CVE-2026-72329","aliases":[],"title":"Linux liquidio driver (Marvell/Cavium, cached VF pci_dev lookup table): TENANT ISOLATION: the LiquidIO PF caches VF `pci_dev` pointers without taking a reference, so the cached…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux liquidio driver (Marvell/Cavium, cached VF pci_dev lookup table)","year":"2026","cvss_score":9.3,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: the LiquidIO PF caches VF `pci_dev` pointers without taking a reference, so the cached pointers can dangle and are then dereferenced when handling a VF function-level-reset request. A VF triggering an FLR — something a tenant does simply by resetting their own device — drives a use-after-free in the host kernel. FLR is the operation an operator relies on to *clean up* between tenants, so the mechanism intended to enforce the handoff boundary is the one that breaks it.","attack_vector":"A tenant holding a LiquidIO VF issuing a function-level reset, or any path that triggers `OCTEON_VF_FLR_REQUEST` handling on the PF.","remediation":"Kernel upgrade plus host reboot on nodes with Marvell LiquidIO adapters. Rolling drain. Note that LiquidIO is end-of-life hardware still present in older inference fleets — if you are running it, weigh replacing the adapters against maintaining kernel patches for a driver that is no longer actively developed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72329"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-47243","cve":"CVE-2026-47243","aliases":[],"title":"Kata Containers: runtime-rs standalone virtio-fs path is vulnerable to a guest-to-host escape","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":9.2,"severity":"critical","kev":false,"impact":"runtime-rs standalone virtio-fs path is vulnerable to a guest-to-host escape","attack_vector":"Any tenant VM under Kata","remediation":"Emergency Kata upgrade; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47243"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-49445","cve":"CVE-2026-49445","aliases":[],"title":"Cilium: With L7 enabled, the embedded Envoy exposes a world-accessible admin.sock on the cluster","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":9.2,"severity":"critical","kev":false,"impact":"With L7 enabled, the embedded Envoy exposes a world-accessible admin.sock on the cluster; full dataplane control","attack_vector":"Any pod on the cluster network","remediation":"Emergency rolling Cilium upgrade; no GPU drain but expect brief dataplane blips per node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49445"],"status":"curated"},{"id":"CVE-2017-16727","cve":"CVE-2017-16727","aliases":["ICSA-17-355-01"],"title":"Moxa NPort W2150A / W2250A wireless device server: The device ships with an empty default password, so anyone who can reach it on the network can log in as an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Moxa NPort W2150A / W2250A wireless device server","year":"2017","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The device ships with an empty default password, so anyone who can reach it on the network can log in as an unauthorized user with no credential at all and take over the serial device server.","attack_vector":"No authentication required — just network reachability to a device still running the default (blank) password.","remediation":"Firmware upgrade past 1.11 plus setting a real administrator password on every device — the firmware fix stops shipping the box in an unauthenticated state, but existing deployed units also need someone to actually set a password during the upgrade. Track this as a fleet-wide credential-rotation project, not just a flash.","references":["http://www.securityfocus.com/bid/102254","https://ics-cert.us-cert.gov/advisories/ICSA-17-355-01"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-6440","cve":"CVE-2018-6440","aliases":[],"title":"Brocade Fabric OS (proxy service information disclosure): Unauthenticated remote attackers can obtain sensitive information from the Fabric OS proxy service. Pre-auth…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS (proxy service information disclosure)","year":"2018","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated remote attackers can obtain sensitive information from the Fabric OS proxy service. Pre-auth information disclosure on a SAN switch typically yields fabric topology and configuration — which is the reconnaissance an attacker needs to know which zone to attack to reach a specific tenant's storage.","attack_vector":"Unauthenticated, remote to the FOS proxy service.","remediation":"Fabric OS upgrade plus reboot per fabric. Immediate control is management-network isolation for all FC switch management interfaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6440"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-16261","cve":"CVE-2019-16261","aliases":[],"title":"Tripp Lite PDUMH15AT / SU750XL PDU: TENANT ISOLATION: the PDU accepts unauthenticated POST requests to its /Forms/ endpoints, which can be used…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tripp Lite PDUMH15AT / SU750XL PDU","year":"2019","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: the PDU accepts unauthenticated POST requests to its /Forms/ endpoints, which can be used to change the manager or admin password, or directly shut off power to an outlet. Anyone who can reach the PDU's web interface can cut power to whatever rack that outlet feeds, with no login at all.","attack_vector":"Fully remote and unauthenticated — a crafted POST request to the /Forms/ directory is all that's needed to flip an outlet or take over the admin account.","remediation":"Software/firmware upgrade — Tripp Lite (now Eaton) shipped a fixed release after this was reported; confirm every deployed unit is past 12.04.0053 (PDUMH15AT) / 12.04.0052 (SU750XL). Flash each PDU; this directly powers a rack, so schedule around a window where a brief power-monitoring interruption is acceptable, and audit for units still on vulnerable firmware since these are frequently forgotten 'fringe' infrastructure.","references":["https://blog.korelogic.com/blog/2019/08/19/unpatched_fringe_infrastructure_bits"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-4169","cve":"CVE-2019-4169","aliases":["IBM X-Force 158702"],"title":"IBM OpenPower firmware OP910/OP920 - OpenBMC IPMI credential handling: The original default BMC password kept working over IPMI after an operator changed it. Every hardening…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenPower firmware OP910/OP920 - OpenBMC IPMI credential handling","year":"2019","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The original default BMC password kept working over IPMI after an operator changed it. Every hardening runbook says 'change the default BMC password' and on these firmware levels doing so accomplished nothing for the IPMI path - the fleet stayed openable with a credential printed in the vendor documentation. This is the cleanest example in the cluster of why BMC posture cannot be assessed from configuration intent: you have to test that the old credential is actually dead.","attack_vector":"Network access to the BMC's IPMI interface with the publicly documented default credential. No prior foothold required.","remediation":"Fixed in later OpenPower firmware; delivery is a per-node system firmware update with a maintenance window. The transferable lesson is a test, not a patch: after any BMC credential rotation, actively attempt an IPMI and Redfish login with the old and default credentials and alert if either succeeds. Make that a recurring fleet check, not a one-off - it catches this class of bug on any vendor.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-4169","https://exchange.xforce.ibmcloud.com/vulnerabilities/158702"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-4926","cve":"CVE-2020-4926","aliases":[],"title":"IBM Spectrum Scale 5.1 core / IBM Elastic Storage System 6.1: TENANT ISOLATION: unauthorized access to user data, or injection of arbitrary data into the communication…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale 5.1 core / IBM Elastic Storage System 6.1","year":"2020","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: unauthorized access to user data, or injection of arbitrary data into the communication between cluster nodes. This is the core GPFS daemon protocol, not an add-on layer — an attacker positioned on the storage cluster network can read other tenants' data or write data that nodes accept as legitimate. For a shared training filesystem, data injection is also a training-data poisoning vector.","attack_vector":"An attacker with access to the inter-node communication path of the Spectrum Scale cluster — the back-end storage network.","remediation":"Upgrade Spectrum Scale / ESS to a fixed level. This is a coordinated cluster upgrade; GPFS supports rolling node upgrades but the version-compatibility window means planning, and a full-cluster restart is sometimes unavoidable. Enable and verify GPFS cluster-level authentication and encryption in transit, which is a config change and the thing that actually removes the exposure.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-4926"],"status":"curated"},{"id":"CVE-2021-1577","cve":"CVE-2021-1577","aliases":[],"title":"Cisco APIC / Cloud APIC (API endpoint): Unauthenticated arbitrary file read and write on the APIC — the ACI fabric controller. File write on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco APIC / Cloud APIC (API endpoint)","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated arbitrary file read and write on the APIC — the ACI fabric controller. File write on the controller is effectively fabric takeover: you can plant credentials, alter policy state, and reprogram forwarding for every tenant on the pod.","attack_vector":"Unauthenticated, remote to the APIC's API endpoint. No credentials.","remediation":"APIC software upgrade across the controller cluster (rolling, one APIC at a time, fabric keeps forwarding). Restrict APIC API reachability to a dedicated management network as a durable control — a config change you should make regardless of this CVE.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1577"],"status":"curated"},{"id":"CVE-2021-22794","cve":"CVE-2021-22794","aliases":["SEVD-2022-095-01"],"title":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior: Path traversal to remote code execution on the DCIM appliance. DCE is the aggregation point for a site's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Path traversal to remote code execution on the DCIM appliance. DCE is the aggregation point for a site's entire power and cooling estate - it holds working credentials for every UPS, PDU, NMC, CRAC and sensor it polls. Compromise here is not one device; it is the keys to the whole facility layer, and from there PHYSICAL control of power and cooling.","attack_vector":"Network access to the DCE appliance. DCE typically sits on the facility management network with broad reachability by design, since it must poll everything.","remediation":"Upgrade DCE to v7.9.0 or later. Appliance upgrade with a service restart - hours, not a power maintenance window. Afterwards, rotate every device credential DCE stored, because they were all reachable. Long term, DCE should be the most tightly segmented host you run: inbound access from a jump host only, no internet egress.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-22794"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2021-22795","cve":"CVE-2021-22795","aliases":["SEVD-2022-095-01"],"title":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior: OS command injection over the network on the DCIM appliance - same blast radius as the path traversal above…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric StruxureWare Data Center Expert (DCE) v7.8.1 and prior","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"OS command injection over the network on the DCIM appliance - same blast radius as the path traversal above, reached by a different route. An attacker running commands on DCE inherits its trust relationship with every piece of power and cooling gear at the site.","attack_vector":"Remote, over the network to the DCE appliance.","remediation":"Upgrade to DCE v7.9.0 or later and rotate all stored device credentials. If the appliance was reachable from an untrusted network, rebuild it - DCE keeps polling credentials in a form an attacker with shell can read.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-22795"],"status":"curated"},{"id":"CVE-2021-26731","cve":"CVE-2021-26731","aliases":[],"title":"Lanner IAC-AST2500A BMC firmware: An authenticated BMC user escalates to root code execution on the controller. It is included alongside its…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC firmware","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"An authenticated BMC user escalates to root code execution on the controller. It is included alongside its unauthenticated siblings because it closes a different door: an operator who mitigates the unauthenticated bugs by putting the BMC behind a bastion still has this one live for anyone holding a valid BMC credential, including monitoring accounts and any credential shared across the fleet. The outcome is the same - out-of-band power, console, virtual media and firmware-level persistence. Command injection and stack buffer overflows in the modifyUserb_func handler of spx_restservice, reachable after authentication through the user-modification path.","attack_vector":"An authenticated attacker reaching the BMC REST service. Given how commonly BMC credentials are shared across a whitebox fleet, one leaked password is fleet-wide reach.","remediation":"Firmware flash from Lanner or the board integrator, with the same sourcing difficulty as the rest of this cluster. Config-only actions that help now: rotate BMC credentials to per-node unique values so a single leak is not fleet-wide, remove any BMC accounts issued to tenants or third parties, and keep the BMC REST service off any network segment reachable from tenant workloads. Note that this is one of at least nine published CVEs against this one BMC module's REST service, which is itself the signal - the codebase was not audited before shipping.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26731","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-26731/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-28506","cve":"CVE-2021-28506","aliases":[],"title":"Arista EOS (gNOI): gNOI APIs bypass authentication, allowing an unauthenticated factory reset of the switch — instant…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (gNOI)","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"gNOI APIs bypass authentication, allowing an unauthenticated factory reset of the switch — instant fabric-wide outage primitive","attack_vector":"Network","remediation":"EOS upgrade; interim mitigation is a service ACL restricting gNOI, though CVE-2021-28507 shows those ACLs were themselves bypassable","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28506"],"status":"curated"},{"id":"CVE-2021-47348","cve":"CVE-2021-47348","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Memory is handed to a consumer without being initialised or cleared in the amdgpu display core (DC/DM).…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu display core (DC/DM). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amd/display: Avoid HDCP over-read and corruption","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47348","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-0670","cve":"CVE-2022-0670","aliases":[],"title":"Ceph Manager (volumes plugin): Owner of one CephFS share can read/write any share or the entire file system - cross-tenant","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph Manager (volumes plugin)","year":"2022","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Owner of one CephFS share can read/write any share or the entire file system - cross-tenant","attack_vector":"Network (remote)","remediation":"Data-plane: ceph-mgr upgrade across the cluster; audit CephFS share access","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0670"],"status":"curated"},{"id":"CVE-2022-0715","cve":"CVE-2022-0715","aliases":["TLStorm"],"title":"APC Smart-UPS SMT/SMC/SMX/SCL/SMTL series - firmware update signing: PHYSICAL and PERSISTENT. Firmware images are signed with a key that leaked, so an attacker can flash…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Smart-UPS SMT/SMC/SMX/SCL/SMTL series - firmware update signing","year":"2022","cvss_score":9.1,"severity":"critical","kev":false,"impact":"PHYSICAL and PERSISTENT. Firmware images are signed with a key that leaked, so an attacker can flash arbitrary firmware onto the UPS and have it accepted as genuine. This is the worst of the three TLStorm bugs for an operator: the implant lives in the power device, survives every reimage of every server behind it, and is invisible to anything you run on the compute plane. An attacker who owns the UPS owns a kill switch for the racks it feeds, on a timer of their choosing.","attack_vector":"Network access to the UPS - either directly on the management network or via the intercepted cloud channel. Also reachable by anyone who can push a firmware image through the vendor's own update path.","remediation":"Flash to fixed firmware, which changes the signing scheme. Because the compromise persists in firmware, patching alone does not prove cleanliness on a unit you believe was targeted - the honest answer there is re-flash from a known-good image and verify the reported firmware version out of band. Treat UPS firmware as part of your supply chain: version-inventory it, and refuse units that cannot report a verifiable firmware version.","references":["https://www.se.com/ww/en/download/document/SEVD-2022-067-02/","https://www.armis.com/research/tlstorm/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-23131","cve":"CVE-2022-23131","aliases":[],"title":"Zabbix: Unverified user login in session data (SAML SSO enabled)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2022","cvss_score":9.1,"severity":"critical","kev":true,"impact":"[KEV] Unverified user login in session data (SAML SSO enabled) -> unauthenticated admin takeover","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate the Zabbix session secret","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23131"],"status":"curated"},{"id":"CVE-2022-31247","cve":"CVE-2022-31247","aliases":[],"title":"Rancher: Anyone who can create role template bindings escalates privileges cluster-wide","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Anyone who can create role template bindings escalates privileges cluster-wide","attack_vector":"Cluster user with cluster-owner or project-owner rights","remediation":"Upgrade Rancher; audit RoleTemplateBindings","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31247"],"status":"curated"},{"id":"CVE-2023-23947","cve":"CVE-2023-23947","aliases":[],"title":"Argo CD: Improper authorization lets a user modify resources outside their permitted projects","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Improper authorization lets a user modify resources outside their permitted projects","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-23947"],"status":"curated"},{"id":"CVE-2023-25132","cve":"CVE-2023-25132","aliases":["JVN#95119483"],"title":"CyberPower PowerPanel Business - default.cmd file upload: Unrestricted upload of a dangerous file type into default.cmd - the script PowerPanel runs on a power event.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Business - default.cmd file upload","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unrestricted upload of a dangerous file type into default.cmd - the script PowerPanel runs on a power event. Same nasty shape as the PowerChute issue: the attacker's code runs at the moment the UPS signals, across every host the software controls, with elevated privilege. The trigger is a power event, which an attacker with UPS access can also cause.","attack_vector":"An attacker who can write to the PowerPanel Business installation - reachable via the default-credential issue in the same advisory.","remediation":"Upgrade past v4.8.6. Independently: shutdown scripts invoked by power-management software are privileged code and should be under change control with restricted write permissions, on every host, regardless of vendor.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25132"],"status":"curated"},{"id":"CVE-2023-25725","cve":"CVE-2023-25725","aliases":[],"title":"HAProxy (before 2.7.3): HAProxy's HTTP/1 header parser accepts empty header field names, which can be used to make legitimate headers…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HAProxy (before 2.7.3)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"HAProxy's HTTP/1 header parser accepts empty header field names, which can be used to make legitimate headers silently disappear after parsing. An attacker can use this to smuggle requests past access-control rules that were supposed to inspect those headers — bypassing ACLs meant to keep unauthorized traffic away from backend inference/storage services.","attack_vector":"Remote — an attacker sends a crafted HTTP/1 request with an empty header field name to a HAProxy instance doing header-based ACL enforcement.","remediation":"Software upgrade to HAProxy 2.7.3 or later, then reload/restart the process. HAProxy typically runs as a software component rather than an appliance, so this is a package upgrade + service restart across whichever hosts run it in front of the cluster; a graceful reload avoids dropping in-flight connections if your HAProxy version supports it.","references":["https://git.haproxy.org/?p=haproxy-2.7.git%3Ba=commit%3Bh=a0e561ad7f29ed50c473f5a9da664267b60d1112","https://www.haproxy.org/"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2023-28863","cve":"CVE-2023-28863","aliases":[],"title":"AMI MegaRAC SPx12/SPx13: Insufficient verification of data authenticity — firmware image signature can be subverted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx12/SPx13","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Insufficient verification of data authenticity — firmware image signature can be subverted; enables persistent implant","attack_vector":"Local/network firmware update path","remediation":"BMC flash; also requires operational control so that only signed, vendor-verified images reach the update endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28863"],"status":"curated"},{"id":"CVE-2023-31127","cve":"CVE-2023-31127","aliases":["libspdm session establishment bypass"],"title":"DMTF libspdm - SPDM session establishment (reference implementation used in GPU/device attestation): MULTI-TENANT ISOLATION: a device supporting both DHE and PSK sessions with mutual…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DMTF libspdm - SPDM session establishment (reference implementation used in GPU/device attestation)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: a device supporting both DHE and PSK sessions with mutual authentication can be driven into establishing a session with KEY_EXCHANGE plus PSK_FINISH, bypassing mutual authentication entirely. Scored 9.1 with a changed scope. This matters far beyond libspdm itself: SPDM is the protocol underneath device attestation and the encrypted host-to-device channel in confidential GPU computing and in PCIe/CXL IDE, and libspdm is the reference code many silicon and firmware vendors derived from. If your accelerator's attestation stack forked libspdm before 2.3.1, ask the vendor directly.","attack_vector":"An attacker on an adjacent path with low privileges - in the device context, something able to speak SPDM to the responder, which means the platform firmware, a DPU, or a compromised device on the link.","remediation":"Update to libspdm 2.3.1 or later. Cost for an operator is indirect and slow: you do not deploy libspdm, your GPU/DPU/NIC firmware vendor does, so the action is a supply-chain question to each vendor and then a firmware update on their schedule. Firmware flash means node drain. Track it as a vendor-management item rather than a patch you can apply.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-qw76-4v8p-xq9f","https://nvd.nist.gov/vuln/detail/CVE-2023-31127"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-3267","cve":"CVE-2023-3267","aliases":["ZDI-23-1149"],"title":"CyberPower PowerPanel Enterprise DCIM - remote backup location username field: OS command injection through the remote-backup username field, executing as NT AUTHORITY\\SYSTEM. Chained…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM - remote backup location username field","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"OS command injection through the remote-backup username field, executing as NT AUTHORITY\\SYSTEM. Chained behind either authentication bypass above, this is unauthenticated SYSTEM on the host that manages your UPS estate.","attack_vector":"Authenticated user adding a remote backup location - trivially reachable given the two auth bypasses in the same advisory batch.","remediation":"Upgrade PowerPanel Enterprise. If the host was reachable from an untrusted network, rebuild it rather than patch - SYSTEM-level RCE on a Windows DCIM host means credential theft from the whole box.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3267"],"status":"curated"},{"id":"CVE-2023-34329","cve":"CVE-2023-34329","aliases":[],"title":"AMI MegaRAC SPx12 (BMC&C): Auth bypass by spoofing the HTTP header","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx12 (BMC&C)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Auth bypass by spoofing the HTTP header; combined with CVE-2023-34330 gives unauthenticated RCE as root on the BMC. Below-OS persistence","attack_vector":"Network, BMC web/Redfish","remediation":"BMC firmware update per node via ODM build; combined-chain risk means treat as critical even where the BMC is on a management VLAN","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34329"],"status":"curated","fleet":{"ubiquity":"universal - same MegaRAC OEM footprint","remediation_pain":"firmware-flash - full BMC image update per node, out-of-band, cannot run while a tenant job holds the host","pain_class":"firmware-flash","why_fleet_wide":"Chaining the HTTP-header auth spoof with the code-injection primitive yields pre-OS code execution on the BMC; identical firmware across a homogeneous GPU fleet means one exploit works on every node."}},{"id":"CVE-2023-3961","cve":"CVE-2023-3961","aliases":[],"title":"Samba: Path traversal in client pipe names","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Path traversal in client pipe names -> connect to Unix sockets outside the private directory","attack_vector":"Network (remote)","remediation":"Data-plane: smbd upgrade on file-serving nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3961"],"status":"curated"},{"id":"CVE-2023-48023","cve":"CVE-2023-48023","aliases":[],"title":"Ray (`/log_proxy`): SSRF from the dashboard","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (`/log_proxy`)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"SSRF from the dashboard","attack_vector":"Unauthenticated network to the dashboard","remediation":"No vendor patch (same disputed position); isolate","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48023"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-51786","cve":"CVE-2023-51786","aliases":[],"title":"Lustre (incorrect access control, 2.13.x-2.15.x before 2.15.4): TENANT ISOLATION: incorrect access control in Lustre lets an attacker escalate privileges and obtain…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre (incorrect access control, 2.13.x-2.15.x before 2.15.4)","year":"2023","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: incorrect access control in Lustre lets an attacker escalate privileges and obtain sensitive information. This is the modern-release equivalent of the 2019 family and it lands on the versions actually deployed in current AI-training clusters — 2.15.x is the long-term-support line most sites run today. On a shared training filesystem, 'obtain sensitive information' means another tenant's datasets, checkpoints and model weights.","attack_vector":"An attacker with Lustre client access on versions 2.13.x, 2.14.x, or 2.15.x before 2.15.4.","remediation":"Upgrade to Lustre 2.15.4 or later — client and server packages, with a coordinated restart of the storage cluster. Enable Lustre nodemap with `admin`/`trusted` set to off for tenant clients and squash root, so a compromised client cannot act as root against the filesystem; that is a config change you can make ahead of the upgrade and it is the durable control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-51786"],"status":"curated"},{"id":"CVE-2024-12378","cve":"CVE-2024-12378","aliases":["Arista Security Advisory 0116"],"title":"Arista EOS (secure VXLAN / Tunnelsec agent): TENANT ISOLATION: after the Tunnelsec agent restarts, traffic that should be encrypted inside secure VXLAN…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (secure VXLAN / Tunnelsec agent)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: after the Tunnelsec agent restarts, traffic that should be encrypted inside secure VXLAN tunnels goes out in cleartext. You configured tunnel encryption between sites or between pods, the agent bounced, and now every tenant's overlay traffic is readable by anyone with a tap on the underlay — with no alarm, because the tunnels still look up. This is the worst kind of fabric bug: the control plane reports healthy while the confidentiality guarantee is silently gone.","attack_vector":"Passive attacker with access to the underlay path the VXLAN tunnels traverse — a dark-fibre tap, a transit provider, a compromised intermediate switch, or another tenant with underlay visibility. Requires the Tunnelsec agent to have restarted, which happens on upgrade, crash, or config change.","remediation":"EOS upgrade plus a switch reload on every VTEP running secure VXLAN. Until then, monitor Tunnelsec agent restarts and treat any restart as a confidentiality incident for traffic since that moment. There is no config workaround that keeps encryption on across a restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-12378"],"status":"curated"},{"id":"CVE-2024-21887","cve":"CVE-2024-21887","aliases":[],"title":"Ivanti Connect Secure: Command injection in web components","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure","year":"2024","cvss_score":9.1,"severity":"critical","kev":true,"impact":"[KEV] Command injection in web components; chained with CVE-2023-46805 for unauthenticated root RCE","attack_vector":"Network (remote)","remediation":"Control-plane: rebuild from a factory image; rotate every VPN secret and certificate","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21887"],"status":"curated"},{"id":"CVE-2024-22120","cve":"CVE-2024-22120","aliases":[],"title":"Zabbix: Unsanitized clientip in the audit log","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unsanitized clientip in the audit log -> time-based blind SQL injection around script execution","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate stored host and IPMI credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22120"],"status":"curated"},{"id":"CVE-2024-32752","cve":"CVE-2024-32752","aliases":["CVE-2017-17704","ICSA-24-158-04"],"title":"Software House iSTAR door controllers (firmware before 6.6.B) and the IP-ACM Ethernet Door Module link: The iSTAR controllers do not support authenticated communications with the iSTAR…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Software House iSTAR door controllers (firmware before 6.6.B) and the IP-ACM Ethernet Door Module link","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The iSTAR controllers do not support authenticated communications with the iSTAR Configuration Utility, and the older IP-ACM door-module link used a fixed AES key and a fixed IV restarted on every message - which leaks enough structure that door-unlock commands can be replayed or forged outright. Both bugs land in the same place: the wire between the controller and its door hardware or configuration tool is not trustworthy, so an attacker on that network can open doors without ever authenticating to anything. Standing in the hall or the cage, that person reaches drives, console ports, the out-of-band management switch and the physical console of machines running customer workloads. For a multi-tenant GPU operator this collapses the cage boundary that the whole bare-metal isolation story rests on, and it does so without generating a failed-authentication event anywhere, because there is no authentication to fail.","attack_vector":"A host on the physical-security network that can see traffic between the iSTAR controller and the IP-ACM modules or the configuration utility. That means passive capture and replay, or active injection, from the security VLAN - typically shared with CCTV and the integrator's tooling. Physical access to structured cabling in a back-of-house space is an equally valid vector and is not covered by most tenants' security policies because the cabling is the landlord's.","remediation":"Update iSTAR controller firmware to 6.6.B or later, which introduces authenticated ICU communications, and retire the affected IP-ACM generation in favour of the IP-ACM v2 hardware that supports proper encryption. That is firmware plus a hardware refresh of the door modules - a capital project with a security integrator and per-door downtime, not a patch. Until it is done: physically secure every cable run and panel between the controller and the door modules, put the physical-security VLAN behind a firewall with an allow-list, and enable 802.1X or port security on the switch ports serving door hardware so a rogue device cannot join the segment. Treat any hall where door-module cabling runs through space you do not control as a hall with a broken cage boundary.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-24-158-04","https://nvd.nist.gov/vuln/detail/CVE-2024-32752","https://nvd.nist.gov/vuln/detail/CVE-2017-17704","https://systemoverlord.com/2017/12/18/cve-2017-17704-broken-cryptography-in-istar-ultra-ip-acm-by-software-house.html"],"status":"curated"},{"id":"CVE-2024-3411","cve":"CVE-2024-3411","aliases":["VU#163057"],"title":"The IPMI 2.0 authenticated-session mechanism as specified and as implemented across multiple vendors: An attacker hijacks an established IPMI session by spoofing packets with a predicted session…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"The IPMI 2.0 authenticated-session mechanism as specified and as implemented across multiple vendors","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"An attacker hijacks an established IPMI session by spoofing packets with a predicted session ID, bypassing authentication entirely. What they inherit is the privilege of whoever's session they stole - typically an operator or automation account with power control, boot-device selection, sensor access and often virtual media. This is a spec-level weakness rather than one vendor's bug, so it is one of the few entries here that an operator should assume applies to their whole heterogeneous fleet, not just the boards from one manufacturer. Session IDs are predictable and BMC random number generation is weak, so the values that are supposed to make a session unforgeable are guessable. CERT/CC tracks it as VU#163057 and the affected list spans vendor BMC implementations built to the Intel IPMI specification.","attack_vector":"Network reachability to the BMC's IPMI-over-LAN port (UDP 623) with the ability to observe or infer session state. Unauthenticated in effect, since the whole point is that authentication is bypassed. Anything on the out-of-band management VLAN qualifies, as does anything that can reach a BMC exposed by a routing mistake.","remediation":"Firmware updates exist from individual vendors, but there is no single fix because the weakness is in the protocol's assumptions - so the durable remediation is to stop using IPMI-over-LAN. Disable IPMI-over-LAN on the BMC and drive management through Redfish over TLS instead; that is a config-only change on most modern BMCs and is the single highest-value action in this entire database for an operator who still has UDP 623 open. Where legacy tooling forces IPMI, restrict UDP 623 to an explicit allowlist of management hosts at the switch, and treat the management VLAN as a network where session hijacking is assumed possible.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3411","https://kb.cert.org/vuls/id/163057"],"status":"curated"},{"id":"CVE-2024-37287","cve":"CVE-2024-37287","aliases":[],"title":"Kibana: Prototype pollution via ML/Alerting connectors + write access to internal ML indices","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kibana","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Prototype pollution via ML/Alerting connectors + write access to internal ML indices -> arbitrary code execution","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the observability UI tier; restrict ML roles","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37287"],"status":"curated"},{"id":"CVE-2024-3829","cve":"CVE-2024-3829","aliases":[],"title":"Qdrant (snapshot recovery): Arbitrary file read and write during snapshot recovery","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Qdrant (snapshot recovery)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Arbitrary file read and write during snapshot recovery","attack_vector":"Customer-supplied snapshot file","remediation":"Upgrade; snapshots must be treated as untrusted archives","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3829"],"status":"curated"},{"id":"CVE-2024-45763","cve":"CVE-2024-45763","aliases":["DSA-2024-449"],"title":"Dell Enterprise SONiC (OS command injection): OS command injection giving arbitrary command execution on the switch's underlying Linux. Chained behind the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (OS command injection)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"OS command injection giving arbitrary command execution on the switch's underlying Linux. Chained behind the authentication-bypass in the same advisory batch, an unauthenticated attacker goes from the network to root on the switch in two steps. On SONiC the 'switch' is a fairly normal Linux box with the ASIC SDK attached, so root there means arbitrary forwarding-table manipulation for every tenant on the device.","attack_vector":"Remote attacker holding high-privilege access — but see CVE-2024-45764, which supplies that access without credentials.","remediation":"Same fix as the rest of DSA-2024-449: NOS image upgrade and reboot on every Dell Enterprise SONiC switch running 4.1.x or 4.2.x. Assume any device that was reachable pre-patch may hold persistence and consider a clean re-image rather than an in-place upgrade.","references":["https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2024-45763"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45765","cve":"CVE-2024-45765","aliases":["DSA-2024-449"],"title":"Dell Enterprise SONiC (privilege boundary in CLI): High-privilege OS commands can be run by users holding less privileged roles. Whatever role separation you…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (privilege boundary in CLI)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"High-privilege OS commands can be run by users holding less privileged roles. Whatever role separation you built for your NOC — read-only, operator, admin — does not hold on the switch. Relevant to any operator who gives tenants or contractors scoped switch access.","attack_vector":"Authenticated user with a low-privilege SONiC role.","remediation":"NOS image upgrade and reboot per switch. Until then, do not issue low-privilege SONiC accounts to anyone you would not give admin, because the distinction is not enforced.","references":["https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2024-45765"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-7776","cve":"CVE-2024-7776","aliases":[],"title":"ONNX (`download_model`): Arbitrary file overwrite","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`download_model`)","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Arbitrary file overwrite","attack_vector":"Customer-supplied model reference","remediation":"Upgrade past 1.16.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7776"],"status":"curated"},{"id":"CVE-2024-9487","cve":"CVE-2024-9487","aliases":[],"title":"GitHub Enterprise Server: Improper signature verification","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"GitHub Enterprise Server","year":"2024","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Improper signature verification -> SAML SSO bypass, unauthorized user provisioning and instance access","attack_vector":"Network (remote)","remediation":"Control-plane: GHES upgrade; audit newly created users","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9487"],"status":"curated"},{"id":"CVE-2025-0108","cve":"CVE-2025-0108","aliases":[],"title":"Palo Alto PAN-OS: Management web interface auth bypass invoking PHP scripts","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2025","cvss_score":9.1,"severity":"critical","kev":true,"impact":"[KEV] Management web interface auth bypass invoking PHP scripts; actively exploited","attack_vector":"Network (remote)","remediation":"Control-plane: patch; management plane on an out-of-band network only","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0108"],"status":"curated"},{"id":"CVE-2025-10263","cve":"CVE-2025-10263","aliases":["TFV-17","AMP-SB-0008","XSA-493","TLBI+DSB completes too early"],"title":"Arm Neoverse N1 / N2 / V1 / V2 / V3 / V3AE, Cortex-A76/A77/A78/A710, Cortex-X1-X925, C1-Ultra/Premium; Trusted Firmware-A up to and including v2.15; Ampere Altra and Altra Max: A store issued on…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Neoverse N1 / N2 / V1 / V2 / V3 / V3AE, Cortex-A76/A77/A78/A710, Cortex-X1-X925, C1-Ultra/Premium; Trusted…","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A store issued on one core can land after another core has already invalidated the translation and completed its TLBI+DSB sequence. The write then commits through a stale translation into memory owned by a higher exception level - a guest writing into hypervisor memory, or a hypervisor writing into EL3/secure memory. In practice that is a write primitive into page tables and other memory-management structures you assumed were fenced off. Neoverse N1 is Ampere Altra and Graviton2; V1/V2 are Graviton3/Graviton4 and Grace. This is close to a worst case for a shared Arm host.","attack_vector":"Code running in a guest or in the host kernel on an affected Arm core, racing another core's TLB maintenance. Purely local, no device or network access needed, but it is a race, so exploitation needs control of scheduling on at least two cores - trivially satisfied by any tenant with more than one vCPU.","remediation":"Requires coordinated firmware and OS updates, and you need both halves. TF-A platforms must build with WORKAROUND_CVE_2025_10263 enabled, which makes the affected TLBI+DSB sequences execute twice; the hypervisor and guest kernels need the matching arm64 errata workaround (Xen shipped XSA-493, Red Hat shipped a long errata chain). Flash + reboot + job drain on every node, and the OEM has to ship the BL31 build first. Ampere published AMP-SB-0008 for Altra and Altra Max. Expect a measurable cost on TLB-shootdown-heavy workloads because the barrier sequence is now doubled.","references":["https://developer.arm.com/documentation/112137","https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-17.html","http://xenbits.xen.org/xsa/advisory-493.html","https://amperecomputing.com/products/product-security"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-1260","cve":"CVE-2025-1260","aliases":["Arista Security Advisory 21098"],"title":"Arista EOS (OpenConfig gNOI authorization): The gNOI equivalent of the gNMI authorization bypass: operations that should have been rejected run anyway.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (OpenConfig gNOI authorization)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"The gNOI equivalent of the gNMI authorization bypass: operations that should have been rejected run anyway. gNOI covers reboot, image install, certificate rotation and factory reset — so an under-privileged caller can reload switches or push images, not merely edit config.","attack_vector":"A client reaching the gNOI endpoint on a switch with OpenConfig configured, holding credentials that should not authorize the operation.","remediation":"EOS upgrade plus reload. Immediately restrict gNOI endpoint reachability by ACL and re-issue any certificates that could have been rotated by an unauthorized caller.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1260"],"status":"curated"},{"id":"CVE-2025-15031","cve":"CVE-2025-15031","aliases":[],"title":"MLflow (pyfunc tar extraction): Arbitrary file write from crafted tar entries","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (pyfunc tar extraction)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Arbitrary file write from crafted tar entries","attack_vector":"Customer-supplied model archive","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-15031"],"status":"curated"},{"id":"CVE-2025-23317","cve":"CVE-2025-23317","aliases":[],"title":"NVIDIA Triton (HTTP server): Attacker can start a reverse shell from the HTTP server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (HTTP server)","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Attacker can start a reverse shell from the HTTP server","attack_vector":"Unauthenticated network to the HTTP port","remediation":"Patch. Chains with CVE-2025-23319/23320 into full takeover of the inference host","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23317"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-122"]},{"id":"CVE-2025-54576","cve":"CVE-2025-54576","aliases":[],"title":"OAuth2-Proxy: skip_auth_routes route matching flaw","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OAuth2-Proxy","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"skip_auth_routes route matching flaw -> authentication bypass on protected paths","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the ingress auth sidecar; re-audit every skip rule","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54576"],"status":"curated"},{"id":"CVE-2025-6000","cve":"CVE-2025-6000","aliases":[],"title":"HashiCorp Vault: Root-namespace operator with write on sys/audit gains code execution on the Vault host via the plugin…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Root-namespace operator with write on sys/audit gains code execution on the Vault host via the plugin directory","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT - Vault holds tenant and BMC creds; upgrade + rotate root token","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-6000"],"status":"curated"},{"id":"CVE-2025-67039","cve":"CVE-2025-67039","aliases":["ICSA-26-069-02"],"title":"Lantronix EDS3000PS serial-to-Ethernet device server: Full bypass of the management-page login. Appending a specific suffix to a management URL, combined with a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix EDS3000PS serial-to-Ethernet device server","year":"2025","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Full bypass of the management-page login. Appending a specific suffix to a management URL, combined with a crafted Authorization header, gets an attacker straight into admin functionality with no valid credentials — including whatever serial console sessions the device is bridging.","attack_vector":"Purely network-based, no credentials required — the attacker just needs to reach the device's web management port and knows the URL/header trick published in the advisory.","remediation":"Firmware flash to the fixed release; this is an auth-check logic bug, not something you can compensate for with a password change. Roll out per device; each flash briefly interrupts the serial bridging that device provides.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-069-02"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-0257","cve":"CVE-2026-0257","aliases":[],"title":"Palo Alto PAN-OS: GlobalProtect portal/gateway auth bypass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2026","cvss_score":9.1,"severity":"critical","kev":true,"impact":"[KEV] GlobalProtect portal/gateway auth bypass -> establish an unauthorized VPN connection","attack_vector":"Network (remote)","remediation":"Control-plane: patch; review VPN session and tunnel logs for rogue connections","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-0257"],"status":"curated"},{"id":"CVE-2026-14890","cve":"CVE-2026-14890","aliases":[],"title":"SGLang (expert-parallel backup ZMQ PULL): Unauthenticated, unvalidated deserialization on a routable interface","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (expert-parallel backup ZMQ PULL)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated, unvalidated deserialization on a routable interface","attack_vector":"Co-tenant on the cluster fabric","remediation":"Upgrade; the EP backup plane binds routable by default","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-14890"],"status":"curated"},{"id":"CVE-2026-35030","cve":"CVE-2026-35030","aliases":[],"title":"LiteLLM (JWT auth): Auth bypass when `enable_jwt_auth` is set","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (JWT auth)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Auth bypass when `enable_jwt_auth` is set","attack_vector":"Unauthenticated network","remediation":"Upgrade to 1.83.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35030"],"status":"curated"},{"id":"CVE-2026-41475","cve":"CVE-2026-41475","aliases":["CVE-2023-51773","CVE-2026-26264","CVE-2025-66624","CVE-2026-41502","CVE-2026-41503","CVE-2026-21878","CVE-2018-10238"],"title":"BACnet Stack open-source C library (bacnet-stack) embedded in third-party controllers and gateways: A run of out-of-bounds reads and length underflows in the decoders for WritePropertyMultiple…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BACnet Stack open-source C library (bacnet-stack) embedded in third-party controllers and gateways","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"A run of out-of-bounds reads and length underflows in the decoders for WritePropertyMultiple, ReadPropertyMultiple, WriteProperty and NPDU handling, all reachable by an unauthenticated attacker sending a truncated or malformed request. The reason this matters more than the individual CVSS scores suggest is supply chain: bacnet-stack is the reference C implementation that a long tail of controller, gateway and sensor vendors embed in their firmware without ever telling the customer. Your CRAH controller, your BACnet router, your rear-door heat exchanger's BACnet interface and your environmental gateway may all be running the same library, and none of them appear in a search for 'bacnet-stack'. One crafted packet can therefore crash a heterogeneous set of devices simultaneously - a fleet-wide loss of thermal control in a hall where 40-140 kW racks have minutes of margin. Earlier issues in the same library reached memory corruption rather than just reads, so treat the class as potentially more than DoS on older embedded builds.","attack_vector":"Unauthenticated BACnet/IP on the facility network. No credentials, no interaction, and in several of these cases the trigger is a single truncated request. Broadcast-reachable services widen this further. The practical problem is that you cannot enumerate affected devices from the outside - you need each vendor to disclose whether they embed the library and at what version, and most will not answer quickly.","remediation":"You cannot patch this yourself in the general case. The library fix is upstream (bacnet-stack 1.4.3 / 1.5.0 and later), but the code is compiled into vendor firmware, so remediation means each device vendor rebuilding and shipping firmware, then a per-device flash by the controls contractor. Expect that most embedded devices in your hall will never get a fixed build. That makes segmentation the actual answer: BACnet on an isolated VLAN, no untrusted hosts on it, no BBMD bridging to anything else. In parallel, use this as a procurement lever - make an SBOM for the BACnet stack a requirement in new controller purchases, because right now operators have no way to answer 'am I affected' and that is the real finding.","references":["https://github.com/bacnet-stack/bacnet-stack/security/advisories/GHSA-cvv4-v3g6-4jmv","https://nvd.nist.gov/vuln/detail/CVE-2026-41475","https://nvd.nist.gov/vuln/detail/CVE-2023-51773","https://github.com/bacnet-stack/bacnet-stack/security/advisories/GHSA-phjh-v45p-gmjj"],"status":"curated"},{"id":"CVE-2026-46043","cve":"CVE-2026-46043","aliases":["RDMA/rxe payload_size underflow","Soft-RoCE BTH pad validation"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) receive path, drivers/infiniband/sw/rxe/rxe_recv.c: FABRIC DOS: rxe_rcv() checked only that an inbound packet was at least header_size bytes, but payload_size()…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) receive path, drivers/infiniband/sw/rxe/rxe_recv.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"FABRIC DOS: rxe_rcv() checked only that an inbound packet was at least header_size bytes, but payload_size() then subtracts the attacker-controlled BTH pad field and the ICRC size from the packet length. A short packet, or one carrying a forged non-zero pad, makes that subtraction underflow and hands a bogus length to everything downstream in the receive path. Since Soft-RoCE rides UDP/4791, a single unauthenticated datagram from anywhere that can reach the node crashes or corrupts the kernel - no connection, no handshake, no credentials. On a shared fabric one packet from one tenant takes down a GPU node and every job on it.","attack_vector":"Send a crafted UDP datagram to port 4791 on any host with rdma_rxe loaded. The BTH pad field is set by the attacker, so even a packet long enough to pass the header check can drive the payload length negative. Entirely pre-authentication and reachable from any source the network permits, including across routed segments if 4791 is not filtered.","remediation":"Host reboot / kernel upgrade. Faster and cheaper: unload and blacklist rdma_rxe on every node that is not deliberately running Soft-RoCE (config change, zero downtime) - this closes the entire rxe family of remote packet bugs. If rxe is genuinely required, firewall UDP/4791 to known peers immediately as a stopgap, then upgrade the kernel on a rolling drain. Note the follow-on CVE-2026-46133 shows the first attempt at this fix was incomplete, so verify you are on a kernel carrying both.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46043.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46043"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53186","cve":"CVE-2026-53186","aliases":["RDMA/srp SRP_RSP sense buffer overrun","SCSI RDMA Protocol initiator"],"title":"Linux kernel - SRP (SCSI RDMA Protocol) initiator, drivers/infiniband/ulp/srp/ib_srp.c: TENANT ISOLATION: The SRP initiator copied the sense data out of an SRP_RSP without bounding the copy by the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - SRP (SCSI RDMA Protocol) initiator, drivers/infiniband/ulp/srp/ib_srp.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: The SRP initiator copied the sense data out of an SRP_RSP without bounding the copy by the length actually received, so a malicious or compromised SRP target can overrun the initiator's buffer and read host kernel memory or crash the node. This inverts the usual threat direction - here the storage array attacks its clients. In a GPU cluster where an array or a software SRP target serves many compute nodes, one compromised target reaches every node that mounts from it, which is a fleet-wide blast radius from a single storage compromise.","attack_vector":"The attacker controls or has compromised an SRP target the victim connects to, and returns an SRP_RSP whose declared sense length exceeds what was received. No credentials on the victim are needed beyond it being a normal client of the target. Also reachable by an attacker who can inject into the SRP connection using the RDMA packet-injection primitives.","remediation":"Host reboot / kernel upgrade on all SRP initiator nodes. Interim: verify SRP targets are on a management-isolated storage fabric with per-tenant partitioning so an untrusted party cannot stand up a rogue target and attract connections; that is a switch/SM config change. If SRP is legacy in your environment and NVMe-oF has replaced it, unload ib_srp and remove the initiator configuration entirely - a config change that eliminates the surface.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-53186.json","https://nvd.nist.gov/vuln/detail/CVE-2026-53186"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64269","cve":"CVE-2026-64269","aliases":["RDMA/rtrs-srv chunk overrun","RTRS server rdma_write_sg unbounded length"],"title":"Linux kernel - RDMA/rtrs server (RDMA Transport, used by RNBD block storage), drivers/infiniband/ulp/rtrs/rtrs-srv.c: TENANT ISOLATION: When the RTRS server answers a READ it builds the RDMA WRITE…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rtrs server (RDMA Transport, used by RNBD block storage), drivers/infiniband/ulp/rtrs/rtrs-srv.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: When the RTRS server answers a READ it builds the RDMA WRITE source scatter/gather entry with a length taken straight off the wire descriptor, which the remote peer filled itself, and the source lkey used is the PD-wide local_dma_lkey rather than a key bound to that chunk's mapping - so the verbs layer never constrains the transfer to the chunk size. A peer advertising a length larger than max_chunk_size makes the NIC read past the chunk's mapped region and ship the result back over the fabric. With no IOMMU or in passthrough mode - a common configuration on high-performance storage nodes chasing latency - that returns adjacent host memory to the attacker. This is exactly the disaggregated-storage tenant-isolation break operators worry about: one client reads memory belonging to the server and, by extension, to other clients.","attack_vector":"The attacker is an RTRS client connected to the server - the normal trust position for a storage tenant. They set desc[0].len in the read descriptor larger than the negotiated chunk size; before the fix only a zero length was rejected. With a translating IOMMU the over-range access faults and drops the connection (denial of service instead of disclosure), so IOMMU configuration decides whether this reads memory or just breaks the session.","remediation":"Host reboot / kernel upgrade on RTRS/RNBD server nodes. Interim config changes that materially reduce impact: enable the IOMMU in translating (not passthrough) mode on storage servers, which converts disclosure into a connection abort - this is a kernel command-line change requiring a reboot anyway, so fold it into the same maintenance window and expect a small latency cost. Restrict RTRS server ports to authenticated client subnets. If RNBD/RTRS is not in use, ensure the modules are not loaded.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64269.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64269"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64319","cve":"CVE-2026-64319","aliases":["nvmet-auth DHCHAP_REPLY bounds","NVMe-oF in-band auth heap overread"],"title":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c: TENANT ISOLATION: nvmet_auth_reply() reads the variable-length array using attacker-supplied…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target DH-HMAC-CHAP authentication, drivers/nvme/target/fabrics-cmd-auth.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: nvmet_auth_reply() reads the variable-length array using attacker-supplied hash-length and DH-value-length fields without checking they fit the allocated transfer length. A malicious initiator sends a DHCHAP_REPLY with a small transfer length but large hl/dhvlen and drives out-of-bounds heap reads of up to 526 bytes past the buffer, with the out-of-bounds pointer handed straight to the crypto layer. The bitter irony for operators: this is in the authentication code you were told to turn on to fix NQN spoofing, and it is exploitable pre-authentication - so enabling in-band auth opens this surface rather than closing it. Still worth enabling auth, but only on a patched kernel.","attack_vector":"Any peer that can reach an auth-enabled NVMe-oF target sends a crafted DHCHAP_REPLY during the authentication exchange. By definition this happens before authentication completes, so no credentials are required. Discovered by an automated vulnerability-discovery engine, which suggests more of this class is coming in the same file.","remediation":"Host reboot / kernel upgrade on NVMe-oF targets before or at the same time as enabling DH-HMAC-CHAP. If you have already enabled in-band auth on an unpatched kernel, prioritise the upgrade; disabling auth is not a good trade because it reopens NQN spoofing. Network-level control in the interim: restrict target ports to known initiator addresses. Sequence the rollout as patch first, then enable auth.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64319.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64319"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64320","cve":"CVE-2026-64320","aliases":["nvmet pre-auth OOB heap read","NVMe-oF Discovery Get Log Page memory disclosure"],"title":"Linux kernel - NVMe-oF target discovery controller, drivers/nvme/target/discovery.c: TENANT ISOLATION: The discovery controller validated only the dword alignment of the host-supplied 64-bit Log…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target discovery controller, drivers/nvme/target/discovery.c","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"TENANT ISOLATION: The discovery controller validated only the dword alignment of the host-supplied 64-bit Log Page Offset before adding it to a small heap buffer and memcpy'ing data straight back over the fabric. The discovery subsystem accepts every Host NQN and explicitly skips DH-HMAC-CHAP, so this is reachable pre-authentication by any TCP, RDMA, or FC peer that can see the target. The kernel CNA's own writeup reports an empirical run on a default nvmet-tcp target leaking 81 canonical kernel pointers in a single Get Log Page response; pointing the offset at unmapped memory panics the target instead. For a neocloud running disaggregated NVMe, one unauthenticated request from any tenant reads the storage node's kernel heap - and repeated requests crash the node serving everyone.","attack_vector":"Connect to the discovery controller (nvme discover, or a raw fabrics command) with any Host NQN and issue Get Log Page with an offset at or beyond the allocated discovery log length. No authentication, no valid identity, no prior connection. Works over NVMe/TCP on port 4420, over NVMe/RDMA, and over Fibre Channel. Purely a read primitive plus a crash primitive, but the leaked kernel pointers defeat KASLR and set up heavier exploitation.","remediation":"Host reboot / kernel upgrade on every nvmet target - urgent, given pre-auth reachability and a 9.1 score. Immediate stopgaps while you schedule it: firewall NVMe/TCP 4420 and the RDMA discovery path to known initiator addresses, and move the discovery controller off any tenant-reachable interface onto a management network (nvmet configfs change, runtime, no reboot, but initiators need reconfiguring). Enabling DH-HMAC-CHAP does not help here because the discovery subsystem bypasses it by design - only the kernel fix or network isolation closes it.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64320.json","https://nvd.nist.gov/vuln/detail/CVE-2026-64320"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-7302","cve":"CVE-2026-7302","aliases":[],"title":"SGLang (multimodal runtime): Unauthenticated path traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (multimodal runtime)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Unauthenticated path traversal → arbitrary file write as the server process","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade; run serving processes non-root with read-only root filesystems","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7302"],"status":"curated"},{"id":"CVE-2026-7482","cve":"CVE-2026-7482","aliases":[],"title":"Ollama (GGUF model loader): Heap out-of-bounds read from an attacker-supplied GGUF via `/api/create`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF model loader)","year":"2026","cvss_score":9.1,"severity":"critical","kev":false,"impact":"Heap out-of-bounds read from an attacker-supplied GGUF via `/api/create`","attack_vector":"Customer-supplied model file to an exposed API","remediation":"Upgrade to 0.17.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7482"],"status":"curated"},{"id":"CVE-2018-8930","cve":"CVE-2018-8930","aliases":["MASTERKEY-1","MASTERKEY-2","MASTERKEY-3"],"title":"AMD EPYC / Ryzen - Hardware Validated Boot enforcement: MULTI-TENANT ISOLATION: Hardware Validated Boot is not properly enforced, so an attacker who can reflash the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD EPYC / Ryzen - Hardware Validated Boot enforcement","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: Hardware Validated Boot is not properly enforced, so an attacker who can reflash the BIOS can install firmware the platform will accept and execute despite it not being legitimately signed. That is a persistent, below-the-OS implant on a server: it survives reimaging, disk replacement and tenant handoff, and nothing running in the OS can see it. On a bare-metal GPU cloud where nodes are recycled between customers, this is the classic 'previous tenant left something behind' scenario.","attack_vector":"Local with BIOS reflash capability - so root plus SPI write access, a compromised BMC, or physical/supply-chain access. Not remote.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. The durable control on a bare-metal fleet is not the patch but the process: measure firmware between tenants, enable platform SPI write protection, and treat any node whose firmware measurement changed as suspect rather than reprovisioning it blindly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8930","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-8931","cve":"CVE-2018-8931","aliases":["RYZENFALL-1"],"title":"AMD Secure Processor (Ryzen / Ryzen Pro / Ryzen Mobile): MULTI-TENANT ISOLATION: Insufficient access control on the Secure Processor lets code already running with OS…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (Ryzen / Ryzen Pro / Ryzen Mobile)","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient access control on the Secure Processor lets code already running with OS administrator rights reach into the AMD Secure Processor and execute there. The ASP sits below the hypervisor and below Secure Boot, so once an attacker is inside it, everything the platform's security rests on - memory encryption keys, fTPM state, boot measurements - is theirs. On a shared host this is the end of any isolation guarantee you were making to tenants, and the compromise survives an OS reinstall.","attack_vector":"Local, requires OS administrator/root plus the ability to flash or load a signed driver. Not remotely reachable. Realistically this is a post-exploitation depth charge: an attacker who already owns the host uses it to get persistence you cannot wash out.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. AMD's fix was a PSP firmware update carried in AGESA. Note this batch (the CTS-Labs disclosures) targeted client Ryzen silicon rather than EPYC; verify against your actual server SKU before spending a maintenance window on it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8931","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-8932","cve":"CVE-2018-8932","aliases":["RYZENFALL-2","RYZENFALL-3","RYZENFALL-4"],"title":"AMD Secure Processor (Ryzen / Ryzen Pro): MULTI-TENANT ISOLATION: The same class of Secure Processor access-control failure as RYZENFALL-1, covering…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (Ryzen / Ryzen Pro)","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: The same class of Secure Processor access-control failure as RYZENFALL-1, covering three further variants. An administrator-level attacker writes into ASP-protected memory and gains execution in the secure coprocessor, defeating the hardware root of trust the rest of the platform is anchored to.","attack_vector":"Local, administrator-privileged. Requires the attacker to already control the OS on the node.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Client-silicon focused; confirm applicability to your EPYC server SKUs before scheduling.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8932","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-8933","cve":"CVE-2018-8933","aliases":["FALLOUT-1","FALLOUT-2","FALLOUT-3"],"title":"AMD EPYC Server - protected memory region access control: MULTI-TENANT ISOLATION: Insufficient access control over protected memory regions on EPYC server parts lets…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD EPYC Server - protected memory region access control","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient access control over protected memory regions on EPYC server parts lets privileged code read and write memory that the platform reserves for security purposes - including regions used by SMM and the secure processor. An attacker uses it to get persistence and to reach data the hardware was supposed to fence off from the OS entirely.","attack_vector":"Local, requires administrator privilege on the host. EPYC server silicon specifically, which is what makes this batch relevant to a datacenter rather than a desktop.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8933","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-8934","cve":"CVE-2018-8934","aliases":["CHIMERA-FW"],"title":"Promontory chipset firmware (AMD Ryzen / Ryzen Pro platforms): MULTI-TENANT ISOLATION: A backdoor in the Promontory chipset firmware. The chipset sits on the DMA path for…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Promontory chipset firmware (AMD Ryzen / Ryzen Pro platforms)","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: A backdoor in the Promontory chipset firmware. The chipset sits on the DMA path for USB, SATA and PCIe, so code running there can read and write host memory independently of the CPU and outside the reach of anything the OS enforces. Included as the canonical example of the risk class rather than as an EPYC issue: third-party chipset silicon in your server has its own firmware, its own DMA capability, and usually no attestation story at all.","attack_vector":"Local, requires the ability to load chipset firmware. Ryzen/Ryzen Pro client platforms rather than EPYC servers.","remediation":"Fixed by a chipset firmware update from the OEM, delivered in a BIOS package - drain plus power cycle. Verify applicability before spending a window: this is client-platform silicon and almost certainly not in your EPYC server fleet. The transferable lesson for a datacenter operator is to ask which non-AMD firmware images your server actually loads at boot and who signs them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8934","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-8936","cve":"CVE-2018-8936","aliases":["CHIMERA-adjacent PSP escalation"],"title":"AMD EPYC / Ryzen - Platform Security Processor privilege escalation: MULTI-TENANT ISOLATION: A direct privilege escalation into the Platform Security Processor on EPYC server…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD EPYC / Ryzen - Platform Security Processor privilege escalation","year":"2018","cvss_score":9,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: A direct privilege escalation into the Platform Security Processor on EPYC server parts. The PSP holds the platform's root of trust, fTPM state and SEV key material, so an attacker who escalates into it owns the security posture of the whole node beneath the hypervisor. Nothing the OS or the hypervisor can do detects or contains it.","attack_vector":"Local, administrator privilege.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-8936","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-15206","cve":"CVE-2020-15206","aliases":[],"title":"TensorFlow (SavedModel protobuf): Mutating a SavedModel protobuf crashes or corrupts the serving process","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TensorFlow (SavedModel protobuf)","year":"2020","cvss_score":9,"severity":"critical","kev":false,"impact":"Mutating a SavedModel protobuf crashes or corrupts the serving process","attack_vector":"Customer-supplied SavedModel served by a shared serving tier","remediation":"Patch; isolate per-model serving processes","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15206"],"status":"curated"},{"id":"CVE-2021-42114","cve":"CVE-2021-42114","aliases":["Blacksmith"],"title":"PC-DDR4 / LPDDR4X DRAM - Target Row Refresh mitigation: Non-uniform Rowhammer patterns triggered bit flips on every one of the 40 DDR4 modules the researchers…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"PC-DDR4 / LPDDR4X DRAM - Target Row Refresh mitigation","year":"2021","cvss_score":9,"severity":"critical","kev":false,"impact":"Non-uniform Rowhammer patterns triggered bit flips on every one of the 40 DDR4 modules the researchers tested, including modules whose TRR implementation had resisted TRRespass. Scored 9.0 with a changed scope. Same operator consequence as TRRespass: an integrity attack on host memory reachable from tenant code.","attack_vector":"Local code on the node able to generate the access pattern. Scored AV:Network by the reporters because remote code paths (JavaScript, network stacks) can drive memory access, but the realistic datacenter path is a tenant workload.","remediation":"No vendor patch. Use ECC and alert on correctable-error rate rather than only on uncorrectable errors; enable increased refresh rate or RFM in BIOS if your platform exposes it, accepting a small memory-bandwidth cost. Cost: BIOS change means drain and reboot. Effectively UNPATCHABLE.","references":["https://comsec.ethz.ch/research/dram/blacksmith/","https://comsec.ethz.ch/wp-content/files/blacksmith_sp22.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-42114"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-22805","cve":"CVE-2022-22805","aliases":["TLStorm"],"title":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - cloud-connected UPS firmware: PHYSICAL. A heap overflow in TLS packet reassembly gives an attacker code execution on the UPS's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - cloud-connected UPS firmware","year":"2022","cvss_score":9,"severity":"critical","kev":false,"impact":"PHYSICAL. A heap overflow in TLS packet reassembly gives an attacker code execution on the UPS's own controller - the device that decides whether your racks get power. From there an attacker can cut output, refuse to transfer to battery during a utility event, or (as Armis demonstrated on the bench) drive the unit until it physically burns. This is not a monitoring card compromise; it is the power train. A single UPS covering a GPU row kills every training job in that row with no checkpoint.","attack_vector":"Unauthenticated. The UPS initiates an outbound TLS connection to Schneider's cloud service, so an attacker who can intercept or MITM that connection - or who is simply on the same network segment as the UPS management port - reaches the vulnerable parser. No credentials, no prior foothold on the compute network.","remediation":"Firmware flash on every affected UPS, pushed through the Schneider update tool or the cloud service. Cost is real: each unit must be updated individually and some models require the load to be transferred or the unit taken to bypass first, so this is a scheduled electrical maintenance window per unit, not a fleet-wide push. If you cannot patch promptly, block the UPS's outbound path to the SmartConnect cloud and put the management port on an isolated VLAN with no route to the internet.","references":["https://www.se.com/ww/en/download/document/SEVD-2022-067-02/","https://www.armis.com/research/tlstorm/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-22806","cve":"CVE-2022-22806","aliases":["TLStorm"],"title":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - TLS state machine: PHYSICAL. A TLS authentication bypass by capture-replay: a malformed handshake puts the UPS into an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Smart-UPS SmartConnect family (SMT, SMC, SMTL, SCL, SMX series) - TLS state machine","year":"2022","cvss_score":9,"severity":"critical","kev":false,"impact":"PHYSICAL. A TLS authentication bypass by capture-replay: a malformed handshake puts the UPS into an unauthenticated-but-connected state, letting an attacker talk to it as if it were the trusted cloud service. Combined with the unsigned-firmware issue below, this is the full chain from network reachability to controlling whether a GPU hall stays energised.","attack_vector":"Unauthenticated, network-adjacent to the UPS management interface or positioned on the path of its outbound cloud connection.","remediation":"Same firmware campaign as CVE-2022-22805 - a per-unit flash with an electrical maintenance window for models that need a bypass transfer. Interim mitigation is network isolation of the UPS management plane and blocking its egress. There is no configuration toggle that removes the vulnerable code path.","references":["https://www.se.com/ww/en/download/document/SEVD-2022-067-02/","https://www.armis.com/research/tlstorm/"],"status":"curated"},{"id":"CVE-2022-31035","cve":"CVE-2022-31035","aliases":[],"title":"Argo CD: Stored XSS via a `javascript:` link executes in an admin's browser","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":9,"severity":"critical","kev":false,"impact":"Stored XSS via a `javascript:` link executes in an admin's browser","attack_vector":"Any user who can create an Application","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31035"],"status":"curated"},{"id":"CVE-2023-22482","cve":"CVE-2023-22482","aliases":[],"title":"Argo CD: Improper authorization causes the API to accept tokens it should reject","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":9,"severity":"critical","kev":false,"impact":"Improper authorization causes the API to accept tokens it should reject","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22482"],"status":"curated"},{"id":"CVE-2023-4299","cve":"CVE-2023-4299","aliases":["ICSA-23-243-04"],"title":"Digi RealPort protocol (Digi console/terminal servers): RealPort is the protocol Digi console servers use to expose their serial ports as virtual COM ports over the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Digi RealPort protocol (Digi console/terminal servers)","year":"2023","cvss_score":9,"severity":"critical","kev":false,"impact":"RealPort is the protocol Digi console servers use to expose their serial ports as virtual COM ports over the network. Its authentication can be replayed — an attacker who captures a legitimate auth exchange (e.g. via a network tap or ARP spoof on the management VLAN) can replay it to open a session on connected serial equipment without knowing the real password.","attack_vector":"Requires network visibility into a RealPort authentication exchange (passive capture is enough) and the ability to send the replayed traffic to the target console server — no credential cracking needed.","remediation":"Firmware/software upgrade on the RealPort driver and the console server firmware to a version with replay-resistant authentication; also a network-segmentation fix — RealPort traffic should never traverse a network segment an untrusted party can sniff. Rollout is a firmware flash per device plus a driver update on every host connecting to RealPort ports.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-243-04","https://www.digi.com/getattachment/resources/security/alerts/realport-cves/Dragos-Disclosure-Statement.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-0087","cve":"CVE-2024-0087","aliases":[],"title":"Triton Inference Server: RCE via path traversal on the model-load API","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"RCE via path traversal on the model-load API","attack_vector":"Unauthenticated/low-priv client of the inference endpoint","remediation":"Upgrade Triton (24.09+); rebuild and redeploy all serving images; restrict model-control API","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0087","https://github.com/NVIDIA/product-security/tree/main/2024/5535"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:C/C:H/I:L/A:H","cwe":["CWE-73"]},{"id":"CVE-2024-0095","cve":"CVE-2024-0095","aliases":[],"title":"Triton Inference Server: Improper logging of security events (audit gap)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Improper logging of security events (audit gap)","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0095","https://github.com/NVIDIA/product-security/tree/main/2024/5546"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:N","cwe":["CWE-117"]},{"id":"CVE-2024-0132","cve":"CVE-2024-0132","aliases":[],"title":"Container Toolkit: Container escape to host filesystem via TOCTOU in the **default** configuration","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Container escape to host filesystem via TOCTOU in the **default** configuration","attack_vector":"Any tenant that can run an arbitrary container image on a GPU node","remediation":"Bump nvidia-container-toolkit to 1.16.2+ and restart the container runtime on every GPU node; upgrade GPU Operator to 24.6.2+; evict and re-admit tenant workloads","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=NVIDIA%20Container%20Toolkit"],"status":"curated","fleet":{"ubiquity":"Universal - all versions <= 1.16.1, i.e. the entire installed base at disclosure","remediation_pain":"`daemon-restart` to 1.16.2 / GPU Operator 24.6.2; hosts that ran untrusted images need `node-drain` + rebuild because the host-filesystem write has already happened","pain_class":"node-drain","why_fleet_wide":"TOCTOU in the mount path lets a crafted container image mount the host filesystem and execute code as root - the canonical \"one CVE, every GPU node in the fleet\" event"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-367"]},{"id":"CVE-2024-28175","cve":"CVE-2024-28175","aliases":[],"title":"Argo CD: Improper URL protocol filtering in link annotations enables client-side attacks against admins","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Improper URL protocol filtering in link annotations enables client-side attacks against admins","attack_vector":"Any user who can annotate an Application","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28175"],"status":"curated"},{"id":"CVE-2024-28179","cve":"CVE-2024-28179","aliases":[],"title":"Jupyter Server Proxy: Authentication weakness in proxied-process access","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server Proxy","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"Authentication weakness in proxied-process access","attack_vector":"Network user of a JupyterHub deployment","remediation":"Upgrade; commonly deployed on managed GPU notebook platforms","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28179"],"status":"curated"},{"id":"CVE-2024-31989","cve":"CVE-2024-31989","aliases":[],"title":"Argo CD: An unprivileged pod in any namespace can reach the unauthenticated Argo CD Redis on 6379 and poison the…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"An unprivileged pod in any namespace can reach the unauthenticated Argo CD Redis on 6379 and poison the cache, which becomes arbitrary deployment","attack_vector":"Any pod on the cluster network","remediation":"Rolling Argo CD upgrade; enable Redis auth and a NetworkPolicy around it. Critical in a multi-tenant neocloud where any tenant pod is on that network","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31989"],"status":"curated"},{"id":"CVE-2024-45764","cve":"CVE-2024-45764","aliases":["DSA-2024-449"],"title":"Dell Enterprise SONiC (authentication): A critical step in authentication is missing, so an unauthenticated remote attacker bypasses the protection…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (authentication)","year":"2024","cvss_score":9,"severity":"critical","kev":false,"impact":"A critical step in authentication is missing, so an unauthenticated remote attacker bypasses the protection mechanism and gets into the switch. SONiC is increasingly the NOS of choice for cost-driven GPU buildouts precisely because it is open and cheap; this is the reminder that the open NOS ecosystem has the same class of front-door bugs as the incumbents, with a shorter advisory history to check against.","attack_vector":"Unauthenticated, remote — reachability to the switch's management services is the only requirement.","remediation":"Upgrade Dell Enterprise SONiC past 4.1.x/4.2.x to a fixed release, which means a NOS image install and switch reboot per device. In a SONiC fabric that is a full image swap, not a patch — budget a maintenance window per leaf and stage it across MLAG pairs. Restrict management-interface reachability in the meantime.","references":["https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2024-45764"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0282","cve":"CVE-2025-0282","aliases":[],"title":"Ivanti Connect Secure: Stack-based buffer overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure","year":"2025","cvss_score":9,"severity":"critical","kev":true,"impact":"[KEV] Stack-based buffer overflow -> unauthenticated remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: emergency patch plus factory reset","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0282"],"status":"curated"},{"id":"CVE-2025-22457","cve":"CVE-2025-22457","aliases":[],"title":"Ivanti Connect Secure/ZTA: Stack-based buffer overflow","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure/ZTA","year":"2025","cvss_score":9,"severity":"critical","kev":true,"impact":"[KEV] Stack-based buffer overflow -> unauthenticated RCE, exploited by a China-nexus actor","attack_vector":"Network (remote)","remediation":"Control-plane: patch; assume compromise","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22457"],"status":"curated"},{"id":"CVE-2025-23266","cve":"CVE-2025-23266","aliases":["NVIDIAScape"],"title":"Container Toolkit: Container escape to host root via malicious image (LD_PRELOAD in OCI hook)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"Container escape to host root via malicious image (LD_PRELOAD in OCI hook)","attack_vector":"Any tenant that can run an arbitrary container image on a GPU node","remediation":"Emergency: bump nvidia-container-toolkit to 1.17.8+, restart container runtime on every GPU node, upgrade GPU Operator Helm chart; evict and re-admit all tenant workloads; audit for prior exploitation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23266","https://github.com/NVIDIA/product-security/tree/main/2025/5659"],"status":"curated","fleet":{"ubiquity":"Very common - the standard way to ship driver + toolkit + DCGM on every K8s-based neocloud","remediation_pain":"`node-drain` in practice: the Operator's driver and toolkit DaemonSets restart per node, and a driver-container reload requires evicting every GPU pod","pain_class":"node-drain","why_fleet_wide":"One Helm version bump has to roll across every GPU node pool; until it finishes, every tenant pod on an un-rolled node still holds the escape primitive"},"cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-426"]},{"id":"CVE-2025-23350","cve":"CVE-2025-23350","aliases":[],"title":"BlueField (GA firmware): RCE on the DPU via firmware buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField (GA firmware)","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"RCE on the DPU via firmware buffer overflow","attack_vector":"Privileged network attacker on the DPU management path","remediation":"Flash BlueField firmware out-of-band; DPU reset drops tenant networking, schedule node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23350","https://github.com/NVIDIA/product-security/tree/main/2026/5699"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-23351","cve":"CVE-2025-23351","aliases":[],"title":"BlueField (GA firmware): RCE on the DPU via firmware buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField (GA firmware)","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"RCE on the DPU via firmware buffer overflow","attack_vector":"Privileged network attacker","remediation":"Flash BlueField firmware out-of-band; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23351","https://github.com/NVIDIA/product-security/tree/main/2026/5699"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-29783","cve":"CVE-2025-29783","aliases":[],"title":"vLLM (Mooncake): Unsafe deserialization over ZMQ/TCP bound to all interfaces","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (Mooncake)","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"Unsafe deserialization over ZMQ/TCP bound to all interfaces","attack_vector":"Unauthenticated network, co-tenant reachable","remediation":"Upgrade; the KV-transfer plane is unauthenticated by design in these versions","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29783"],"status":"curated"},{"id":"CVE-2025-33210","cve":"CVE-2025-33210","aliases":[],"title":"NVIDIA Isaac Lab: A deserialization flaw reaches code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Lab","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"A deserialization flaw reaches code execution; scored 9.0 network with a changed scope. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5733 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33210","https://github.com/NVIDIA/product-security/tree/main/2025/5733"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33244","cve":"CVE-2025-33244","aliases":[],"title":"Apex: Remote RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Apex","year":"2025","cvss_score":9,"severity":"critical","kev":false,"impact":"Remote RCE via unsafe pickle deserialization","attack_vector":"Network attacker / malicious checkpoint","remediation":"Bump Apex in training images; rebuild and redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33244","https://github.com/NVIDIA/product-security/tree/main/2026/5782"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-65094","cve":"CVE-2026-65094","aliases":[],"title":"NVIDIA BlueField - VIRTIO-Net emulation: MULTI-TENANT ISOLATION: a VM user sends a crafted message to the BlueField VIRTIO-Net device and gets a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA BlueField - VIRTIO-Net emulation","year":"2026","cvss_score":9,"severity":"critical","kev":false,"impact":"MULTI-TENANT ISOLATION: a VM user sends a crafted message to the BlueField VIRTIO-Net device and gets a write-what-where primitive, reaching code execution in the VIRTIO-Net context on the DPU. The DPU is the component you offloaded tenant network isolation onto - a tenant VM reaching code execution inside it inverts the trust model of the whole design. Scored 9.0 with a changed scope.","attack_vector":"A user inside a tenant VM talking to the emulated virtio-net device its own hypervisor exposed. No host or DPU credentials needed. This is the guest-to-DPU boundary.","remediation":"Update the BlueField VIRTIO-Net firmware/software per bulletin 5815 across GA, LTS23, LTS24 and LTS25 branches as applicable. Cost: a DPU firmware update takes the DPU's dataplane down, which means the host loses network - treat it as a full node drain, not a live update. Sequence carefully: a half-updated DPU fleet has inconsistent offload behaviour.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65094","https://github.com/NVIDIA/product-security/tree/main/2026/5815"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-123"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-0105","cve":"CVE-2024-0105","aliases":[],"title":"ConnectX / BlueField firmware: Improper certificate validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX / BlueField firmware","year":"2024","cvss_score":8.9,"severity":"high","kev":false,"impact":"Improper certificate validation -> malicious firmware/image accepted","attack_vector":"Network-adjacent attacker in the firmware update path","remediation":"Flash NIC/DPU firmware; verify update chain; node reboot with tenant eviction","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0105","https://github.com/NVIDIA/product-security/tree/main/2024/5562"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-274"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-12060","cve":"CVE-2025-12060","aliases":[],"title":"Keras (`utils.get_file`, tar extract): Path traversal on tar extraction","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`utils.get_file`, tar extract)","year":"2025","cvss_score":8.9,"severity":"high","kev":false,"impact":"Path traversal on tar extraction → arbitrary file write","attack_vector":"Customer-supplied dataset/model URL fetched with `extract=True`","remediation":"Upgrade; a training job with write access to shared mounts can escape into other tenants' paths if mounts are shared","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12060"],"status":"curated"},{"id":"CVE-2025-53630","cve":"CVE-2025-53630","aliases":[],"title":"llama.cpp (`gguf_init_from_file_impl`): Integer overflow in GGUF init","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`gguf_init_from_file_impl`)","year":"2025","cvss_score":8.9,"severity":"high","kev":false,"impact":"Integer overflow in GGUF init","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53630"],"status":"curated"},{"id":"CVE-2018-0395","cve":"CVE-2018-0395","aliases":[],"title":"Cisco NX-OS / FXOS (LLDP parser): FABRIC DOS: a malformed LLDP frame reloads the switch. LLDP is enabled by default on essentially every fabric…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS / FXOS (LLDP parser)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"FABRIC DOS: a malformed LLDP frame reloads the switch. LLDP is enabled by default on essentially every fabric port and is unauthenticated by design, so any host on any port can drop its leaf. In a training cluster a single leaf reload takes out a whole rack of GPUs mid-job.","attack_vector":"Unauthenticated, adjacent — a single crafted frame from a directly connected host.","remediation":"NX-OS upgrade plus reload. Interim: disable LLDP on tenant-facing ports. Note that if you use LLDP for cabling verification (common in GPU builds, to prove the rail-optimized topology is wired right) disabling it costs you that check, so most operators patch rather than disable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0395"],"status":"curated"},{"id":"CVE-2018-1000400","cve":"CVE-2018-1000400","aliases":[],"title":"CRI-O: Ambient-capability mishandling runs containers with elevated privileges","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Ambient-capability mishandling runs containers with elevated privileges","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1000400"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2018-1244","cve":"CVE-2018-1244","aliases":[],"title":"Dell iDRAC7 / iDRAC8 / iDRAC9 (SNMP agent): Command injection in the iDRAC SNMP agent gives an attacker who already holds an iDRAC account with…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7 / iDRAC8 / iDRAC9 (SNMP agent)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Command injection in the iDRAC SNMP agent gives an attacker who already holds an iDRAC account with configuration rights arbitrary command execution as root on the BMC itself. That is the full BMC prize: out-of-band power control, Virtual Media boot, KVM, and an implant that lives on the service processor and survives every host reimage and OS reinstall. The privilege jump matters - it converts a routine monitoring or configuration credential into persistent control of the node beneath the hypervisor.","attack_vector":"An authenticated iDRAC account holding the Configure iDRAC privilege - the kind of account handed to monitoring tooling, an integrator, or a datacenter-remote-hands team, not a full administrator. Reachability is the management VLAN.","remediation":"Flash to iDRAC7/8 2.60.60.60 or iDRAC9 3.21.21.21 or later - out-of-band, per-node, no host reboot and no job drain. Config-only mitigations that reduce blast radius immediately: disable the iDRAC SNMP agent where you are not actually scraping it, and audit which service accounts hold Configure iDRAC rather than read-only. Original Dell TechCenter advisory URL is dead; NVD carries the version data.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1244","https://www.dell.com/support/kbdoc/en-us/000131409/dsa-2018-004-idrac-vulnerabilities"],"status":"curated"},{"id":"CVE-2018-15774","cve":"CVE-2018-15774","aliases":[],"title":"Dell iDRAC9 (Redfish): Redfish interface permission-check flaw enabling privilege escalation to admin","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (Redfish)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Redfish interface permission-check flaw enabling privilege escalation to admin","attack_vector":"Network / Redfish, authenticated low-privilege","remediation":"iDRAC firmware update; illustrates that Redfish RBAC bugs are as common as the IPMI ones they replaced","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15774"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-3628","cve":"CVE-2018-3628","aliases":[],"title":"Intel AMT (HTTP handler) in Intel CSME firmware: A buffer overflow in AMT's HTTP handler allows arbitrary code execution with AMT privileges. AMT's HTTP…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT (HTTP handler) in Intel CSME firmware","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"A buffer overflow in AMT's HTTP handler allows arbitrary code execution with AMT privileges. AMT's HTTP handler is what serves the out-of-band management interface, so exploitation gives the attacker the same below-the-OS control AMT itself has.","attack_vector":"Reachable through the AMT network interface on provisioned machines.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines. Immediate compensating control is to unprovision AMT where you do not use it and firewall the AMT ports where you do.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3628","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00112.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6247","cve":"CVE-2018-6247","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Any local account on a Windows GPU host can call the driver's escape path and hit a NULL dereference. Best…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Any local account on a Windows GPU host can call the driver's escape path and hit a NULL dereference. Best case the box bugchecks and you lose every job on it; NVIDIA also rates privilege escalation as possible, which on a kernel driver means SYSTEM.","attack_vector":"Any local user or service account on the host, including anything running inside a Windows container that has the GPU mapped in.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6247"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6248","cve":"CVE-2018-6248","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Out-of-bounds kernel read/write from an unprivileged escape call - a length value is trusted that should not…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Out-of-bounds kernel read/write from an unprivileged escape call - a length value is trusted that should not be. This is the shape that turns into a full SYSTEM escalation with enough effort, and a bugcheck with none.","attack_vector":"Any local user on the host with access to the GPU device, including low-privilege service accounts.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6248"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6249","cve":"CVE-2018-6249","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): NULL dereference in the kernel-mode layer reachable from unprivileged code on Windows, Linux, FreeBSD and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"NULL dereference in the kernel-mode layer reachable from unprivileged code on Windows, Linux, FreeBSD and Solaris hosts. Loss of the node at minimum; NVIDIA leaves escalation on the table.","attack_vector":"Any local user with a handle on the GPU device node - which on Linux means anyone in the container that got /dev/nvidia*.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://usn.ubuntu.com/3662-1/","https://nvd.nist.gov/vuln/detail/CVE-2018-6249"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6250","cve":"CVE-2018-6250","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Same class as the other early-2018 escape bugs: an unprivileged caller crashes the kernel driver, with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"Same class as the other early-2018 escape bugs: an unprivileged caller crashes the kernel driver, with escalation to SYSTEM not ruled out by the vendor.","attack_vector":"Any local user on the Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6250"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6442","cve":"CVE-2018-6442","aliases":[],"title":"Brocade Fabric OS Webtools (firmware update section): A remote authenticated attacker can abuse the Webtools firmware-update path. Firmware update on a SAN switch…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS Webtools (firmware update section)","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"A remote authenticated attacker can abuse the Webtools firmware-update path. Firmware update on a SAN switch is the persistence mechanism — an attacker who can drive it installs an image that survives every subsequent remediation. Companion issue CVE-2018-6436 does the same through the `firmwaredownload` CLI command for a local attacker.","attack_vector":"Authenticated remote user with Webtools access on FOS before 8.2.1 / 8.1.2f / 8.0.2f / 7.4.2d.","remediation":"Fabric OS upgrade plus reboot. Disable Webtools if your operations are CLI/REST-driven — a live config change. Verify installed firmware digests against Broadcom's published values after any suspicious period, because a patched switch running a tampered image is still compromised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6442","https://nvd.nist.gov/vuln/detail/CVE-2018-6436"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-7807","cve":"CVE-2018-7807","aliases":[],"title":"Schneider Electric Data Center Expert 7.5.0 and earlier - zip upload: A crafted zip uploaded through the DCE UI can path-traverse out of its intended directory and write arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric Data Center Expert 7.5.0 and earlier - zip upload","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"A crafted zip uploaded through the DCE UI can path-traverse out of its intended directory and write arbitrary files on the appliance. The exploit path is social as much as technical: an authenticated operator is tricked into uploading a file that looks like a normal DCIM import.","attack_vector":"An authenticated DCE user performing what looks like a routine upload of a supplied file.","remediation":"Upgrade past 7.5.0. Operationally, the lesson generalises to every DCIM: treat imported bundles, device definition packs and 'helpful' vendor-supplied files as untrusted input.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7807"],"status":"curated"},{"id":"CVE-2018-9281","cve":"CVE-2018-9281","aliases":[],"title":"Eaton UPS 9PX 8000 SP administration panel: CSRF on the change-password function plus reflected XSS: an attacker forces a silent password change on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS 9PX 8000 SP administration panel","year":"2018","cvss_score":8.8,"severity":"high","kev":false,"impact":"CSRF on the change-password function plus reflected XSS: an attacker forces a silent password change on the UPS admin account and takes over the device. Once they hold the UPS admin account they control shutdown behaviour and output for whatever the unit feeds - a PHYSICAL outcome from a web bug.","attack_vector":"Requires a logged-in UPS administrator to load an attacker-controlled page.","remediation":"Firmware update on the UPS network card where available. On units this age, availability of a fix is not guaranteed - if none exists, the mitigation is a dedicated management VLAN, no browser access to UPS interfaces from general-purpose workstations, and disabling the web interface if the unit can be managed another way.","references":["https://www.bishopfox.com/news/2018/10/eaton-ups-9px-8000-sp-multiple-vulnerabilities/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-0140","cve":"CVE-2019-0140","aliases":["INTEL-SA-00255"],"title":"Intel Ethernet 700 Series Controller firmware (X710/XL710/XXV710): Buffer overflow in the adapter firmware of Intel's 700-series NICs allowing an *unauthenticated* user to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet 700 Series Controller firmware (X710/XL710/XXV710)","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"Buffer overflow in the adapter firmware of Intel's 700-series NICs allowing an *unauthenticated* user to escalate privilege. This is the rare NIC-firmware bug where the published impact is escalation rather than denial of service, and it needs no host account — the attack surface is the network the card is plugged into. A compromised NIC sits below the OS, persists across reinstall, and on many server designs carries the NC-SI sideband to the BMC, so the blast radius extends past the host it lives in.","attack_vector":"Unauthenticated attacker able to reach the adapter over the network. X710/XL710/XXV710 are the standard 10/25/40GbE management and storage NICs in the server generations that host GPUs.","remediation":"Flash 700-series NVM firmware to 7.0 or later via Intel's NVM Update Utility or the OEM firmware bundle. Requires a **cold power cycle** to activate, so it is a per-node drain across the fleet. Audit installed NVM versions first (`ethtool -i`) — 700-series cards are old enough that many fleets have never updated them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0140"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-0169","cve":"CVE-2019-0169","aliases":[],"title":"Intel CSME / TXE: A heap overflow in a CSME subsystem reachable by an unauthenticated attacker for privilege escalation. Same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME / TXE","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"A heap overflow in a CSME subsystem reachable by an unauthenticated attacker for privilege escalation. Same shape and same consequence as the other network-reachable CSME overflows: code execution in the engine that sits under the OS.","attack_vector":"Unauthenticated attacker with access to the affected interface.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0169","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00241.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-13071","cve":"CVE-2019-13071","aliases":[],"title":"CyberPower PowerPanel Business Edition 3.4.0 Agent/Center: Cross-site request forgery across all forms in the web application. An authenticated facility user visiting…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Business Edition 3.4.0 Agent/Center","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"Cross-site request forgery across all forms in the web application. An authenticated facility user visiting an attacker-controlled page silently submits changes to the UPS management configuration - including shutdown behaviour. Old and unglamorous, but this software persists in production far longer than anything on the compute side.","attack_vector":"Requires an authenticated PowerPanel user to visit a page the attacker controls.","remediation":"Upgrade past 3.4.0. If you are still running 3.x, the more useful action is a full inventory of power-management software versions - CyberPower's 2023-2025 CVE history means anything this old is carrying many more unlisted problems.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13071"],"status":"curated"},{"id":"CVE-2019-1901","cve":"CVE-2019-1901","aliases":[],"title":"Cisco Nexus 9000 ACI Mode (LLDP subsystem): A buffer overflow in the LLDP subsystem of Nexus 9000 switches in ACI mode gives an adjacent unauthenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus 9000 ACI Mode (LLDP subsystem)","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"A buffer overflow in the LLDP subsystem of Nexus 9000 switches in ACI mode gives an adjacent unauthenticated attacker denial of service or arbitrary code execution with root privileges on the switch. In ACI, LLDP is not optional — it is how the fabric discovers and validates its own topology — so you cannot simply turn it off the way you can on a standalone NX-OS leaf.","attack_vector":"Unauthenticated, adjacent — a crafted LLDP frame from a device on a leaf port.","remediation":"ACI software upgrade across the fabric (APIC plus switches), staged. There is no good config workaround because ACI depends on LLDP; the compensating control is strict physical and port-admission control on leaf front-panel ports. Related memory-leak issue in the same subsystem: CVE-2023-20089.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1901","https://nvd.nist.gov/vuln/detail/CVE-2023-20089"],"status":"curated"},{"id":"CVE-2019-19023","cve":"CVE-2019-19023","aliases":[],"title":"Harbor: Privilege escalation in the Harbor registry","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"Privilege escalation in the Harbor registry","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19023"],"status":"curated"},{"id":"CVE-2019-19642","cve":"CVE-2019-19642","aliases":[],"title":"Supermicro BMC virtual media subsystem on X8STi-F with IPMI firmware 2.06: The researcher's own description is the operator-relevant one: a persistent backdoor on the BMC. Virtual…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC virtual media subsystem on X8STi-F with IPMI firmware 2.06","year":"2019","cvss_score":8.8,"severity":"high","kev":false,"impact":"The researcher's own description is the operator-relevant one: a persistent backdoor on the BMC. Virtual media is the single most dangerous BMC feature to lose control of, because it is the mechanism by which an attacker attaches their own boot image to a node and reboots into it. Combined with command injection in the same subsystem, the attacker gets both code on the controller and the ability to boot the host into whatever they want - the complete out-of-band takeover, surviving any host-side remediation. Shell metacharacters in the ShareHost and ShareName fields of the /rpc/setvmdrive.asp handler reach a command interpreter on the controller.","attack_vector":"An authenticated attacker who can send HTTP requests to the BMC's IP address. Any valid BMC credential plus a route to the management network is sufficient.","remediation":"This is legacy X8-generation hardware and firmware updates for it are effectively unavailable, so treat this as a case where flashing is not a real option. The controls that work are structural: disable virtual media on the BMC where the platform allows it, isolate these BMCs onto a management VLAN with no route from tenant or general corporate networks, and plan the hardware out. If X8-era Supermicro boards are still carrying production workloads, that is a fleet-lifecycle decision, not a patching decision.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19642","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2019/19xxx/CVE-2019-19642.json"],"status":"curated"},{"id":"CVE-2020-10676","cve":"CVE-2020-10676","aliases":[],"title":"Rancher: Incorrectly applied authorization check lets a namespace be moved into a different project","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrectly applied authorization check lets a namespace be moved into a different project; cross-tenant boundary break","attack_vector":"Cluster user with namespace access","remediation":"Upgrade Rancher","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-10676"],"status":"curated"},{"id":"CVE-2020-11485","cve":"CVE-2020-11485","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): CSRF in the BMC web application. An operator with a BMC session open in a browser can be made to execute BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"CSRF in the BMC web application. An operator with a BMC session open in a browser can be made to execute BMC actions by loading an attacker's page - disclosure or code execution against out-of-band management without the attacker ever needing network reach to the BMC themselves. DGX-1 before BMC 3.38.30.","attack_vector":"Any attacker who can get a logged-in BMC administrator to visit a web page. The attacker needs no route to the management network at all - the admin's browser is the route.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11485"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11953","cve":"CVE-2020-11953","aliases":[],"title":"Rittal PDU-3C002DEC rack PDU firmware (through 5.15.40): Arbitrary code execution on the rack PDU. Once code runs on the PDU, an attacker has a persistent presence on…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal PDU-3C002DEC rack PDU firmware (through 5.15.40)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary code execution on the rack PDU. Once code runs on the PDU, an attacker has a persistent presence on the OOB network that survives every host reimage in the rack, and direct control of outlet state - so this is both a persistence problem and a PHYSICAL availability problem.","attack_vector":"Network access to the PDU management interface.","remediation":"Firmware flash per PDU. Because code execution means possible implantation, a unit you believe was targeted should be re-flashed from vendor image and its stored credentials rotated, not merely updated.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11953"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-12138","cve":"CVE-2020-12138","aliases":[],"title":"AMD ATI atillk64.sys - physical memory mapping driver: MULTI-TENANT ISOLATION: The AMD ATI atillk64.sys driver exposes routines that map physical memory into a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD ATI atillk64.sys - physical memory mapping driver","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The AMD ATI atillk64.sys driver exposes routines that map physical memory into a caller's virtual address space, and it lets low-privileged users call them. That is arbitrary physical memory read and write handed to any local user - complete bypass of kernel memory protection with no memory-corruption exploit required, because the driver simply offers the capability. Drivers like this are a favourite BYOVD (bring-your-own-vulnerable-driver) primitive precisely because they are signed and they work as designed.","attack_vector":"Local, low-privileged user with the driver loaded. Windows driver; note that an attacker can also *bring* this driver to a host that never shipped it, which is why it matters even if you do not deploy AMD's Windows tooling.","remediation":"Remove or update the driver. On Windows fleets, add atillk64.sys to your vulnerable-driver blocklist (Microsoft's WDAC blocklist covers this class) rather than relying on it not being installed - the BYOVD path means an attacker supplies the driver themselves. Not applicable to Linux ROCm nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12138"],"status":"curated"},{"id":"CVE-2020-12347","cve":"CVE-2020-12347","aliases":[],"title":"Intel Data Center Manager Console: Improper input validation in the DCM Console lets an authenticated user escalate privilege over the network.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center Manager Console","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper input validation in the DCM Console lets an authenticated user escalate privilege over the network. Same management-plane exposure profile as the later DCM issues.","attack_vector":"Authenticated user with network access to the DCM console.","remediation":"Upgrade the Intel Data Center Manager software. This is a management-plane application, so the update is an application upgrade and service restart - no node drain, no firmware, no reboot of managed hosts. The real work is deciding what DCM is allowed to reach: it holds credentials for platform power and telemetry across the fleet, so its network exposure matters more than its version. Upgrade to 3.6.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12347","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00430"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2020-14156","cve":"CVE-2020-14156","aliases":[],"title":"OpenBMC phosphor-host-ipmid (user_channel/passwd_mgr.cpp, /etc/ipmi-pass): The file holding IPMI account passwords is written with permissions that let unprivileged BMC processes read…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC phosphor-host-ipmid (user_channel/passwd_mgr.cpp, /etc/ipmi-pass)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"The file holding IPMI account passwords is written with permissions that let unprivileged BMC processes read it. Because IPMI passwords are stored in a form the daemon can recover (they have to be, for RMCP+ key derivation), reading this file yields usable BMC administrator credentials rather than hashes to crack. Any minor foothold on the BMC promotes straight to BMC admin, and if the fleet reuses BMC credentials across nodes - which most do - one node's compromise becomes the whole fleet's.","attack_vector":"Requires some code execution on the BMC as any local user. That bar is met by any of the unauthenticated network-daemon bugs in this cluster. Not directly reachable from the host or the network.","remediation":"Fixed upstream in phosphor-host-ipmid in April 2020; on your nodes it means a BMC firmware flash, per node, out-of-band, ODM-gated. The compensating control matters more than the patch: stop reusing BMC credentials across the fleet, rotate them per node, and prefer Redfish local accounts or LDAP over IPMI accounts so the ipmi-pass file has nothing valuable in it. If you disable IPMI over LAN for CVE-2021-39296, that also drains most of the value out of this file.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14156","https://github.com/openbmc/phosphor-host-ipmid/commit/b265455a2518ece7c004b43c144199ec980fc620","https://github.com/openbmc/openbmc/issues/3670"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-15046","cve":"CVE-2020-15046","aliases":[],"title":"Supermicro BMC web UI user management (cgi/config_user.cgi, X10DRH-iT): An attacker who gets a logged-in BMC administrator to load a page silently gains a permanent administrator…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web UI user management (cgi/config_user.cgi, X10DRH-iT)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker who gets a logged-in BMC administrator to load a page silently gains a permanent administrator account on that BMC. The account persists after the victim's session ends, which converts a transient phishing-grade interaction into standing out-of-band control of the node - power, console, virtual media, and a launch point for the firmware-level bugs elsewhere in this list. In fleets where one operator's browser session spans many BMCs, a single page load can seed accounts across dozens of controllers. BIOS 2.0a and IPMI firmware 03.40 - no CSRF protection on the call that creates administrator accounts.","attack_vector":"No network position on the management VLAN needed by the attacker directly - instead they need a BMC administrator with an active session to visit attacker-controlled content. That is an unusually low bar in datacenter operations, where staff routinely have BMC tabs open alongside general browsing.","remediation":"Firmware flash to BMC 03.88 or later and BIOS 3.2 or later on X10DRH-iT. Beyond the flash, the durable control is operational rather than technical: BMC administration should happen from a dedicated management workstation or bastion that does not browse the general internet, and BMC accounts should be audited on a schedule so that an injected administrator gets caught. Audit existing BMC user lists now - this bug leaves evidence, in the form of accounts nobody created.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15046","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2020/15xxx/CVE-2020-15046.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-17389","cve":"CVE-2020-17389","aliases":["ZDI QConvergeConsole"],"title":"Marvell QConvergeConsole (QLogic adapter management): Remote code execution on QConvergeConsole, the management application for QLogic/Marvell FastLinQ and Fibre…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Marvell QConvergeConsole (QLogic adapter management)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Remote code execution on QConvergeConsole, the management application for QLogic/Marvell FastLinQ and Fibre Channel adapters. QConvergeConsole is the tool that flashes adapter firmware and configures boot-from-SAN across a fleet, so code execution there is a route to pushing adapter firmware to every server it manages. One of a cluster of near-identical ZDI-reported issues (CVE-2020-17387, CVE-2020-17388, CVE-2020-15642 through CVE-2020-15645) in the same version.","attack_vector":"Remote attacker with authentication to the QConvergeConsole service — the advisory notes the existing authentication mechanism can be bypassed.","remediation":"Upgrade QConvergeConsole past 5.5.0.64. Application upgrade on the management host, no server or switch impact. Better: do not leave a fleet-wide adapter-management console running continuously — stand it up for firmware campaigns and shut it down afterwards, which is a process change that removes a permanently exposed high-value target.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-17389","https://nvd.nist.gov/vuln/detail/CVE-2020-15645"],"status":"curated"},{"id":"CVE-2020-5208","cve":"CVE-2020-5208","aliases":[],"title":"ipmitool (IPMI LAN response parsing): Reverses the usual direction of BMC risk: here the management station is the victim. ipmitool does not…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ipmitool (IPMI LAN response parsing)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Reverses the usual direction of BMC risk: here the management station is the victim. ipmitool does not validate data returned by the remote BMC, so a malicious or already-compromised BMC overflows buffers in the tool and executes code on the machine running it. Because operators run ipmitool from automation hosts, in loops, across the whole fleet, and usually as root, one compromised BMC escalates into ownership of the box that holds credentials for every other BMC. That is the fastest path from a single node to the entire management plane.","attack_vector":"Any BMC that your tooling talks to. A single compromised or spoofed BMC on the management network is enough - and BMC-to-management-host is a trust direction almost nobody models.","remediation":"Package update to ipmitool 1.8.19 or later on every management/automation host - no firmware, no reboot, just the package and restarting any long-running collector. Cheap fix, high leverage. Also worth reviewing whether your fleet automation really needs to run ipmitool as root, and whether the credentials it holds are scoped per rack rather than fleet-wide.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5208","https://github.com/ipmitool/ipmitool/security/advisories/GHSA-g659-9qxw-p7cp"],"status":"curated"},{"id":"CVE-2020-7526","cve":"CVE-2020-7526","aliases":["SEVD-2020-192-01"],"title":"APC PowerChute Business Edition (v9.0.x and earlier): PHYSICAL. PowerChute runs the shutdown script that fires when a UPS reports a power event. Improper input…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC PowerChute Business Edition (v9.0.x and earlier)","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"PHYSICAL. PowerChute runs the shutdown script that fires when a UPS reports a power event. Improper input validation means an attacker can get arbitrary code executed at exactly that moment - during a shutdown, with elevated privilege, on every host running the agent. It is a rare shape of bug: the trigger is a power event you cannot prevent, and the payload runs fleet-wide simultaneously.","attack_vector":"Requires the ability to influence the shutdown script content or the event that invokes it - which in practice means access to the PowerChute management server or the UPS that signals it.","remediation":"Upgrade PowerChute. Agent upgrade across every host that runs it, so this is a fleet-wide package push - schedulable, but it touches every node. Separately, treat shutdown scripts as privileged code: version them, restrict who can edit them, and do not let the UPS management network write to them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7526"],"status":"curated"},{"id":"CVE-2020-7569","cve":"CVE-2020-7569","aliases":["CVE-2020-7572","CVE-2020-7573","CVE-2020-28210"],"title":"Schneider Electric EcoStruxure Building Operation WebReports / WebStation V1.9-V3.1: Authenticated file upload with dangerous type on the WebReports component gives code execution on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric EcoStruxure Building Operation WebReports / WebStation V1.9-V3.1","year":"2020","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated file upload with dangerous type on the WebReports component gives code execution on the EcoStruxure Building Operation server, with an XXE and a broken access-control issue alongside it that help an attacker get there and read server-side files on the way. EBO is Schneider's BAS supervisory platform and it commonly integrates the cooling plant, metering and sometimes access control for the site. Code execution on the EBO server means write authority over the Automation Servers (AS/AS-P) beneath it, which are the devices actually commanding air handling. So the physical consequence is the same as any BAS-server compromise: setpoints and fan commands under attacker control, alarms suppressible, and a hall of high-density GPU racks minutes from thermal shutdown. Worth flagging that Schneider is also the vendor for a lot of the power side in the same buildings, so a compromised EBO server frequently sits inside the same trust boundary as the electrical monitoring.","attack_vector":"Requires an authenticated EBO session for the upload path, so the realistic chain is credential theft or a weak/default operator account followed by upload. The reflected/stored XSS and access-control issues in the same cluster provide the credential-capture step. EBO WebStation is browser-based and frequently published to the corporate network for facilities staff, which is where the credentials get phished.","remediation":"Software upgrade of the EcoStruxure Building Operation server and WebReports to a fixed release per Schneider's SEVD advisories - a server-side change in a normal window, no controller firmware, no cooling downtime. Also enforce MFA on the path to WebStation, remove shared facilities accounts, and take the server off the general corporate segment. In a leased colo the EBO server is the landlord's and shared across tenants: ask for the version, and note that 'authenticated' is a weak barrier when a dozen contractor accounts exist on that server.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7569","https://nvd.nist.gov/vuln/detail/CVE-2020-7572","https://www.se.com/ww/en/work/support/cybersecurity/security-notifications.jsp"],"status":"curated"},{"id":"CVE-2021-21974","cve":"CVE-2021-21974","aliases":[],"title":"VMware ESXi (OpenSLP): OpenSLP heap overflow - the ESXiArgs ransomware entry point that mass-encrypted thousands of hosts","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi (OpenSLP)","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"OpenSLP heap overflow - the ESXiArgs ransomware entry point that mass-encrypted thousands of hosts [KEV]","attack_vector":"Unauthenticated network on the same segment","remediation":"Patch + disable SLP. The canonical example of why the hypervisor management plane must never share a broadcast domain with tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21974"],"status":"curated"},{"id":"CVE-2021-25296","cve":"CVE-2021-25296","aliases":[],"title":"Nagios XI: OS command injection in the windowswmi config wizard (authenticated)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Nagios XI","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"[KEV] OS command injection in the windowswmi config wizard (authenticated) -> shell on the Nagios host","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; management tooling should be VPN-only","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25296"],"status":"curated"},{"id":"CVE-2021-25297","cve":"CVE-2021-25297","aliases":[],"title":"Nagios XI: OS command injection in the switch config wizard","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Nagios XI","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"[KEV] OS command injection in the switch config wizard -> server compromise","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25297"],"status":"curated"},{"id":"CVE-2021-25298","cve":"CVE-2021-25298","aliases":[],"title":"Nagios XI: OS command injection in the cloud-vm config wizard","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Nagios XI","year":"2021","cvss_score":8.8,"severity":"high","kev":true,"impact":"[KEV] OS command injection in the cloud-vm config wizard -> server compromise","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25298"],"status":"curated"},{"id":"CVE-2021-25741","cve":"CVE-2021-25741","aliases":[],"title":"Kubernetes (kubelet): subPath volume mount symlink race gives access to host files and directories outside the volume","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"subPath volume mount symlink race gives access to host files and directories outside the volume; host escape","attack_vector":"Cluster user able to create a pod with a subPath mount","remediation":"Rolling kubelet upgrade with node drain; interim mitigation is an admission policy banning subPath","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25741"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-31215","cve":"CVE-2021-31215","aliases":[],"title":"Slurm: Environment mishandling in PrologSlurmctld/EpilogSlurmctld gives remote code execution as SlurmUser","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Environment mishandling in PrologSlurmctld/EpilogSlurmctld gives remote code execution as SlurmUser","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm; audit prolog/epilog scripts","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-31215"],"status":"curated"},{"id":"CVE-2021-34824","cve":"CVE-2021-34824","aliases":[],"title":"Istio: Gateway/DestinationRule credentialName can read TLS secrets from other namespaces","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Gateway/DestinationRule credentialName can read TLS secrets from other namespaces; cross-tenant private key theft","attack_vector":"Cluster user with namespace access who can create Gateways","remediation":"Rolling istiod upgrade; rotate all mesh TLS keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34824"],"status":"curated"},{"id":"CVE-2021-3570","cve":"CVE-2021-3570","aliases":[],"title":"linuxptp / ptp4l (PTP message forwarding): A missing length check when ptp4l forwards a PTP message between ports leaks memory contents to a remote…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"linuxptp / ptp4l (PTP message forwarding)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"A missing length check when ptp4l forwards a PTP message between ports leaks memory contents to a remote attacker and can be pushed into a crash. PTP is the time source for AI clusters that do distributed tracing, RDMA telemetry correlation, or lockstep checkpointing — and ptp4l typically runs as root with raw socket access on every node. An information leak out of that process is a leak out of a root-privileged daemon reachable from the fabric.","attack_vector":"Remote — any host that can send PTP messages to a node running ptp4l as a boundary/transparent clock. PTP is unauthenticated by default, so no credentials are involved and any tenant on the same segment qualifies.","remediation":"Package upgrade of linuxptp and a restart of ptp4l — no reboot needed, and the restart costs a brief loss of clock discipline rather than a node outage. The durable control is to run PTP on a dedicated VLAN that tenant workloads cannot source traffic onto, which is a switch config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3570"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2021-36230","cve":"CVE-2021-36230","aliases":[],"title":"Terraform Enterprise: Missing authorization on a subset of run-token API requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Terraform Enterprise","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Missing authorization on a subset of run-token API requests -> privilege escalation to organization owner","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade TFE to v202107-1+; review org owner membership","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36230"],"status":"curated"},{"id":"CVE-2021-39298","cve":"CVE-2021-39298","aliases":[],"title":"AMD System Management Mode (SMM) interrupt handler: MULTI-TENANT ISOLATION: A flaw in the AMD SMM interrupt handler lets a high-privilege attacker reach System…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD System Management Mode (SMM) interrupt handler","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A flaw in the AMD SMM interrupt handler lets a high-privilege attacker reach System Management Mode and execute arbitrary code there. SMM sits above the hypervisor and above the OS - code running in SMM can read and write all physical memory including SEV-protected regions in some configurations, and it is invisible to every security tool you run. This is the classic 'ring -2' compromise: persistent, undetectable from the OS, and it survives reinstalling everything above it.","attack_vector":"Local, requires high privilege (root) on the host first. Not a tenant-reachable bug, but the payoff for an attacker who already has root is enormous - it converts a recoverable host compromise into an unrecoverable one.","remediation":"Fixed in AMD reference firmware (AGESA) and delivered only as an OEM SBIOS package - the OEM rebuild and requalification means **one to six months of lag**, and on end-of-support platforms possibly never. Applying it is a drain plus full power cycle. There is no OS-level mitigation for an SMM handler bug. On a bare-metal fleet, the compensating control is firmware measurement between tenants: if you cannot attest that SMM code is unchanged, you cannot honestly claim a node was cleaned by reimaging.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39298","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-43858","cve":"CVE-2021-43858","aliases":[],"title":"MinIO: Hand-crafted admin API call updates a user's policy","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Hand-crafted admin API call updates a user's policy -> privilege escalation","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate admin credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43858"],"status":"curated"},{"id":"CVE-2021-44142","cve":"CVE-2021-44142","aliases":[],"title":"Samba (SMB gateway): Out-of-bounds heap read/write in vfs_fruit","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba (SMB gateway)","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"Out-of-bounds heap read/write in vfs_fruit -> code execution as root on the file server","attack_vector":"Network (remote)","remediation":"Data-plane: emergency smbd upgrade on every SMB gateway/file node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-44142"],"status":"curated"},{"id":"CVE-2021-47390","cve":"CVE-2021-47390","aliases":[],"title":"Linux KVM x86 - stack out-of-bounds in ioapic_write_indirect(): MULTI-TENANT ISOLATION: A guest write to the virtual IOAPIC causes a stack out-of-bounds access in the host…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM x86 - stack out-of-bounds in ioapic_write_indirect()","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A guest write to the virtual IOAPIC causes a stack out-of-bounds access in the host kernel, reported by KASAN. At CVSS 8.8 this is a guest-to-host memory corruption primitive reachable by writing to an emulated device every VM has - stack corruption in the hypervisor is the shortest path from one tenant's VM to owning the machine and everything else on it.","attack_vector":"From inside a guest VM, by writing to the emulated IOAPIC. Tenant-reachable with no privilege beyond running a VM.","remediation":"Fixed in the Linux kernel. Distro kernel update plus host reboot - no firmware, no VBIOS. Treat as top priority on any host running untrusted guest VMs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47390"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-0811","cve":"CVE-2022-0811","aliases":[],"title":"CRI-O: \"cr8escape\": kernel sysctl injection via pod spec gives container escape and arbitrary code execution as root…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"\"cr8escape\": kernel sysctl injection via pod spec gives container escape and arbitrary code execution as root on the node","attack_vector":"Any cluster user who can deploy a pod","remediation":"Emergency CRI-O upgrade on all nodes; drain and recreate pods. Add admission policy blocking sysctl values with newlines","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0811"],"status":"curated","fleet":{"ubiquity":"Common - CRI-O is the default runtime on OpenShift and on several K8s distros neoclouds resell; less common than containerd but not niche","remediation_pain":"`daemon-restart` - upgrading CRI-O restarts the node's container runtime, which on most configurations restarts all pods on that node, so effectively `node-drain` for GPU workloads mid-training","pain_class":"node-drain","why_fleet_wide":"Anyone who can create a pod (i.e. any tenant with namespace access) sets arbitrary host kernel parameters via `sysctls`, abuses `kernel.core_pattern` and gets root code execution on *any* node in the cluster"}},{"id":"CVE-2022-1025","cve":"CVE-2022-1025","aliases":[],"title":"Argo CD: Improper access control lets any user escalate to admin-level","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper access control lets any user escalate to admin-level","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1025"],"status":"curated"},{"id":"CVE-2022-1227","cve":"CVE-2022-1227","aliases":[],"title":"Podman: Malicious image causes privilege escalation when a user runs `podman top`","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malicious image causes privilege escalation when a user runs `podman top`","attack_vector":"Malicious image","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1227"],"status":"curated"},{"id":"CVE-2022-1552","cve":"CVE-2022-1552","aliases":[],"title":"PostgreSQL: Autovacuum, REINDEX, CLUSTER etc. apply protections too late","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Autovacuum, REINDEX, CLUSTER etc. apply protections too late -> user code runs privileged","attack_vector":"Network (remote)","remediation":"Control-plane: minor-version upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1552"],"status":"curated"},{"id":"CVE-2022-2031","cve":"CVE-2022-2031","aliases":[],"title":"Samba (AD DC): KDC and kpasswd share keys","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba (AD DC)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"KDC and kpasswd share keys -> a user forced to change password can obtain tickets to other services","attack_vector":"Network (remote)","remediation":"Control-plane: DC-only upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2031"],"status":"curated"},{"id":"CVE-2022-20824","cve":"CVE-2022-20824","aliases":[],"title":"Cisco NX-OS / FXOS (Cisco Discovery Protocol): Root code execution on the switch from a crafted CDP frame sent by anything on an adjacent link. CDP is a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS / FXOS (Cisco Discovery Protocol)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Root code execution on the switch from a crafted CDP frame sent by anything on an adjacent link. CDP is a layer-2 protocol that is on by default and is not authenticated, so a compromised server NIC — or a tenant's bare-metal node — can attack the leaf it is cabled to directly. This is the classic 'the switch trusts the host' failure and it is why CDP/LLDP should be off on tenant-facing ports.","attack_vector":"Unauthenticated, adjacent — layer-2 reachability to the switch port. Any host on the link, including a tenant's own machine.","remediation":"NX-OS upgrade plus reload. Immediate config mitigation: disable CDP globally or per-interface on all host-facing ports (`no cdp enable`), which is a live change with no reload and is good hygiene independent of the CVE. Same treatment applies to the older CDP RCEs (CVE-2020-3119, CVE-2020-3172, CVE-2018-0303).","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-20824","https://nvd.nist.gov/vuln/detail/CVE-2020-3119"],"status":"curated"},{"id":"CVE-2022-23182","cve":"CVE-2022-23182","aliases":[],"title":"Intel Data Center Manager: Improper access control in Data Center Manager lets an unauthenticated attacker with adjacent network access…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center Manager","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Improper access control in Data Center Manager lets an unauthenticated attacker with adjacent network access escalate privilege. DCM is the fleet-wide power and thermal management plane - it holds credentials to platform management across every node it monitors, so compromising it is a lateral-movement jackpot rather than a single-host problem.","attack_vector":"Unauthenticated attacker on the same network segment as the DCM server. Whether that is a realistic position depends entirely on your management network segmentation.","remediation":"Upgrade the Intel Data Center Manager software. This is a management-plane application, so the update is an application upgrade and service restart - no node drain, no firmware, no reboot of managed hosts. The real work is deciding what DCM is allowed to reach: it holds credentials for platform power and telemetry across the fleet, so its network exposure matters more than its version. Upgrade to 4.1 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23182","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00662.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2022-24842","cve":"CVE-2022-24842","aliases":[],"title":"MinIO: Non-admin user can create service accounts for root/admin users and assume their policies","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Non-admin user can create service accounts for root/admin users and assume their policies","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + rotate all service-account credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24842"],"status":"curated"},{"id":"CVE-2022-28639","cve":"CVE-2022-28639","aliases":["HPESBHF04365"],"title":"HPE iLO 5 (adjacent-network code execution / DoS): Arbitrary code execution on the iLO from an adjacent network position, with denial of service as the softer…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 (adjacent-network code execution / DoS)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary code execution on the iLO from an adjacent network position, with denial of service as the softer outcome. Code execution on the service processor means the attacker owns power, Virtual Media, console and firmware for that node, and can leave an implant that outlives any host reinstall. The DoS variant is its own operational problem on a GPU fleet: losing iLO means losing the only way to power-cycle or console into a wedged training node, so an outage turns into a truck roll.","attack_vector":"Adjacent network - an attacker already on the same management segment as the iLO. That is a compromised jump host, a monitoring collector, another node's BMC, or anything else sharing the OOB VLAN. Not internet-reachable by design, but flat management networks make 'adjacent' mean 'the entire datacenter'.","remediation":"Flash iLO 5 to v2.72 or later (v2.71 and earlier are affected). Out-of-band, per-node, no host reboot and no job drain. The durable control is segmentation: this class of bug is only exploitable from the management network, so the value of putting each rack's BMCs behind their own segment with an explicit allowlist is high and it is a network change rather than a per-node campaign.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04365en_us","https://nvd.nist.gov/vuln/detail/CVE-2022-28639"],"status":"curated"},{"id":"CVE-2022-29178","cve":"CVE-2022-29178","aliases":[],"title":"Cilium: Incorrect default permissions on Cilium-managed host paths allow privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect default permissions on Cilium-managed host paths allow privilege escalation","attack_vector":"Any tenant workload on the node","remediation":"Rolling Cilium agent DaemonSet upgrade; brief per-node dataplane interruption but no GPU pod eviction","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29178"],"status":"curated"},{"id":"CVE-2022-29500","cve":"CVE-2022-29500","aliases":[],"title":"Slurm: Incorrect access control leading to information disclosure across users' jobs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect access control leading to information disclosure across users' jobs","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29500"],"status":"curated","fleet":{"ubiquity":"Very common - Slurm is the default scheduler on HPC-style GPU clouds and on most bare-metal H100/GB200 clusters sold to AI labs","remediation_pain":"`daemon-restart` fleet-wide - SchedMD is explicit that the cluster stays vulnerable until *every* slurmdbd, slurmctld and slurmd has restarted, i.e. a coordinated restart across all compute nodes","pain_class":"daemon-restart","why_fleet_wide":"Credential-handling flaw lets an unprivileged user impersonate SlurmUser and then run arbitrary processes as root; affects every Slurm release since 1.0.0, so one bug covers the whole scheduler estate"}},{"id":"CVE-2022-29501","cve":"CVE-2022-29501","aliases":[],"title":"Slurm: Incorrect access control leading to privilege escalation and code execution","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect access control leading to privilege escalation and code execution","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29501"],"status":"curated"},{"id":"CVE-2022-30243","cve":"CVE-2022-30243","aliases":["CVE-2022-30242","CVE-2022-30244","CVE-2022-30245"],"title":"Honeywell Alerton Visual Logic, Ascent Control Module (ACM) and Compass 1.6.5: Unauthenticated program writes to the controller. Not configuration - program. An attacker on the network can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Honeywell Alerton Visual Logic, Ascent Control Module (ACM) and Compass 1.6.5","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Unauthenticated program writes to the controller. Not configuration - program. An attacker on the network can push new control logic onto an Alerton controller and stop or replace the running program with no verification of who they are. This is the deepest form of BMS compromise available: the attacker is not sending a bad setpoint that an operator might notice and override, they are rewriting the control algorithm so the controller itself now does the wrong thing and reports the right thing. Applied to the controllers sequencing CRAHs or chilled water in a GPU hall, an attacker can write logic that holds fans low, ignores high-temperature alarms, or trips the plant on a delay so the failure looks like a mechanical fault. Thermal shutdown of a 40-140 kW rack row follows in minutes, and the forensics point at the HVAC contractor rather than at an intrusion. The companion issues let configuration be changed the same way, unauthenticated.","attack_vector":"A crafted packet from any host on the controller's network - no authentication exists on the programming path at all. Alerton gear sits on the building/facility VLAN, and the Alerton BACtalk ecosystem is BACnet-based, so anything that can route BACnet to the controller can do this. Physical access to a mechanical room panel is an equally valid path.","remediation":"Effectively unpatchable as a class - the vendor's guidance for this family is defensive positioning rather than a fix that adds authentication to the programming path, because the programming protocol was never designed with any. The only real control is to make the controllers unreachable: isolated VLAN carrying BACnet only, explicit allow-list from the Alerton supervisor (Compass/Envision) and nothing else, no internet path, port security on the switch ports feeding mechanical rooms, and physical locks on control panels. Where Honeywell offers a newer controller generation with authenticated programming, replacing the controllers is the actual fix and it is a capital project. In a leased site, you cannot touch these - require the landlord to attest that BACnet programming traffic is not routable from any tenant or corporate network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30243","https://nvd.nist.gov/vuln/detail/CVE-2022-30244","https://github.com/scadafence/Honeywell-Alerton-Vulnerabilities","https://www.honeywell.com/us/en/product-security"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-32744","cve":"CVE-2022-32744","aliases":[],"title":"Samba (AD DC): KDC accepts kpasswd requests encrypted with any key it knows","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba (AD DC)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"KDC accepts kpasswd requests encrypted with any key it knows -> change any user's password, full domain takeover","attack_vector":"Network (remote)","remediation":"Control-plane: DC upgrade; force org-wide credential reset","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32744"],"status":"curated"},{"id":"CVE-2022-33183","cve":"CVE-2022-33183","aliases":[],"title":"Brocade Fabric OS CLI: A remote authenticated attacker can act beyond their role through the Fabric OS CLI. SAN switches are…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS CLI","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"A remote authenticated attacker can act beyond their role through the Fabric OS CLI. SAN switches are routinely given shared operator accounts across a storage team, so role separation on FOS is thin to begin with; this removes what is left. Companion privilege escalation: CVE-2022-33182.","attack_vector":"Authenticated remote user on FOS before 9.1.0 / 9.0.1e / 8.2.3c / 8.2.0cbn5 / 7.4.2.j.","remediation":"Fabric OS upgrade plus reboot per fabric. Move to individual named accounts with RBAC roles rather than a shared switch login — a config and process change worth doing regardless.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33183","https://nvd.nist.gov/vuln/detail/CVE-2022-33182"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-34669","cve":"CVE-2022-34669","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: The user mode driver layer lets an unprivileged user read or modify files critical to the driver, which is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"The user mode driver layer lets an unprivileged user read or modify files critical to the driver, which is a straightforward path to SYSTEM on the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34669","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-73"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-36348","cve":"CVE-2022-36348","aliases":[],"title":"Intel Server Platform Services (SPS) firmware: Active debug code left enabled in shipped SPS firmware lets an authenticated user escalate privilege. SPS is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Platform Services (SPS) firmware","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Active debug code left enabled in shipped SPS firmware lets an authenticated user escalate privilege. SPS is the server variant of the management engine - it is the component running on your Xeon nodes, not the consumer CSME - so debug code in production firmware is squarely a datacenter problem.","attack_vector":"Authenticated local access on an affected server.","remediation":"Fixed in Intel SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36348","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00718.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-42309","cve":"CVE-2022-42309","aliases":["XSA-414"],"title":"Xen (xenstored): Guest can crash xenstored, taking down control-plane services for all guests on the host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (xenstored)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"Guest can crash xenstored, taking down control-plane services for all guests on the host","attack_vector":"Tenant VM guest","remediation":"Patch xenstored + restart; noisy-neighbour availability risk rather than a confidentiality break","references":["https://xenbits.xen.org/xsa/advisory-414.html"],"status":"curated"},{"id":"CVE-2022-4886","cve":"CVE-2022-4886","aliases":[],"title":"ingress-nginx: `log_format` directive bypasses path sanitization","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"`log_format` directive bypasses path sanitization; read arbitrary files including the SA token","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; scope the controller's RBAC down from cluster-wide secret read","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4886"],"status":"curated"},{"id":"CVE-2022-48864","cve":"CVE-2022-48864","aliases":["vdpa/mlx5 VIRTIO_NET_CTRL_MQ_VQ_PAIRS_SET missing validation"],"title":"Linux kernel drivers/vdpa/mlx5 (mlx5 vDPA net device): TENANT ISOLATION: a guest with an assigned mlx5 vDPA net device sends an unvalidated queue-pair-count control…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel drivers/vdpa/mlx5 (mlx5 vDPA net device)","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a guest with an assigned mlx5 vDPA net device sends an unvalidated queue-pair-count control command and panics the host kernel. CVSS scope is Changed - this is a guest breaking out of its own blast radius into the hypervisor. On a multi-tenant node using mlx5 vDPA for accelerated guest networking, one tenant VM takes down every other VM on that host.","attack_vector":"A malicious virtio driver inside a guest VM with an mlx5 vDPA device, or any local process with access to /dev/vhost-vdpa (typically the qemu/kvm group). No host root required.","remediation":"Upgrade the host kernel to 5.17, or a stable backport (5.15.29, 5.16.15) or your distro's patched kernel. Kernel upgrade means a rolling reboot of every hypervisor node using mlx5 vDPA, with live migration or workload drain per node. If you cannot reboot soon, stop exposing vDPA devices to untrusted guests - fall back to SR-IOV VFs or software virtio.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48864","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48864.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0184","cve":"CVE-2023-0184","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container; also vGPU guest","remediation":"Driver + vGPU Manager upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0184","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-822"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0189","cve":"CVE-2023-0189","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node, evict tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0189","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-822"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-22612","cve":"CVE-2023-22612","aliases":["INSYDE-SA-2023019"],"title":"Insyde InsydeH2O (IhisiSmm SMI handler): A malicious host OS calls an Insyde SMI handler with malformed arguments and corrupts SMM memory. IHISI is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IhisiSmm SMI handler)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"A malicious host OS calls an Insyde SMI handler with malformed arguments and corrupts SMM memory. IHISI is Insyde's own firmware-services interface, used by BIOS update and configuration tooling, so it is reachable by design from the OS - the bug is that it trusts what it is told. Successful exploitation is ring -2 code execution: firmware persistence, attestation you can no longer believe, and a foothold under the hypervisor.","attack_vector":"Local admin/root on the host OS invoking the IHISI SMI interface with crafted arguments.","remediation":"OEM BIOS update built on the fixed Insyde kernel (5.0-5.5 affected). Firmware flash plus one reboot per node. No config workaround - the interface is a supported firmware service and cannot be disabled. Reduce blast radius by restricting which host-side tools can issue SMIs and by not granting untrusted tenants root on bare metal running unpatched firmware. NCC Group published the underlying research, which is worth reading before assessing your exposure.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22612","https://www.insyde.com/security-pledge/SA-2023019","https://research.nccgroup.com/2023/04/11/stepping-insyde-system-management-mode/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-23583","cve":"CVE-2023-23583","aliases":[],"title":"Intel CPU (Reptar): Redundant REX-prefix MOVSB causes unpredictable behaviour - a guest can hang or crash the entire host, with…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Reptar)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Redundant REX-prefix MOVSB causes unpredictable behaviour - a guest can hang or crash the entire host, with possible privilege escalation","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode update + reboot. Pure availability risk for a neocloud: one tenant can take down a whole node","references":["https://access.redhat.com/security/cve/CVE-2023-23583"],"status":"curated","fleet":{"ubiquity":"very common - broad recent Intel Xeon coverage, the host CPU under most NVIDIA GPU nodes","remediation_pain":"microcode+reboot - Intel shipped microcode, but on many platforms it only lands via a BIOS/UEFI update from the OEM, adding an OEM-image dependency on top of the reboot","pain_class":"microcode + reboot","why_fleet_wide":"An unprivileged instruction sequence can hang or machine-check the host, or potentially escalate privilege - so a single tenant can crash a whole GPU node, and the same instruction works on every node of that Xeon generation."}},{"id":"CVE-2023-25194","cve":"CVE-2023-25194","aliases":[],"title":"Apache Kafka Connect: Attacker able to create/modify a connector sets a SASL JAAS JndiLoginModule config","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Kafka Connect","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Attacker able to create/modify a connector sets a SASL JAAS JndiLoginModule config -> RCE / JNDI SSRF","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade Connect workers; block arbitrary connector configuration","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25194"],"status":"curated"},{"id":"CVE-2023-25528","cve":"CVE-2023-25528","aliases":[],"title":"DGX H100 BMC (openBMC): RCE on BMC (web server plugin stack overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (openBMC)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE on BMC (web server plugin stack overflow)","attack_vector":"Network-adjacent attacker on the management LAN","remediation":"Flash BMC to 23.08.18 out-of-band; segregate BMC network; no tenant eviction required","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2023-25548","cve":"CVE-2023-25548","aliases":["SEVD-2023-101-04"],"title":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - device credential endpoints: Incorrect authorisation lets a low-privileged DCE user read device credentials from endpoints…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - device credential endpoints","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Incorrect authorisation lets a low-privileged DCE user read device credentials from endpoints that were never meant to expose them. The read-only NOC account you gave a monitoring contractor becomes the credential set for the entire power and cooling estate.","attack_vector":"Any low-privileged authenticated user on the DCE appliance - including accounts issued to third-party monitoring and maintenance vendors.","remediation":"Upgrade past V7.9.2. Then rotate device credentials and audit who holds DCE accounts. In most operators this audit is the finding: DCE accounts accumulate for vendors, integrators and former staff, and nobody owns the list.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25548"],"status":"curated"},{"id":"CVE-2023-28410","cve":"CVE-2023-28410","aliases":[],"title":"Intel i915 graphics driver for Linux (kernel < 6.2.10): MULTI-TENANT ISOLATION: A memory-buffer bounds failure in the i915 kernel driver that an authenticated local…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel i915 graphics driver for Linux (kernel < 6.2.10)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A memory-buffer bounds failure in the i915 kernel driver that an authenticated local user can drive into privilege escalation. i915 is the driver behind Intel integrated and Data Center GPU nodes, and it is reachable from inside any container that has been granted /dev/dri - so this is a container-to-host kernel escape on Intel-GPU nodes.","attack_vector":"Any local user or container with access to the DRM render node. No special hardware access beyond having been scheduled a GPU.","remediation":"Update to a Linux kernel with the fix (6.2.10 or later upstream, or your distro's backport) and reboot. i915 cannot be reloaded under running GPU workloads, so drain the node. No BIOS, firmware or microcode component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28410","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00886.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-28433","cve":"CVE-2023-28433","aliases":[],"title":"MinIO: Windows deployments fail to filter `\\`","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Windows deployments fail to filter `\\` -> arbitrary object placement across buckets by a low-privilege key","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; only relevant to Windows-hosted MinIO","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28433"],"status":"curated"},{"id":"CVE-2023-28434","cve":"CVE-2023-28434","aliases":[],"title":"MinIO: Crafted request bypasses PostPolicyBucket metadata bucket-name check","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2023","cvss_score":8.8,"severity":"high","kev":true,"impact":"[KEV] Crafted request bypasses PostPolicyBucket metadata bucket-name check -> object write to any bucket","attack_vector":"Network (remote)","remediation":"Control-plane: gateway upgrade; audit for cross-tenant object writes","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28434"],"status":"curated"},{"id":"CVE-2023-32460","cve":"CVE-2023-32460","aliases":["DSA-2023-361"],"title":"Dell PowerEdge Server BIOS (privilege management): An improper privilege-management flaw in PowerEdge BIOS that an unauthenticated local attacker can use to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS (privilege management)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An improper privilege-management flaw in PowerEdge BIOS that an unauthenticated local attacker can use to escalate. 'Unauthenticated local' is the sharp part: it does not require a valid OS login, so a tenant with any code execution path on the box, or someone with brief physical access during a rack move or RMA, can reach it. The prize is the firmware layer - an implant there sits below the hypervisor, survives reimaging and disk replacement, and is invisible to every host-level agent you run.","attack_vector":"Local to the host with no authentication required - a tenant workload that escapes its boundary, a technician with console access during maintenance, or anyone with the box open. Does not touch the management VLAN at all, so OOB network segmentation buys you nothing here.","remediation":"System BIOS update, which is materially more expensive than a BMC flash: the payload can be staged out-of-band through iDRAC or OME, but it only applies on the next host reboot. That means draining running training jobs, or waiting for a natural maintenance window - realistically a scheduled rolling campaign across the fleet, not a same-day fix. Version floors differ per platform; take them from the advisory's per-model table. No config-only mitigation.","references":["https://www.dell.com/support/kbdoc/en-us/000219550/dsa-2023-361-security-update-for-dell-poweredge-server-bios-for-an-improper-privilege-management-security-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2023-32460"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-33412","cve":"CVE-2023-33412","aliases":[],"title":"Supermicro BMC web interface CGI endpoints on X11 and M11 based boards with BMC firmware before 3.17.02: An authenticated BMC user runs arbitrary commands on the controller. Because X11 boards are…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web interface CGI endpoints on X11 and M11 based boards with BMC firmware before 3.17.02","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated BMC user runs arbitrary commands on the controller. Because X11 boards are the ones most likely to be running years-behind firmware in a rented or secondhand fleet, this is a realistic path from 'we gave the customer an IPMI login' to full out-of-band ownership of the node: power, console, virtual media, and firmware persistence. Any operator who has ever handed a tenant a BMC account on X11 hardware should assume this is exploitable. The X11 generation is the bulk of the older Supermicro GPU and storage fleet still in production at neoclouds.","attack_vector":"An authenticated BMC session over the network, at ordinary user privilege rather than administrator. Reachable by anything routable to the OOB management VLAN that holds any valid credential.","remediation":"Firmware flash to BMC 3.17.02 or later, per board, from Supermicro's December 2023 BMC advisory. On X11 boards this is often a multi-hop upgrade because the fleet is far behind, and some very old X11 SKUs stopped receiving images - for those the honest answer is that there is no fix and the only control is network isolation plus never issuing tenant-facing BMC accounts. Revoke any BMC credentials previously issued to tenants or contractors as part of the same change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-33412","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/33xxx/CVE-2023-33412.json"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-33413","cve":"CVE-2023-33413","aliases":[],"title":"Supermicro BMC configuration functionality on X11 and M11 based boards through firmware 3.17.02: Arbitrary command execution on the BMC from an authenticated session, reached through configuration…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC configuration functionality on X11 and M11 based boards through firmware 3.17.02","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary command execution on the BMC from an authenticated session, reached through configuration rather than a web handler. The operational significance of it being a separate path is that disabling or firewalling one interface does not close both - an operator who patched around CVE-2023-33412 without flashing still has this one. Outcome is BMC-level control of the node: power, boot device, console, and persistent firmware residency. A second, distinct command-execution path from the same December 2023 disclosure, this one in the settings surface rather than the CGI endpoints.","attack_vector":"An authenticated remote BMC user reaching the controller's configuration surface over the management network, at ordinary user privilege rather than administrator. Any valid credential on an X11 or M11 BMC below firmware 3.17.02 is enough, which includes the read-only accounts operators hand to monitoring systems and the shared credentials that most whitebox fleets still use across every node.","remediation":"Firmware flash to 3.17.02 or later per board SKU from the December 2023 Supermicro BMC advisory - the same image that fixes CVE-2023-33411 and CVE-2023-33412, so treat all three as one flash campaign rather than three. For X11 boards past end of firmware support, network isolation of the management VLAN is the only remaining control, and that should be documented as an accepted permanent risk rather than a temporary workaround.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-33413","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/33xxx/CVE-2023-33413.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-3676","cve":"CVE-2023-3676","aliases":[],"title":"Kubernetes: Command injection via pod spec on Windows nodes","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Command injection via pod spec on Windows nodes; escalation to node admin","attack_vector":"Cluster user able to create pods on a Windows node","remediation":"Rolling control-plane and kubelet upgrade; Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3676"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-3893","cve":"CVE-2023-3893","aliases":[],"title":"kubernetes-csi-proxy: Insufficient input sanitisation in csi-proxy leads to Windows node admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"kubernetes-csi-proxy","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Insufficient input sanitisation in csi-proxy leads to Windows node admin","attack_vector":"Cluster user able to create pods with CSI volumes on Windows","remediation":"Upgrade csi-proxy; Windows node drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-3955","cve":"CVE-2023-3955","aliases":[],"title":"Kubernetes: Second Windows-node input-sanitisation escalation to admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Second Windows-node input-sanitisation escalation to admin","attack_vector":"Cluster user able to create pods on a Windows node","remediation":"Rolling upgrade; Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-3955"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-42130","cve":"CVE-2023-42130","aliases":["ZDI-23-1496"],"title":"A10 Thunder ADC (FileMgmtExport): An authenticated attacker can walk outside the intended export directory via the FileMgmtExport component…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"A10 Thunder ADC (FileMgmtExport)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated attacker can walk outside the intended export directory via the FileMgmtExport component, reading or deleting arbitrary files on the ADC — enough to pull sensitive configuration data or sabotage the device by deleting files it depends on.","attack_vector":"Requires authentication to the management interface, then a crafted path sent to the file-export functionality.","remediation":"Software upgrade to the fixed ACOS release per A10's security advisory. Standard upgrade-and-reboot per Thunder ADC instance; coordinate with failover if this device is in the active traffic path for a cluster.","references":["https://support.a10networks.com/support/security_advisory/a10-acos-file-access-vulnerability/","https://www.zerodayinitiative.com/advisories/ZDI-23-1496/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-45234","cve":"CVE-2023-45234","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (DHCPv6 DNS Servers option handling): A crafted DNS Servers option inside a DHCPv6 Advertise overflows a firmware buffer, giving memory corruption…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (DHCPv6 DNS Servers option handling)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"A crafted DNS Servers option inside a DHCPv6 Advertise overflows a firmware buffer, giving memory corruption and a plausible route to pre-boot code execution. Same class of loss as the Server ID overflow: an attacker who lands here is executing inside DXE, above the OS and outside anything the tenant's EDR or attestation agent can see.","attack_vector":"Unauthenticated attacker able to answer DHCPv6 on the provisioning network during the node's PXE boot. No physical access, no host credentials.","remediation":"Firmware flash from the server OEM, not from Tianocore - the upstream edk2 patch has to be rebased by your IBV and then re-qualified by the OEM, which historically takes one to two BIOS release cycles. One reboot per node. Immediate config workaround: disable IPv6 in the UEFI network boot stack, or disable network boot on nodes that do not need it, and restrict who can emit DHCPv6/RA on the deployment VLAN.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45234","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45235","cve":"CVE-2023-45235","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (DHCPv6 proxy Advertise, Server ID option): Buffer overflow in the proxy-DHCPv6 path - the exact path a PXE/HTTP-boot provisioning flow uses when the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (DHCPv6 proxy Advertise, Server ID option)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Buffer overflow in the proxy-DHCPv6 path - the exact path a PXE/HTTP-boot provisioning flow uses when the boot server and the address server are different boxes. Memory corruption in DXE means the attacker can potentially take the node before it has an OS, which for a multi-tenant bare-metal GPU fleet means a persistent foothold that outlives the tenant lease and the reimage.","attack_vector":"Unauthenticated, on-link attacker impersonating or racing the proxy DHCPv6 server on the provisioning segment. Pre-OS.","remediation":"OEM BIOS update, flash and reboot each node. There is no OS-level patch and no runtime mitigation - the vulnerable code runs before the OS. Until the OEM ships, the practical controls are network-side: segment the provisioning VLAN away from tenant traffic, enforce DHCPv6 guard on access ports, and disable UEFI network boot on any node whose boot source is local NVMe.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45235","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-46229","cve":"CVE-2023-46229","aliases":[],"title":"LangChain (recursive URL loader): SSRF — crawling proceeds to internal hosts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (recursive URL loader)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"SSRF — crawling proceeds to internal hosts","attack_vector":"Attacker-supplied or crawled external page","remediation":"Upgrade past 0.0.317; egress-restrict RAG ingestion workers","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46229"],"status":"curated"},{"id":"CVE-2023-49935","cve":"CVE-2023-49935","aliases":[],"title":"Slurm: slurmd message-integrity bypass permits reuse of root-level authentication tokens","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"slurmd message-integrity bypass permits reuse of root-level authentication tokens","attack_vector":"Any user who can reach slurmd","remediation":"Upgrade Slurm; rotate munge keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49935"],"status":"curated"},{"id":"CVE-2023-5178","cve":"CVE-2023-5178","aliases":[],"title":"Linux NVMe-oF (nvmet-tcp): Use-after-free/double-free in nvmet_tcp_free_crypto","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux NVMe-oF (nvmet-tcp)","year":"2023","cvss_score":8.8,"severity":"high","kev":false,"impact":"Use-after-free/double-free in nvmet_tcp_free_crypto -> may permit remote code execution on the target","attack_vector":"Network (remote)","remediation":"Data-plane: kernel patch on NVMe-oF targets; storage-fabric maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5178"],"status":"curated"},{"id":"CVE-2024-0130","cve":"CVE-2024-0130","aliases":[],"title":"UFM Enterprise / UFM Appliance / UFM CyberAI: Fabric-manager privesc, data corruption, service disruption via improper authentication on the Ethernet mgmt…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"UFM Enterprise / UFM Appliance / UFM CyberAI","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Fabric-manager privesc, data corruption, service disruption via improper authentication on the Ethernet mgmt interface","attack_vector":"Network-adjacent attacker on the fabric management network","remediation":"Upgrade UFM; isolate the UFM mgmt interface to a dedicated VLAN; no tenant eviction, but InfiniBand fabric control is at stake","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=NVIDIA%20UFM"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-287"]},{"id":"CVE-2024-10979","cve":"CVE-2024-10979","aliases":[],"title":"PostgreSQL: PL/Perl lets an unprivileged DB user change process env vars (e.g. PATH)","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"PL/Perl lets an unprivileged DB user change process env vars (e.g. PATH) -> arbitrary code execution","attack_vector":"Network (remote)","remediation":"Control-plane: minor-version upgrade of the control-plane DB","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10979"],"status":"curated"},{"id":"CVE-2024-21802","cve":"CVE-2024-21802","aliases":[],"title":"llama.cpp / GGUF library: Heap buffer overflow in GGUF `info","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF library","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap buffer overflow in GGUF `info->ne` parsing","attack_vector":"Customer-supplied GGUF model file","remediation":"Rebuild any llama.cpp-derived binary; no host patch exists because the parser is statically linked into each tenant build","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21802"],"status":"curated"},{"id":"CVE-2024-21807","cve":"CVE-2024-21807","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Improper initialisation in the Linux kernel-mode driver for Intel 800-series…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper initialisation in the Linux kernel-mode driver for Intel 800-series Ethernet, reachable by an authenticated user for privilege escalation. The 800 series (E810) is the NIC under most RoCE/RDMA AI fabrics, so a kernel-mode driver escalation here is host compromise reached from whoever can talk to the network stack - and on nodes exposing SR-IOV VFs or RDMA verbs to tenants, that includes tenants.","attack_vector":"Authenticated local user; on nodes that expose VFs or RDMA devices into containers, that extends to tenant workloads.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21807","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-21810","cve":"CVE-2024-21810","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (improper input…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (improper input validation) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21810","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-21825","cve":"CVE-2024-21825","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `GGUF_TYPE_ARRAY`/`GGUF_TYPE_STRING` parsing","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `GGUF_TYPE_ARRAY`/`GGUF_TYPE_STRING` parsing","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild from a patched commit","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21825"],"status":"curated"},{"id":"CVE-2024-21836","cve":"CVE-2024-21836","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `header.n_tensors` handling","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `header.n_tensors` handling","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21836"],"status":"curated"},{"id":"CVE-2024-23496","cve":"CVE-2024-23496","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `gguf_fread_str`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `gguf_fread_str`","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23496"],"status":"curated"},{"id":"CVE-2024-23497","cve":"CVE-2024-23497","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (an out-of-bounds write) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23497","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-23605","cve":"CVE-2024-23605","aliases":[],"title":"llama.cpp / GGUF: Heap overflow in `header.n_kv`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp / GGUF","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Heap overflow in `header.n_kv`","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23605"],"status":"curated"},{"id":"CVE-2024-23898","cve":"CVE-2024-23898","aliases":[],"title":"Jenkins: No origin validation on the CLI WebSocket endpoint","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"No origin validation on the CLI WebSocket endpoint -> cross-site WebSocket hijacking, attacker runs CLI commands","attack_vector":"Network (remote)","remediation":"Control-plane: same patch window as CVE-2024-23897","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23898"],"status":"curated"},{"id":"CVE-2024-23918","cve":"CVE-2024-23918","aliases":[],"title":"Intel Xeon memory controller configuration (with SGX): MULTI-TENANT ISOLATION: An improper conditions check in Xeon memory controller configuration under SGX gives…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon memory controller configuration (with SGX)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An improper conditions check in Xeon memory controller configuration under SGX gives a privileged local user privilege escalation. Highest-scored of the memory-controller-plus-SGX family and the one to prioritise if you run SGX on Xeon.","attack_vector":"Privileged local access on the host.","remediation":"OEM platform BIOS update plus TCB recovery and re-attestation. Drain and reboot per node; OEM availability is the long pole.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23918","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01079.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-23981","cve":"CVE-2024-23981","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (a wrap-around…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (a wrap-around (integer) error) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23981","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-24747","cve":"CVE-2024-24747","aliases":[],"title":"MinIO: Access keys inherit the parent's `admin:*` actions, not just `s3:*`","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Access keys inherit the parent's `admin:*` actions, not just `s3:*` -> silent admin rights","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; explicitly deny admin actions in the access-key hierarchy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24747"],"status":"curated"},{"id":"CVE-2024-24986","cve":"CVE-2024-24986","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode driver flaw in the Intel 800-series Ethernet Linux driver (an access-control failure) giving an authenticated user privilege escalation. Same family and same fix train as the other 2024 ice driver issues; the reason to care is that E810 is the fabric NIC on most GPU nodes and its driver runs in the kernel on the host.","attack_vector":"Authenticated local user, including tenant workloads on nodes that map VFs or RDMA devices into containers.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24986","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-25744","cve":"CVE-2024-25744","aliases":["Heckler"],"title":"Linux guest kernel - hypervisor-injected int 0x80 on the 32-bit syscall path (SEV-SNP / SEV-ES, AMD-SB-3008): MULTI-TENANT ISOLATION: The higher-scoring half of Heckler. A malicious hypervisor…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux guest kernel - hypervisor-injected int 0x80 on the 32-bit syscall path (SEV-SNP / SEV-ES, AMD-SB-3008)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The higher-scoring half of Heckler. A malicious hypervisor injects interrupt 0x80 - the legacy 32-bit syscall gate - into a confidential guest at a chosen instruction, and the guest kernel services it as a real syscall. The researchers turned this into an OpenSSH authentication bypass and a sudo-to-root escalation *inside* the confidential VM, with the host never touching guest memory. At 8.8 with a changed scope this is the most severe confidential-guest issue in the set, and it also covers Intel TDX, so it is a property of the interrupt-delivery design rather than an AMD-only bug.","attack_vector":"Malicious or compromised hypervisor injecting interrupts into its own guest. No guest vulnerability and no tenant cooperation needed.","remediation":"**Guest kernel fix, not a host fix** - Linux 6.6.7 / 6.9 and later carry the int80 hardening series. That inverts the usual rollout: patching every hypervisor you own does nothing, because the protection has to live in the tenant's VM image. As the operator, publish a minimum guest kernel for confidential workloads and gate admission on it. A quicker guest-side control is to disable 32-bit x86 emulation in the guest kernel entirely, which removes the int 0x80 gate. The hardware answer - protected/restricted interrupt delivery - had no mainline Linux support at disclosure. No host reboot, no firmware, no BIOS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25744","https://ahoi-attacks.github.io/heckler/","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3008.html"],"status":"curated"},{"id":"CVE-2024-2961","cve":"CVE-2024-2961","aliases":[],"title":"glibc (iconv): Out-of-bounds write in the ISO-2022-CN-EXT iconv converter - turns PHP/app file-read bugs into RCE","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"glibc (iconv)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in the ISO-2022-CN-EXT iconv converter - turns PHP/app file-read bugs into RCE","attack_vector":"Unauthenticated network (via an app) or local user","remediation":"Package update + restart all consuming services","references":["https://access.redhat.com/security/cve/CVE-2024-2961"],"status":"curated"},{"id":"CVE-2024-30368","cve":"CVE-2024-30368","aliases":["ZDI-24-524"],"title":"A10 Thunder ADC (CsrRequestView): An authenticated attacker can inject a system-call payload through the CsrRequestView component (used for…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"A10 Thunder ADC (CsrRequestView)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"An authenticated attacker can inject a system-call payload through the CsrRequestView component (used for certificate-signing-request handling), running arbitrary code on the load balancer with the privileges of the vulnerable process.","attack_vector":"Requires authentication first — the advisory doesn't specify a high privilege tier, meaning even a lower-privileged operator account may be enough to trigger it.","remediation":"Software upgrade to the fixed ACOS release per A10's advisory for CVE-2024-30368/CVE-2024-30369 (the two ship together). Upgrade and reboot each Thunder ADC instance; if it's fronting inference traffic, plan for a failover to a standby unit during the upgrade rather than a hard outage.","references":["https://support.a10networks.com/support/security_advisory/cve-2024-30368-cve-2024-30369","https://www.zerodayinitiative.com/advisories/ZDI-24-524/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-31856","cve":"CVE-2024-31856","aliases":[],"title":"CyberPower PowerPanel MQTT message handling: An attacker with MQTT publish permissions can craft messages to all managed PowerPanel devices, achieving SQL…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel MQTT message handling","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker with MQTT publish permissions can craft messages to all managed PowerPanel devices, achieving SQL injection, arbitrary file write and remote code execution. Combined with the shared-certificate flaw, obtaining those permissions is not a high bar. One compromised device becomes code execution across the power-management estate.","attack_vector":"Anyone able to publish on the PowerPanel MQTT broker - reachable via the shared device certificate, or from any device already on the management network.","remediation":"Vendor upgrade. Separately, put the MQTT broker on a segment reachable only by the devices that must use it, and monitor for publishers you do not recognise. MQTT wildcards should be blocked at the broker (see the companion issue CVE-2024-31409).","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31856"],"status":"curated"},{"id":"CVE-2024-36324","cve":"CVE-2024-36324","aliases":[],"title":"AMD Graphics Driver - crafted pointer leading to arbitrary code execution: MULTI-TENANT ISOLATION: Improper input validation in the AMD graphics driver lets an attacker supply a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Graphics Driver - crafted pointer leading to arbitrary code execution","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper input validation in the AMD graphics driver lets an attacker supply a crafted pointer and reach arbitrary code execution. At 8.8 this is the most severe of AMD's own graphics-driver advisories in the window: a pointer that crosses the driver boundary unvalidated means kernel-level execution from whatever context can issue the call.","attack_vector":"Local, via the graphics driver interface - reachable by a process holding the GPU device.","remediation":"Update the AMD graphics driver package, then reload the driver or reboot. Confirm the fixed version is present in the ROCm/amdgpu build you actually deploy, since the packaged AMD driver and the mainline kernel driver move on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36324","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-37032","cve":"CVE-2024-37032","aliases":["Probllama"],"title":"Ollama: Path traversal in the digest field","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path traversal in the digest field → arbitrary file overwrite → RCE","attack_vector":"Unauthenticated network to the Ollama API (`/api/pull` from an attacker-controlled registry)","remediation":"Upgrade to 0.1.34+. Ollama binds 11434 with no auth by default — never expose to a tenant network","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37032"],"status":"curated","fleet":{"ubiquity":"Common - Ollama is the standard single-node LLM server on GPU dev boxes and small inference tenants; Docker installs run the API as root","remediation_pain":"`daemon-restart` - upgrade to 0.1.34+; trivial on the daemon, but every compromised host is root-owned and needs rebuild","pain_class":"daemon-restart","why_fleet_wide":"Unvalidated `digest` in OCI manifests gives path traversal; a rogue registry writes `/etc/ld.so.preload` and gets unauthenticated RCE as root - a poisoned model pull compromises every node that pulled it"}},{"id":"CVE-2024-37052","cve":"CVE-2024-37052","aliases":[],"title":"MLflow (model flavors): Deserialization RCE from a maliciously uploaded model (one of a family: 37052–37060)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (model flavors)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Deserialization RCE from a maliciously uploaded model (one of a family: 37052–37060)","attack_vector":"Customer-supplied model artifact loaded via `mlflow.pyfunc.load_model`","remediation":"No format fix — every MLflow model flavor wraps pickle. Restrict who can register models and sandbox model-load workers","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37052"],"status":"curated"},{"id":"CVE-2024-37061","cve":"CVE-2024-37061","aliases":[],"title":"MLflow (recipes / pyfunc): RCE via a maliciously crafted MLproject","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (recipes / pyfunc)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via a maliciously crafted MLproject","attack_vector":"Customer-supplied MLproject/recipe run by the platform","remediation":"Upgrade. If the provider runs a managed MLflow, tenant projects execute in the provider plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37061"],"status":"curated"},{"id":"CVE-2024-43044","cve":"CVE-2024-43044","aliases":[],"title":"Jenkins: Agent processes can read arbitrary controller files via ClassLoaderProxy#fetchJar","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Agent processes can read arbitrary controller files via ClassLoaderProxy#fetchJar","attack_vector":"Network (remote)","remediation":"Control-plane: controller upgrade; treat build agents as untrusted","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43044"],"status":"curated"},{"id":"CVE-2024-48013","cve":"CVE-2024-48013","aliases":[],"title":"Dell SmartFabric OS10 (execution with unnecessary privileges): A low-privileged attacker escalates through an OS10 component running with more privilege than it needs. OS10…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (execution with unnecessary privileges)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A low-privileged attacker escalates through an OS10 component running with more privilege than it needs. OS10 is a Linux-based NOS, so escalation here is root on a box that programs the forwarding ASIC — arbitrary control over which tenant's traffic goes where. Sits alongside a long series of OS10 command-injection findings (CVE-2024-48830, CVE-2024-49557, CVE-2024-49560, CVE-2025-22472, CVE-2025-22473, CVE-2025-46427, CVE-2025-46428) that all give a low-privileged local or remote user a path to root.","attack_vector":"Low-privileged attacker with access to the switch, versions 10.5.4.x through 10.6.0.x.","remediation":"OS10 upgrade plus reload. Because so many of these share the same precondition — a low-privileged account on the switch — the highest-leverage control is eliminating low-privilege switch accounts entirely and driving all changes through an automation account on a bastion.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48013","https://nvd.nist.gov/vuln/detail/CVE-2025-46427","https://nvd.nist.gov/vuln/detail/CVE-2025-46428"],"status":"curated"},{"id":"CVE-2024-49559","cve":"CVE-2024-49559","aliases":[],"title":"Dell SmartFabric OS10 (default password): A default password in SmartFabric OS10 across 10.5.4.x through 10.6.0.x, usable remotely by a low-privileged…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (default password)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A default password in SmartFabric OS10 across 10.5.4.x through 10.6.0.x, usable remotely by a low-privileged attacker to escalate. Default credentials on a datacenter switch are the same failure the BMC world has been fighting for a decade — the switch arrives with a working account nobody in the deployment checklist knows to remove.","attack_vector":"Low-privileged attacker with remote access to the switch.","remediation":"OS10 upgrade plus reload. Also add a switch-intake step that enumerates and disables every non-provisioned local account, the same way you rotate BMC credentials at rack intake — a process change that catches the next one of these before an advisory does.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49559"],"status":"curated"},{"id":"CVE-2024-50627","cve":"CVE-2024-50627","aliases":[],"title":"Digi ConnectPort LTS (before 1.4.12): An attacker who can reach the ConnectPort LTS's file-upload feature with only limited/local-network…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Digi ConnectPort LTS (before 1.4.12)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker who can reach the ConnectPort LTS's file-upload feature with only limited/local-network privileges can upload and execute a malicious file, escalating straight to full control of the device — which puts them on the cellular/serial gateway path some fleets use for out-of-band access when the primary network is down.","attack_vector":"Requires local-area-network reach to the device's web upload feature and some baseline permission level (not full admin); no internet-facing exposure needed if the device is on a segmented OOB VLAN, but that's exactly the network this appliance is meant to be reachable from.","remediation":"Software upgrade to ConnectPort LTS firmware 1.4.12 or later. This is a firmware flash per unit; because ConnectPort LTS is frequently the fallback OOB path when the primary network is unavailable, schedule the upgrade during planned maintenance rather than waiting for an outage to force it.","references":["https://www.digi.com/getattachment/Resources/Security/Alerts/Digi-ConnectPort-LTS-Firmware-Update/ConnectPort-LTS-KB.pdf","https://www.digi.com/resources/security"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-5187","cve":"CVE-2024-5187","aliases":[],"title":"ONNX (`download_model_with_test_data`): Arbitrary file overwrite from a crafted model archive","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`download_model_with_test_data`)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary file overwrite from a crafted model archive","attack_vector":"Customer-supplied model URL/archive","remediation":"Upgrade; do not run model-zoo download helpers as root or with shared mounts","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5187"],"status":"curated"},{"id":"CVE-2024-53224","cve":"CVE-2024-53224","aliases":[],"title":"Linux kernel mlx5_ib (pkey change notifier): A race between InfiniBand device deregistration and the pkey-change work item leaves the handler running…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (pkey change notifier)","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"A race between InfiniBand device deregistration and the pkey-change work item leaves the handler running after the device is gone, causing a NULL dereference. Pkey changes are pushed by the subnet manager, so a fabric-side event - benign reconfiguration or a hostile SM - can crash hosts that are cycling their RDMA devices.","attack_vector":"Adjacent, unauthenticated: an actor able to cause partition-key change events on the subnet (a compromised or spoofed subnet manager) combined with device teardown on the target.","remediation":"Upgrade the host kernel to 6.13 or a stable backport (6.6.64, 6.11.11, 6.12.2). Rolling reboot of RDMA hosts. Complementary control: lock down who can run a subnet manager on the fabric and set SM priority so a rogue SM cannot take over.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53224","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-53224.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-5948","cve":"CVE-2024-5948","aliases":["CVE-2024-5947","CVE-2024-5949","CVE-2024-5950","CVE-2024-5951","CVE-2024-5952","ZDI-24-672"],"title":"Deep Sea Electronics DSE855 generator communications gateway: Six unauthenticated flaws in one device: two stack-based buffer overflows giving remote code execution, an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Deep Sea Electronics DSE855 generator communications gateway","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Six unauthenticated flaws in one device: two stack-based buffer overflows giving remote code execution, an infinite-loop DoS, a configuration-backup disclosure that leaks the device's stored settings and credentials, and missing authentication on both factory-reset and restart. The reset and restart issues deserve particular attention because they need no exploitation skill at all - a network-adjacent attacker simply asks the gateway to factory-reset itself, and the link between the generator controllers and the monitoring system is gone along with its configuration. Code execution gives persistence on a device sitting inside the electrical infrastructure segment. For an operator running GPU racks, the consequence is the same class as any standby-power monitoring failure: you find out your generators did not start when the hall goes dark and the cooling stops, and the accelerators cross thermal shutdown minutes later. The configuration disclosure additionally hands over credentials that are usually reused across the site's other DSE and BMS gear.","attack_vector":"Network-adjacent and unauthenticated for all six - no credentials required for any of them, including the destructive ones. The device's web service on the facility or electrical-infrastructure VLAN is the entire attack surface. As with the newer DSE855 issue, the generator contractor's remote-monitoring path is the most likely route in from outside.","remediation":"Firmware update from Deep Sea Electronics. The unit is small and the flash is fast, but ownership is the friction: these are usually specified, installed and maintained by the generator vendor, not by the datacenter's own team, so the change has to go through that contract. Given that a factory reset can be triggered by anyone who can reach the device, an ACL restricting the gateway's web port to the monitoring server is a same-day control worth taking regardless of patch status. Keep an offline copy of the gateway configuration so a triggered factory reset is a ten-minute restore rather than a contractor callout.","references":["https://www.zerodayinitiative.com/advisories/ZDI-24-672/","https://www.zerodayinitiative.com/advisories/ZDI-24-675/","https://nvd.nist.gov/vuln/detail/CVE-2024-5948","https://nvd.nist.gov/vuln/detail/CVE-2024-5951"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-6983","cve":"CVE-2024-6983","aliases":[],"title":"LocalAI: RCE — the backend accepts inputs beyond the config path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LocalAI","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE — the backend accepts inputs beyond the config path","attack_vector":"Unauthenticated/low-privilege network to the LocalAI API","remediation":"Upgrade past 2.17.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6983"],"status":"curated"},{"id":"CVE-2024-7348","cve":"CVE-2024-7348","aliases":[],"title":"PostgreSQL: TOCTOU race in pg_dump","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"TOCTOU race in pg_dump -> an object creator runs arbitrary SQL as the (often superuser) dump operator","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; stop running scheduled pg_dump as superuser","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7348"],"status":"curated"},{"id":"CVE-2024-7646","cve":"CVE-2024-7646","aliases":[],"title":"ingress-nginx: Annotation validation bypass reaching config injection and cluster-wide secret access","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2024","cvss_score":8.8,"severity":"high","kev":false,"impact":"Annotation validation bypass reaching config injection and cluster-wide secret access","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade, no GPU drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-0657","cve":"CVE-2025-0657","aliases":["CVE-2025-0658"],"title":"Automated Logic / Carrier i-Vu Gen5 BACnet router (drv_gen5_106-01-2380) and i-Vu Zone Controller: Malformed BACnet MS/TP frames put the router and the zone controllers into a fault state, and the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Automated Logic / Carrier i-Vu Gen5 BACnet router (drv_gen5_106-01-2380) and i-Vu Zone Controller","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malformed BACnet MS/TP frames put the router and the zone controllers into a fault state, and the vendor is explicit that recovery requires a manual power cycle - on the zone controller a second packet after reset leaves it permanently unresponsive until someone physically touches it. That is the important operator detail: this is not a reboot-and-recover DoS, it is a bricking-until-truck-roll DoS against the devices that command air handling in the hall. Lose the zone controllers and the affected zone stops modulating; depending on the failsafe wiring you either get fans stuck at last-known state or dampers closed. Either way you have lost closed-loop thermal control over a GPU hall and you are now running on whatever the mechanical failsafe does, with a technician on the way. For a 40 kW+ rack density that is a race you can lose. Expect this to hit during the worst possible moment because an attacker will fire it while the plant is already at high load.","attack_vector":"An attacker on the BACnet MS/TP segment, or on any BACnet/IP segment that routes onto it. MS/TP is an RS-485 serial bus, so pure MS/TP access means physical proximity to the field wiring - but the Gen5 router exists precisely to bridge IP to MS/TP, and that is the exposed side. Anyone who can send BACnet/IP to the router can reach the serial devices behind it. On the facility VLAN this needs no credentials because BACnet has no authentication to begin with.","remediation":"Driver/firmware update from Automated Logic or Carrier for the Gen5 router and controllers. Realistically this means the controls contractor on site with a laptop touching every controller, during a maintenance window, on live cooling - which is exactly the work most operators defer. Because the fix is slow, the compensating control is the one that matters: no arbitrary host should be able to originate BACnet/IP toward the router. Put the BACnet segment behind an allow-list at the switch or a firewall that permits only the BMS supervisor's address, and disable BACnet routing between building zones that do not need to talk. If you lease, you cannot flash the landlord's controllers - get the driver version in writing and require them to schedule it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0657","https://nvd.nist.gov/vuln/detail/CVE-2025-0658","https://www.corporate.carrier.com/product-security/advisories-resources/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-1097","cve":"CVE-2025-1097","aliases":[],"title":"ingress-nginx: Config injection via unsanitized auth-tls-match-cn annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via unsanitized auth-tls-match-cn annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-1098","cve":"CVE-2025-1098","aliases":[],"title":"ingress-nginx: Config injection via unsanitized mirror annotations","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via unsanitized mirror annotations","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-15566","cve":"CVE-2025-15566","aliases":[],"title":"ingress-nginx: Config injection via the auth-proxy-set-headers annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via the auth-proxy-set-headers annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-22086","cve":"CVE-2025-22086","aliases":["RDMA/mlx5 fix mlx5_poll_one() cur_qp update flow"],"title":"Linux kernel mlx5_ib (InfiniBand/RoCE completion queue polling): mlx5_poll_one() compares the firmware's QP number against the wrong structure's QP number, so the wrong queue…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_ib (InfiniBand/RoCE completion queue polling)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"mlx5_poll_one() compares the firmware's QP number against the wrong structure's QP number, so the wrong queue pair is used to handle a completion, leading to a NULL dereference. Anyone already on the InfiniBand subnet can drive it: unsolicited SMP/GMP/CM management datagrams land on QP0/QP1 and generate the completions, and MAD reception on an IB fabric is entirely unauthenticated. On a shared InfiniBand fabric this is a way for any attached node to crash other nodes' RDMA stacks.","attack_vector":"Any node on the same InfiniBand subnet, unauthenticated - no login on the target, just fabric attachment. Adjacent-network attack vector.","remediation":"Upgrade the host kernel to 6.15 or a stable backport (5.4.292, 5.10.236, 5.15.180, 6.1.134, 6.6.87, 6.12.23, 6.13.11, 6.14.2). Rolling reboot of every InfiniBand/RoCE host. Complementary control: enforce partition keys and restrict which nodes can attach to the subnet - this bug is a strong argument for not treating an IB fabric as a trusted flat network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22086","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-22086.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23120","cve":"CVE-2025-23120","aliases":[],"title":"Veeam Backup & Replication: Remote code execution reachable by any domain user on a domain-joined backup server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Remote code execution reachable by any domain user on a domain-joined backup server","attack_vector":"Network (remote)","remediation":"Control-plane: patch; take the backup server off the production AD domain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23120"],"status":"curated"},{"id":"CVE-2025-23121","cve":"CVE-2025-23121","aliases":[],"title":"Veeam Backup & Replication: Authenticated domain user achieves remote code execution on the Backup Server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated domain user achieves remote code execution on the Backup Server","attack_vector":"Network (remote)","remediation":"Control-plane: patch; workgroup-isolate the backup infrastructure","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23121"],"status":"curated"},{"id":"CVE-2025-23253","cve":"CVE-2025-23253","aliases":[],"title":"NVIDIA App: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA App","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23253","https://github.com/NVIDIA/product-security/tree/main/2025/5644"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/UI:N/PR:L/S:C/C:H/I:H/A:H","cwe":["CWE-547"]},{"id":"CVE-2025-23254","cve":"CVE-2025-23254","aliases":[],"title":"TensorRT-LLM: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious model / untrusted engine file","remediation":"Bump TensorRT-LLM; rebuild serving images; enforce signed model artifacts","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23254","https://github.com/NVIDIA/product-security/tree/main/2025/5648"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-24325","cve":"CVE-2025-24325","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Improper input validation in the 800-series Linux kernel driver allowing an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper input validation in the 800-series Linux kernel driver allowing an authenticated user to escalate. Highest-scored of the 2025 ice driver batch.","attack_vector":"Authenticated local user; extends to tenants on VF/RDMA-exposing nodes.","remediation":"Fixed in the Intel out-of-tree ice driver (or in-kernel equivalent). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24325","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-24514","cve":"CVE-2025-24514","aliases":[],"title":"ingress-nginx: Config injection via unsanitized auth-url annotation (part of the IngressNightmare set)","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via unsanitized auth-url annotation (part of the IngressNightmare set)","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-32431","cve":"CVE-2025-32431","aliases":[],"title":"Traefik: Path matcher flaw in PathPrefix/Path/PathRegex routing enables route and authorization bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path matcher flaw in PathPrefix/Path/PathRegex routing enables route and authorization bypass","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32431"],"status":"curated"},{"id":"CVE-2025-33186","cve":"CVE-2025-33186","aliases":[],"title":"NVIDIA AIStore - AuthN: A flaw in the AIStore authentication component reaches privilege escalation, information disclosure and data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA AIStore - AuthN","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"A flaw in the AIStore authentication component reaches privilege escalation, information disclosure and data tampering - effectively an authentication bypass in front of your training data store. Scored 8.8 network with no privileges required.","attack_vector":"Network, unauthenticated, one user-interaction step. Anyone with a route to AIStore's AuthN service.","remediation":"Upgrade AIStore per bulletin 5724 as a priority and rotate AuthN tokens afterwards - the patch does not invalidate credentials an attacker may already hold. Cost: rolling control-plane restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33186","https://github.com/NVIDIA/product-security/tree/main/2025/5724"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-798"]},{"id":"CVE-2025-33208","cve":"CVE-2025-33208","aliases":[],"title":"NVIDIA TAO Toolkit: An uncontrolled search path loads an attacker-planted resource, reaching privilege escalation and code…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA TAO Toolkit","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"An uncontrolled search path loads an attacker-planted resource, reaching privilege escalation and code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5730 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33208","https://github.com/NVIDIA/product-security/tree/main/2025/5730"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"]},{"id":"CVE-2025-33213","cve":"CVE-2025-33213","aliases":[],"title":"NVIDIA Merlin Transformers4Rec: The Trainer component deserializes untrusted data, reaching code execution. In an AI datacenter this is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Merlin Transformers4Rec","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The Trainer component deserializes untrusted data, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5739 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33213","https://github.com/NVIDIA/product-security/tree/main/2025/5739"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33214","cve":"CVE-2025-33214","aliases":[],"title":"NVIDIA NVTabular: The Workflow component deserializes untrusted data, reaching code execution. Feature-engineering workflows…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVTabular","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The Workflow component deserializes untrusted data, reaching code execution. Feature-engineering workflows are commonly shared as artifacts, which is the delivery path. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5739 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33214","https://github.com/NVIDIA/product-security/tree/main/2025/5739"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-3928","cve":"CVE-2025-3928","aliases":[],"title":"Commvault Web Server: Remote authenticated attacker creates and executes webshells","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Commvault Web Server","year":"2025","cvss_score":8.8,"severity":"high","kev":true,"impact":"[KEV] Remote authenticated attacker creates and executes webshells","attack_vector":"Network (remote)","remediation":"Control-plane: patch; sweep the web server for webshells","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3928"],"status":"curated"},{"id":"CVE-2025-39961","cve":"CVE-2025-39961","aliases":[],"title":"Linux iommu/amd - race while increasing host page table level: MULTI-TENANT ISOLATION: The AMD IOMMU host page table implementation supports growing the page table…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux iommu/amd - race while increasing host page table level","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The AMD IOMMU host page table implementation supports growing the page table dynamically, and the code that increases the level races with concurrent users. At CVSS 8.8 this is the most severe AMD IOMMU issue in the set. The AMD IOMMU is what constrains device DMA on a GPU host - it is the boundary that stops a passed-through or SR-IOV accelerator from reading memory belonging to another tenant - so a race that corrupts its page tables is a direct threat to device-level isolation, and a corruption primitive in host kernel memory besides.","attack_vector":"Local, triggered by concurrent DMA mapping activity. On a GPU node with high-rate accelerator and RDMA NIC DMA, the concurrency needed to hit this arises from normal workload behaviour, and a tenant can drive it deliberately by hammering mapping operations.","remediation":"Fixed in the Linux kernel AMD IOMMU driver. Take the distro kernel update and reboot the host - no firmware, VBIOS or AGESA step. Prioritise this on nodes doing GPU passthrough, SR-IOV or heavy RDMA: those are the configurations that exercise the dynamic page-table growth path hardest and depend most on IOMMU correctness for tenant isolation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39961"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-49847","cve":"CVE-2025-49847","aliases":[],"title":"llama.cpp (vocab): Attacker-supplied GGUF vocabulary triggers memory corruption","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (vocab)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Attacker-supplied GGUF vocabulary triggers memory corruption","attack_vector":"Customer-supplied model file","remediation":"Rebuild past b5662","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-49847"],"status":"curated"},{"id":"CVE-2025-51480","cve":"CVE-2025-51480","aliases":[],"title":"ONNX (`save_external_data`): Path traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`save_external_data`)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path traversal → overwrite arbitrary files","attack_vector":"Customer-supplied ONNX model with crafted external-data entries","remediation":"Upgrade past 1.17.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-51480"],"status":"curated"},{"id":"CVE-2025-62164","cve":"CVE-2025-62164","aliases":[],"title":"vLLM (multimodal embeddings): Memory corruption","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (multimodal embeddings)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"Memory corruption → crash and possible RCE","attack_vector":"Unauthenticated request to an exposed OpenAI-compatible serving port carrying crafted embeddings","remediation":"Upgrade to 0.11.1+. If the provider hosts the endpoint, this is provider-owned; if the tenant runs their own vLLM, it is tenant-owned and the provider can only enforce network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-62164"],"status":"curated","fleet":{"ubiquity":"Very common - vLLM is the default open-source inference server on nearly every neocloud's serverless-inference and BYO-endpoint product","remediation_pain":"`daemon-restart` - upgrade to 0.11.1+ and roll the serving fleet; every replica of every model endpoint must be restarted","pain_class":"daemon-restart","why_fleet_wide":"The Completions API `torch.load`s user-supplied prompt embeddings; with PyTorch 2.8's sparse-tensor checks off by default a crafted tensor causes an out-of-bounds write - network-only, unauthenticated-adjacent, on every serving replica"}},{"id":"CVE-2025-6685","cve":"CVE-2025-6685","aliases":["ZDI-25-650"],"title":"ATEN eco DC (DCIM/environmental management platform): The web interface doesn't check a user's assigned role before acting on their requests, so an authenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATEN eco DC (DCIM/environmental management platform)","year":"2025","cvss_score":8.8,"severity":"high","kev":false,"impact":"The web interface doesn't check a user's assigned role before acting on their requests, so an authenticated low-privileged user can escalate to actions normally reserved for administrators on the datacenter-infrastructure-management platform — which typically has visibility and control hooks into PDUs and environmental sensors across the facility.","attack_vector":"Requires a valid but low-privileged account on ATEN eco DC; the attacker sends requests for admin-level functions that the server fails to gate on role.","remediation":"Software upgrade to the patched eco DC release per ATEN's advisory. This is a server-side application (not per-rack firmware), so it's a single upgrade rather than a fleet-wide rollout, but audit who has any account on it since the bug turns any low-privileged login into an admin one.","references":["https://www.aten.com/global/en/supportcenter/info/security-advisory/25/","https://www.zerodayinitiative.com/advisories/ZDI-25-650/"],"status":"curated"},{"id":"CVE-2025-8876","cve":"CVE-2025-8876","aliases":[],"title":"N-able N-central: Improper input validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2025","cvss_score":8.8,"severity":"high","kev":true,"impact":"[KEV] Improper input validation -> OS command injection","attack_vector":"Network (remote)","remediation":"Control-plane: patch to 2025.3.1+ immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8876"],"status":"curated"},{"id":"CVE-2026-1580","cve":"CVE-2026-1580","aliases":[],"title":"ingress-nginx: Config injection via the auth-method annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via the auth-method annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-22622","cve":"CVE-2026-22622","aliases":["eaton-va-2026-1005"],"title":"Eaton Tripp Lite series PADM firmware, session management interface: A low-privilege authenticated user escalates to unrestricted device access. In practice that means a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Tripp Lite series PADM firmware, session management interface","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A low-privilege authenticated user escalates to unrestricted device access. In practice that means a read-only monitoring account - the kind operators hand to a DCIM tool or an NOC vendor - becomes full control of rack power.","attack_vector":"Any authenticated user on the PDU, including the shared read-only accounts typically configured for monitoring integrations.","remediation":"PADM firmware update or replacement for EOL SKUs. Separately, audit which third parties hold PDU accounts - monitoring integrations are the usual source of the low-privilege credential this bug needs.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/eaton-va-2026-1005.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-22807","cve":"CVE-2026-22807","aliases":[],"title":"vLLM (HF `auto_map`): Loads Hugging Face dynamic modules during model resolution","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (HF `auto_map`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Loads Hugging Face dynamic modules during model resolution → code execution","attack_vector":"Customer-supplied or poisoned Hub model repo","remediation":"Upgrade to 0.14.0+. Any `trust_remote_code`-adjacent path means the model repo is executable content","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22807"],"status":"curated"},{"id":"CVE-2026-24054","cve":"CVE-2026-24054","aliases":[],"title":"Kata Containers: Malformed or layer-less container image breaks Kata's handling","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malformed or layer-less container image breaks Kata's handling","attack_vector":"Malicious image","remediation":"Upgrade Kata to 3.26.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24054"],"status":"curated"},{"id":"CVE-2026-24164","cve":"CVE-2026-24164","aliases":[],"title":"BioNeMo Framework: Remote RCE via untrusted serialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Remote RCE via untrusted serialization","attack_vector":"Malicious model artifact / network payload","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24164","https://github.com/NVIDIA/product-security/tree/main/2026/5808"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24186","cve":"CVE-2026-24186","aliases":[],"title":"NVIDIA FLARE SDK: RCE via unsafe deserialization in message handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA FLARE SDK","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization in message handling","attack_vector":"Network peer in the federation","remediation":"Upgrade FLARE; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24186","https://github.com/NVIDIA/product-security/tree/main/2026/5819"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24187","cve":"CVE-2026-24187","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free in context handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free in context handling)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node, evict all tenant workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24187","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24217","cve":"CVE-2026-24217","aliases":[],"title":"BioNeMo Framework: Arbitrary file read/write via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Arbitrary file read/write via path traversal","attack_vector":"Malicious model archive","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24217","https://github.com/NVIDIA/product-security/tree/main/2026/5831"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-29"]},{"id":"CVE-2026-24512","cve":"CVE-2026-24512","aliases":[],"title":"ingress-nginx: Config injection via rules.http.paths.path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Config injection via rules.http.paths.path","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-24747","cve":"CVE-2026-24747","aliases":[],"title":"PyTorch (`weights_only` unpickler): Bypass of the `weights_only` allowlist","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`weights_only` unpickler)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Bypass of the `weights_only` allowlist → arbitrary code execution","attack_vector":"Customer-supplied checkpoint; defeats the mitigation shipped for CVE-2025-32434","remediation":"Rebuild images on torch >= 2.10.0. Proves the format itself, not the flag, is the control — push checkpoint-format policy (safetensors-only ingest) rather than version-chasing","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24747"],"status":"curated"},{"id":"CVE-2026-27893","cve":"CVE-2026-27893","aliases":[],"title":"vLLM (hardcoded `trust_remote_code`, second instance): Same class, two more model files, through 0.18.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (hardcoded `trust_remote_code`, second instance)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Same class, two more model files, through 0.18.0","attack_vector":"Customer-supplied model repo","remediation":"Upgrade to 0.18.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27893"],"status":"curated"},{"id":"CVE-2026-3288","cve":"CVE-2026-3288","aliases":[],"title":"ingress-nginx: rewrite-target annotation injects nginx config","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"rewrite-target annotation injects nginx config; arbitrary code execution in the controller context","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade, no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3288"],"status":"curated"},{"id":"CVE-2026-33175","cve":"CVE-2026-33175","aliases":[],"title":"JupyterHub OAuthenticator: Authenticated user bypasses the intended identity check","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterHub OAuthenticator","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated user bypasses the intended identity check","attack_vector":"Notebook user","remediation":"Upgrade to 17.4.0+; the auth plugin is the tenant boundary on a managed notebook service","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33175"],"status":"curated"},{"id":"CVE-2026-34197","cve":"CVE-2026-34197","aliases":[],"title":"Apache ActiveMQ: Improper input validation and code injection in the broker","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache ActiveMQ","year":"2026","cvss_score":8.8,"severity":"high","kev":true,"impact":"[KEV] Improper input validation and code injection in the broker","attack_vector":"Network (remote)","remediation":"Control-plane: broker upgrade; ActiveMQ is a recurring ransomware entry point","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34197"],"status":"curated"},{"id":"CVE-2026-35029","cve":"CVE-2026-35029","aliases":[],"title":"LiteLLM (`/config/update`): Endpoint does not enforce admin authorization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (`/config/update`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Endpoint does not enforce admin authorization","attack_vector":"Any authenticated proxy user","remediation":"Upgrade to 1.83.0+; a tenant key can rewrite the gateway config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35029"],"status":"curated"},{"id":"CVE-2026-35044","cve":"CVE-2026-35044","aliases":[],"title":"BentoML (Dockerfile generation): Injection into generated Dockerfile","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (Dockerfile generation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Injection into generated Dockerfile → build-time code execution","attack_vector":"Customer-supplied `bentofile.yaml` built by a provider-run build service","remediation":"Provider-owned if the provider offers managed builds: a tenant's build manifest executes in the provider's builder","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35044"],"status":"curated"},{"id":"CVE-2026-35397","cve":"CVE-2026-35397","aliases":[],"title":"Jupyter Server: Path traversal in the REST API","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Path traversal in the REST API","attack_vector":"Notebook user or anyone reaching an exposed notebook port","remediation":"Upgrade past 2.17.0. Notebook servers on GPU nodes are frequently exposed with token auth only","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35397"],"status":"curated"},{"id":"CVE-2026-3821","cve":"CVE-2026-3821","aliases":[],"title":"Supermicro SMASH service (X14DBG-DAP, X14DBI): An attacker with any authorised BMC login escalates through the SMASH shell to arbitrary code execution…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro SMASH service (X14DBG-DAP, X14DBI)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An attacker with any authorised BMC login escalates through the SMASH shell to arbitrary code execution against the BMC, or knocks the controller offline entirely. On the DBG/DBI platform boards this is a current-generation GPU node, so the payoff is out-of-band control of live accelerator hardware: power cycling to disrupt long training runs, virtual-media boot into an attacker image, and a firmware-resident implant that persists across the node being returned to the pool. The CLI management shell exposed over SSH on the BMC of Supermicro's newest GPU-platform boards.","attack_vector":"An authenticated low-privilege BMC account with SSH reachability to the controller. Read-only or operator-tier BMC accounts handed to monitoring systems, support staff or tenants are enough - the privilege bar is low, and SMASH-over-SSH is enabled by default on these boards.","remediation":"Firmware flash from Supermicro's July 2026 BMC/IPMI advisory batch. There is a genuine config-only mitigation here that most operators should apply regardless of patch state: disable the SSH/SMASH service on the BMC entirely if your tooling uses Redfish or IPMI-over-LAN, which removes this and the whole SMASH overflow family from your attack surface at zero rollout cost. Otherwise restrict SSH to the BMC to a management-host allowlist and audit every non-admin BMC account you have handed out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3821","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/3xxx/CVE-2026-3821.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-40217","cve":"CVE-2026-40217","aliases":[],"title":"LiteLLM (`/guardrails/test_custom_code`): RCE via bytecode rewriting","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (`/guardrails/test_custom_code`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"RCE via bytecode rewriting","attack_vector":"Authenticated network user of the proxy","remediation":"Upgrade; the guardrail-test endpoint is an eval sink","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-40217"],"status":"curated"},{"id":"CVE-2026-41486","cve":"CVE-2026-41486","aliases":[],"title":"Ray Data (Arrow extension types): Custom Arrow extension types deserialized unsafely","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray Data (Arrow extension types)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Custom Arrow extension types deserialized unsafely","attack_vector":"Customer-supplied Arrow/Parquet dataset","remediation":"Upgrade to 2.55.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41486"],"status":"curated"},{"id":"CVE-2026-42271","cve":"CVE-2026-42271","aliases":[],"title":"LiteLLM proxy: Two endpoints allow privilege escalation / unauthorized action","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM proxy","year":"2026","cvss_score":8.8,"severity":"high","kev":true,"impact":"Two endpoints allow privilege escalation / unauthorized action","attack_vector":"Authenticated low-privilege proxy user","remediation":"**[KEV]** Patch to 1.83.7+ immediately","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-42271"],"status":"curated"},{"id":"CVE-2026-4342","cve":"CVE-2026-4342","aliases":[],"title":"ingress-nginx: Comment-based nginx configuration injection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Comment-based nginx configuration injection","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-44346","cve":"CVE-2026-44346","aliases":[],"title":"BentoML (`bentofile.yaml`): Malicious build manifest","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (`bentofile.yaml`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Malicious build manifest → code execution in the build pipeline","attack_vector":"Customer-supplied build config","remediation":"Upgrade to 1.4.39+; isolate build workers per tenant","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44346"],"status":"curated"},{"id":"CVE-2026-45831","cve":"CVE-2026-45831","aliases":[],"title":"ChromaDB (SimpleRBAC): Authorization provider evaluates permissions incorrectly","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB (SimpleRBAC)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authorization provider evaluates permissions incorrectly","attack_vector":"Authenticated low-privilege tenant","remediation":"Upgrade past 0.5.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45831"],"status":"curated"},{"id":"CVE-2026-45832","cve":"CVE-2026-45832","aliases":[],"title":"ChromaDB (V1 endpoints): Tenant/database passed as `None` to the authz layer","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB (V1 endpoints)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Tenant/database passed as `None` to the authz layer → cross-tenant access","attack_vector":"Authenticated tenant on a shared instance","remediation":"Upgrade. Direct cross-tenant data access — the multi-tenancy control simply does not apply","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45832"],"status":"curated"},{"id":"CVE-2026-45833","cve":"CVE-2026-45833","aliases":[],"title":"ChromaDB: Authenticated code injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Authenticated code injection","attack_vector":"Any authenticated tenant of a shared Chroma instance","remediation":"Upgrade; do not multi-tenant a single Chroma","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45833"],"status":"curated"},{"id":"CVE-2026-47101","cve":"CVE-2026-47101","aliases":[],"title":"LiteLLM (key generation): internal_user can mint keys with routes their role forbids","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (key generation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"internal_user can mint keys with routes their role forbids","attack_vector":"Authenticated tenant user","remediation":"Upgrade to 1.83.14+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47101"],"status":"curated"},{"id":"CVE-2026-47102","cve":"CVE-2026-47102","aliases":[],"title":"LiteLLM (`/user/update`): User can self-elevate `user_role`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM (`/user/update`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"User can self-elevate `user_role`","attack_vector":"Authenticated tenant user","remediation":"Upgrade to 1.83.10+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47102"],"status":"curated"},{"id":"CVE-2026-4944","cve":"CVE-2026-4944","aliases":[],"title":"vLLM (hardcoded `trust_remote_code=True`): Two model implementation files force remote code execution regardless of operator setting","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (hardcoded `trust_remote_code=True`)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Two model implementation files force remote code execution regardless of operator setting","attack_vector":"Customer-supplied model repo","remediation":"Upgrade past 0.14.1. Operator-level `--trust-remote-code=false` does not protect you — no config fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-4944"],"status":"curated"},{"id":"CVE-2026-53360","cve":"CVE-2026-53360","aliases":[],"title":"Linux KVM - GHCB v2+ scratch area location enforcement: MULTI-TENANT ISOLATION: KVM did not require the GHCB software scratch area to live inside the GHCB's own…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - GHCB v2+ scratch area location enforcement","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: KVM did not require the GHCB software scratch area to live inside the GHCB's own shared buffer when GHCB v2+ is in use, as the spec demands. A guest can therefore point the scratch area at memory outside the shared region and get the host to read or write there on its behalf - a confused-deputy path from a confidential guest into host memory. Guest-to-host escape shape, and the CVSS 8.8 reflects it.","attack_vector":"From inside an SEV-ES/SNP guest via the GHCB protocol - tenant-reachable with no host privilege.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Top-of-queue for any node hosting tenant-supplied confidential VMs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53360"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53374","cve":"CVE-2026-53374","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu GEM/VM/command-submission ioctl surface. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: zero-initialize GART table on allocation","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53374","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53375","cve":"CVE-2026-53375","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vce): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vce)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vce: Prevent partial address patches","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53375","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-56340","cve":"CVE-2026-56340","aliases":[],"title":"vLLM (sparse tensor validation): Missing sparse-tensor invariant checks in multimodal embeddings","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (sparse tensor validation)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Missing sparse-tensor invariant checks in multimodal embeddings → memory corruption","attack_vector":"Unauthenticated request to the serving port","remediation":"Upgrade past 0.13.0; PyTorch disables sparse invariant checks by default, so the fix must be in vLLM","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-56340"],"status":"curated"},{"id":"CVE-2026-57516","cve":"CVE-2026-57516","aliases":[],"title":"Ray (WebDataset reader): Unsafe deserialization","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (WebDataset reader)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Unsafe deserialization → RCE from a malicious tar archive","attack_vector":"Customer-supplied dataset consumed by a Ray Data pipeline","remediation":"Upgrade to 2.56.0+. Dataset files are now an RCE vector, not just model files","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-57516"],"status":"curated"},{"id":"CVE-2026-59093","cve":"CVE-2026-59093","aliases":[],"title":"Weaviate: RBAC role assignment does not verify the assigner holds the granted permissions","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Weaviate","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"RBAC role assignment does not verify the assigner holds the granted permissions → privilege escalation","attack_vector":"Authenticated tenant of a shared Weaviate","remediation":"Upgrade to 1.38.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-59093"],"status":"curated"},{"id":"CVE-2026-64516","cve":"CVE-2026-64516","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce1): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce1)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vce1: Fix VCE 1 firmware size and offsets","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64516","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64522","cve":"CVE-2026-64522","aliases":["net/mlx5e eswitch mode block underflow on IPsec acquire SA"],"title":"Linux kernel mlx5_core IPsec offload / eswitch mode interlock: TENANT ISOLATION: the acquire-SA path unconditionally calls mlx5_eswitch_unblock_mode() without a matching…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core IPsec offload / eswitch mode interlock","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: the acquire-SA path unconditionally calls mlx5_eswitch_unblock_mode() without a matching block, underflowing the counter that is supposed to prevent eswitch mode transitions while IPsec offload is live. Once that interlock is broken, switchdev/legacy mode can be flipped out from under active offloads - and eswitch mode is what defines VF steering and isolation on the NIC. A remote packet is enough to start unwinding the enforcement point for tenant separation.","attack_vector":"Network-reachable, unauthenticated: a remote TCP SYN routed through an administrator-configured outbound IPsec policy reaches the vulnerable acquire-SA callback. No account on the host.","remediation":"Upgrade the host kernel to 7.1 or a stable backport (6.18.34, 7.0.11). Rolling reboot of every node running mlx5 IPsec full offload. Interim: if you are not depending on hardware IPsec offload, disable it on the mlx5 interfaces (config change, no reboot) to take the path out of reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64522","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-64522.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68107","cve":"CVE-2026-68107","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn4: avoid rereading IB param length","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68107","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68108","cve":"CVE-2026-68108","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vce)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vce: fix integer overflow in image size","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68108","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68432","cve":"CVE-2026-68432","aliases":[],"title":"Linux VXLAN driver (CAP_NET_ADMIN check on changelink across netns): TENANT ISOLATION: a VXLAN tunnel's `changelink()` operates across two network namespaces — the device's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux VXLAN driver (CAP_NET_ADMIN check on changelink across netns)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a VXLAN tunnel's `changelink()` operates across two network namespaces — the device's namespace and the sticky underlay namespace — but the capability check only covers the device's. Once a VXLAN device has been created in or moved to a different namespace, a caller with CAP_NET_ADMIN in only one of them can reconfigure the tunnel's underlay side. In a container platform, network namespaces are the tenant boundary and CAP_NET_ADMIN inside a namespace is something you grant routinely; this turns namespace-local privilege into control over the underlay encapsulation that other tenants share.","attack_vector":"A container or tenant holding CAP_NET_ADMIN in its own network namespace, against a VXLAN device whose underlay namespace differs from its device namespace.","remediation":"Kernel upgrade plus host reboot across container hosts. Interim: do not grant CAP_NET_ADMIN to tenant containers — a container-runtime policy change and one of the highest-value single restrictions available on a shared GPU host, since it also closes a long tail of similar netlink-reachable issues.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68432"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72045","cve":"CVE-2026-72045","aliases":[],"title":"Linux octeontx2-af (Marvell OCTEON CN10K, LMTLINE mailbox handler): TENANT ISOLATION: the OCTEON CN10K admin-function mailbox handler uses a caller-supplied `base_pcifunc` as a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux octeontx2-af (Marvell OCTEON CN10K, LMTLINE mailbox handler)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: the OCTEON CN10K admin-function mailbox handler uses a caller-supplied `base_pcifunc` as a direct index into the LMT map table, reading *another* PCI function's LMTLINE physical base address and copying it into the caller's own map-table entry. The mailbox dispatcher authenticates the requesting function, then ignores that authentication for the field that selects whose memory window you get. A VF assigned to one tenant can therefore point itself at another function's doorbell region on a shared OCTEON DPU. This is a textbook SR-IOV isolation break: the hardware isolation exists, the software hands out the key.","attack_vector":"A tenant holding an OCTEON VF — an SR-IOV virtual function passed into a VM or container — sending a crafted mailbox request to the admin function.","remediation":"Kernel upgrade plus host reboot on every node with Marvell OCTEON CN10K networking. Rolling drain across the fleet; nothing to flash. Until patched, do not assign OCTEON VFs to untrusted tenants — the isolation you are relying on is not being enforced.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72045"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72286","cve":"CVE-2026-72286","aliases":[],"title":"Linux KVM - intra-host migration/mirroring of SEV-SNP VMs: MULTI-TENANT ISOLATION: KVM allowed intra-host migration and mirroring of SEV-SNP VMs even though the feature…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - intra-host migration/mirroring of SEV-SNP VMs","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: KVM allowed intra-host migration and mirroring of SEV-SNP VMs even though the feature was never fully implemented for SNP - SNP-specific state such as the guest request mutex and message counters is not carried across. Migrating or mirroring an SNP VM through this path lands it in an inconsistent confidential state, which is a route to breaking the guest's isolation and to host-side memory corruption. At 8.8 this is the most severe SEV-related kernel issue in the current set.","attack_vector":"Reachable through the KVM ioctl surface used for intra-host migration/mirroring - so a process with access to /dev/kvm and the ability to drive migration, i.e. the VMM. In a multi-tenant control plane that is your orchestrator, which makes control-plane compromise the realistic path in.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. The upstream fix simply rejects the operation for SNP VMs. Until you are patched, disable intra-host migration and VM mirroring for SEV-SNP guests in your VMM configuration - that removes the reachable path without a reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72286"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72499","cve":"CVE-2026-72499","aliases":[],"title":"Linux bnxt_re RoCE driver (CQ toggle page use-after-free): TENANT ISOLATION: the completion-queue variant of the toggle-page use-after-free in the Broadcom RoCE driver.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (CQ toggle page use-after-free)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: the completion-queue variant of the toggle-page use-after-free in the Broadcom RoCE driver. Same reachability — the RDMA object lifecycle, which tenants exercise every time they set up or tear down a connection — and the same concern that a stray firmware interrupt writes into memory the kernel has already handed back.","attack_vector":"Local user with RDMA verbs access, racing completion-queue destruction against a notification-queue interrupt.","remediation":"Kernel/driver upgrade plus host reboot. Same rolling-drain cost as its SRQ sibling; both land in the same patch, so plan one reboot, not two.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72499"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72500","cve":"CVE-2026-72500","aliases":[],"title":"Linux bnxt_re RoCE driver (SRQ toggle page use-after-free): TENANT ISOLATION: a use-after-free in the Broadcom RoCE driver — the toggle page backing a shared receive…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (SRQ toggle page use-after-free)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a use-after-free in the Broadcom RoCE driver — the toggle page backing a shared receive queue is freed before firmware teardown completes, so a notification-queue interrupt arriving mid-destroy writes into freed memory. This is in the RDMA path, which means it is reachable from the queue-pair lifecycle that tenants drive directly when they open and close RDMA connections. A write-after-free driven by an interrupt in the RDMA control path is the shape of bug that turns into cross-tenant memory access on a shared RoCE fabric. Companion issue CVE-2026-72499 is the completion-queue equivalent.","attack_vector":"A local user with RDMA verbs access — i.e. any tenant running RoCE workloads — creating and destroying shared receive queues to race the teardown against an incoming NQ interrupt.","remediation":"Kernel/driver upgrade plus host reboot. On a RoCE cluster this is a full rolling reboot of every node using Broadcom RDMA, and jobs must be drained first. If you cannot patch quickly, restricting which containers get RDMA device access (verbs char devices) narrows who can drive the vulnerable lifecycle — a config change in your container runtime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72500","https://nvd.nist.gov/vuln/detail/CVE-2026-72499"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74527","cve":"CVE-2026-74527","aliases":[],"title":"Linux octeontx2-af (VF clobbering shared CGX PKIND state): TENANT ISOLATION: PF and VF NIX logical functions that share a CGX MAC reuse the same hardware packet-parse…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux octeontx2-af (VF clobbering shared CGX PKIND state)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: PF and VF NIX logical functions that share a CGX MAC reuse the same hardware packet-parse (PKIND) programming, and a VF allocating a NIX LF could reset the MAC's RX PKIND and default TX parse configuration set up by the PF. Parse configuration determines how the adapter interprets every frame on that MAC — so a tenant VF can change packet interpretation for everyone sharing the physical port, including the operator. The fix adds an explicit permission check that was simply absent.","attack_vector":"A tenant VF on an OCTEON adapter allocating a NIX logical function on a CGX MAC shared with the PF or with other tenants.","remediation":"Kernel upgrade plus host reboot on OCTEON nodes. Rolling drain. Structurally, avoid sharing a single CGX MAC between an operator PF and tenant VFs where the platform allows dedicating MACs instead — a provisioning-layout decision rather than a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74527"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-8828","cve":"CVE-2026-8828","aliases":[],"title":"ChromaDB (Rust): Missing authorization validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ChromaDB (Rust)","year":"2026","cvss_score":8.8,"severity":"high","kev":false,"impact":"Missing authorization validation → arbitrary read/write across collections","attack_vector":"Authenticated tenant","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-8828"],"status":"curated"},{"id":"NCVD-2021-006-infiniband-subnet-management-sub","cve":null,"aliases":["InfiniBand P_Key enforcement gap","Q_Key weakness","SM MAD trust","subnet manager spoofing","GUID spoofing","M_Key optional"],"title":"InfiniBand subnet management - Subnet Management Packets (SMPs), P_Key/Q_Key partition enforcement, port/node GUIDs: TENANT ISOLATION: InfiniBand's tenant boundary is the partition key, and its…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand subnet management - Subnet Management Packets (SMPs), P_Key/Q_Key partition enforcement, port/node GUIDs","year":"2021","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: InfiniBand's tenant boundary is the partition key, and its fabric control plane is the subnet manager speaking unauthenticated management datagrams. Three weaknesses compound. First, P_Key enforcement is a per-port switch capability that must be explicitly turned on; where it is left off - a common default on smaller fabrics - partition membership is advisory and any node can talk to any other. Second, Q_Keys guarding unreliable-datagram traffic are frequently left at well-known or predictable values, so UD traffic including management traffic is forgeable. Third, subnet management packets are protected only by the optional M_Key, and if M_Key is unset or set to zero any node on the subnet can issue SMPs - reprogramming LIDs and routing tables, or standing up a rogue subnet manager that takes over the fabric. A node with a spoofed GUID can inherit another node's partition membership outright.","attack_vector":"From any host attached to the IB subnet, the attacker sends SMPs on QP0 (unauthenticated when M_Key is unset) to read and rewrite switch forwarding tables and port configuration, or announces a higher-priority subnet manager and wins the SM election. Partition membership is then whatever the attacker says it is, and traffic can be mirrored, redirected, or blackholed. GUID spoofing is a driver/firmware-level parameter on many adapters. None of this requires exploiting a software defect - the specification permits all of it when the optional protections are not configured.","remediation":"Config change, and it is cheap relative to the exposure - do it this quarter. Set a non-zero M_Key with lease protection on every port so SMPs from unauthorised nodes are rejected; enable P_Key enforcement on all switch ports facing tenant hosts; assign a distinct, non-default Q_Key per tenant; and pin the subnet manager by priority with SM handover disabled, ideally running OpenSM or NVIDIA UFM on a management node tenants cannot reach. Applying M_Key and P_Key enforcement is pushed through the SM configuration and takes effect on the next sweep - no switch reload and no host reboot, though a botched M_Key rollout can lock you out of your own fabric, so stage it. Audit with ibnetdiscover/saquery that enforcement is actually on, since it silently defaults off.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://arxiv.org/abs/2202.08080"],"status":"curated"},{"id":"NCVD-2022-001-infiniband-roce-local-rnic-kerne","cve":null,"aliases":["NeVerMore","RDMA local packet injection","Taranov et al., arXiv:2202.08080"],"title":"InfiniBand/RoCE local RNIC - kernel bypass path shared by all local processes: TENANT ISOLATION: NeVerMore showed that an unprivileged local user can inject packets into any RDMA…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand/RoCE local RNIC - kernel bypass path shared by all local processes","year":"2022","cvss_score":8.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: NeVerMore showed that an unprivileged local user can inject packets into any RDMA connection created on the same local network controller, bypassing the operating system and kernel entirely. On a multi-tenant node this is a complete break of the host's process boundary at the fabric layer: a low-privilege container that has been given a /dev/infiniband device can forge traffic belonging to a co-resident tenant's queue pairs, and from there acquire unauthorized block access to NVMe-oF targets that trusted the RDMA connection as the authenticator. In a GPU cluster this converts one compromised training pod into read/write access to the whole tenant's remote datasets.","attack_vector":"The attacker needs only ordinary access to the local RDMA verbs device - which is how RDMA is normally exposed to containers, since kernel bypass is the whole point. They construct raw queue pairs and emit packets whose transport headers name another local process's QP, because the RNIC does not verify that the submitting context owns the source identity it stamps on the wire. The paper implements four RDMA-protocol attacks and seven NVMe-oF attacks and verifies them against both SPDK and the Linux kernel NVMe-oF implementations.","remediation":"No patch closes this in general. Config change with real cost: stop sharing one RNIC/PF across trust boundaries - give each tenant a dedicated SR-IOV VF or a dedicated physical NIC, and never mount /dev/infiniband into an untrusted container. On BlueField DPUs, terminate the RDMA connection on the DPU so the host cannot forge fabric identity - that is a hardware/topology change. Layer real authentication above the transport: enable NVMe-oF in-band DH-HMAC-CHAP so block access does not rest on the RDMA connection alone (config change on target and initiator, no reboot). Treat 'RDMA device in an untrusted container' as equivalent to root on the fabric.","references":["https://arxiv.org/abs/2202.08080","https://arxiv.org/abs/1903.09355"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-1798","cve":"CVE-2022-1798","aliases":[],"title":"KubeVirt: Path traversal lets a user who can configure KubeVirt read arbitrary host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2022","cvss_score":8.7,"severity":"high","kev":false,"impact":"Path traversal lets a user who can configure KubeVirt read arbitrary host files","attack_vector":"Cluster user with VM configuration rights","remediation":"Upgrade KubeVirt; restart virt-handler DaemonSet","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1798"],"status":"curated"},{"id":"CVE-2023-20514","cve":"CVE-2023-20514","aliases":[],"title":"AMD Secure Processor - TEE parameter handling: MULTI-TENANT ISOLATION: A privileged attacker can hand an arbitrary memory value to functions inside the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - TEE parameter handling","year":"2023","cvss_score":8.7,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A privileged attacker can hand an arbitrary memory value to functions inside the trusted execution environment, reaching arbitrary code execution in the ASP. At CVSS 8.7 this is one of the more direct host-root-to-secure-processor escalations in the set: the OS administrator, who is supposed to be outside the ASP trust boundary, gets inside it.","attack_vector":"Local, privileged (host root). No physical access and no signed-TA requirement.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20514","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-0106","cve":"CVE-2024-0106","aliases":[],"title":"BlueField / ConnectX firmware: Improper certificate validation / access control","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField / ConnectX firmware","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"Improper certificate validation / access control","attack_vector":"Network-adjacent","remediation":"Flash NIC/DPU firmware; node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0106","https://github.com/NVIDIA/product-security/tree/main/2024/5562"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-274"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0108","cve":"CVE-2024-0108","aliases":[],"title":"Jetson (Xavier/TX/Nano): Improper error handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson (Xavier/TX/Nano)","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"Improper error handling -> privesc","attack_vector":"Local attacker on the device","remediation":"Flash JetPack; edge fleet only","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0108","https://github.com/NVIDIA/product-security/tree/main/2024/5555"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:H/A:L","cwe":["CWE-755"]},{"id":"CVE-2024-23651","cve":"CVE-2024-23651","aliases":[],"title":"BuildKit: Race between parallel build steps sharing cache mounts with subpaths","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"Race between parallel build steps sharing cache mounts with subpaths; host file access","attack_vector":"Two concurrent builds on a shared builder","remediation":"Upgrade BuildKit; isolate build cache per tenant","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23651"],"status":"curated"},{"id":"CVE-2024-58310","cve":"CVE-2024-58310","aliases":[],"title":"APC Network Management Card 4 (NMC4): An unauthenticated attacker can manipulate URL parameters to walk out of the web root and read arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Network Management Card 4 (NMC4)","year":"2024","cvss_score":8.7,"severity":"high","kev":false,"impact":"An unauthenticated attacker can manipulate URL parameters to walk out of the web root and read arbitrary system files off the card, including files like /etc/passwd — enough to harvest system account information and plan a follow-on attack against the UPS/PDU's management plane.","attack_vector":"Fully remote and unauthenticated — a crafted HTTP request with encoded directory-traversal sequences is enough.","remediation":"Firmware flash of the NMC4 card to the fixed release. Roll out per card; the UPS/PDU keeps serving power to its load during the flash, but remote monitoring/management of that unit drops briefly.","references":["https://www.exploit-db.com/exploits/51897","https://www.vulncheck.com/advisories/apc-network-management-card-path-traversal-via-directory-traversal"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-20105","cve":"CVE-2025-20105","aliases":["INTEL-SA-01234","CVE-2025-20064","CVE-2025-20068","CVE-2025-20027"],"title":"UEFI firmware SMM modules in Intel reference platform firmware (SMM handler, FlashUcAcmSmm, ImcErrorHandler, WheaERST modules): Improper input validation in SMM modules that Intel ships as…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI firmware SMM modules in Intel reference platform firmware","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Improper input validation in SMM modules that Intel ships as reference code and that OEMs build into their server BIOS. SMM is the most privileged execution mode on the platform - above the hypervisor - so an escalation here gives an attacker control of the platform beneath every isolation boundary the fleet relies on, including access to the SPI flash write path. The result is a firmware implant that survives reimage and crosses tenant handoff, and that can neutralise measured boot from underneath. Because this is Intel reference code, the same defect propagates identically across every OEM that consumed that code drop, so exposure is fleet-wide across mixed vendors rather than isolated to one supplier.","attack_vector":"A privileged local user on the host - local root or an existing kernel foothold triggering the SMI. On bare-metal GPU nodes, that is the tenant.","remediation":"BIOS update from each OEM once they pick up Intel's fixed reference code - Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta and Wiwynn all ship independently, and for reference-code advisories the lag from Intel's disclosure to a shipped server BIOS routinely runs one to two quarters, longer for ODM whitebox. Requires host reboot and job drain. Track this by OEM BIOS version rather than by CVE, because OEM release notes often reference only their own advisory ID. No runtime mitigation exists for SMM defects.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20105","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01234.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-20163","cve":"CVE-2025-20163","aliases":[],"title":"Cisco Nexus Dashboard Fabric Controller (SSH host key validation): NDFC does not validate the SSH host keys of the switches it manages, so anyone positioned on the management…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus Dashboard Fabric Controller (SSH host key validation)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"NDFC does not validate the SSH host keys of the switches it manages, so anyone positioned on the management network can impersonate a managed switch and machine-in-the-middle the controller's sessions with it — harvesting the device credentials NDFC presents and feeding back forged state. Affects all NDFC versions regardless of configuration, which means every NDFC deployment ever built has been trusting its management network implicitly.","attack_vector":"Unauthenticated attacker with a position on the path between NDFC and its managed devices — a compromised management-network host, a rogue device, or ARP/route manipulation on the OOB VLAN.","remediation":"Upgrade NDFC to a release that pins host keys. Controller software upgrade. Also rotate every switch credential NDFC holds, since they may have been captured — that credential rotation across a fabric is the expensive part, not the upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20163"],"status":"curated"},{"id":"CVE-2025-23256","cve":"CVE-2025-23256","aliases":[],"title":"BlueField DPU: Access-control bypass on the DPU","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField DPU","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Access-control bypass on the DPU -> control of the tenant network path","attack_vector":"Network-adjacent attacker / tenant on the DPU-served host","remediation":"Flash DPU firmware + upgrade DOCA; DPU reset drops tenant networking, drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23256","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","fleet":{"ubiquity":"very common - ConnectX NICs are in essentially every GPU node; BlueField DPUs increasingly own the tenant network and storage path","remediation_pain":"firmware-flash of NIC/DPU firmware per node (BF-2/BF-3 images 45.1020, 35.4554 LTS22, 39.5050 LTS23, 43.3608 LTS24), often requiring a host reboot to activate","pain_class":"firmware-flash","why_fleet_wide":"The DPU is a full independent computer with DMA to the host and control of the tenant's network and storage offload - compromise is below the hypervisor and invisible to the host OS, on every node carrying the card."},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-863"]},{"id":"CVE-2025-23293","cve":"CVE-2025-23293","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): An unauthorised action against the DLS reaches high integrity and availability impact with a changed scope…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"An unauthorised action against the DLS reaches high integrity and availability impact with a changed scope - scored 8.7, the most serious of the licensing-appliance issues. An attacker on the management network can take the licensing service down or corrupt it, which eventually strands every vGPU guest.","attack_vector":"Adjacent network, low privileges, no user interaction. Any account on the management network reaches it.","remediation":"Update the DLS appliance per bulletin 5705 as a priority, and segment the licensing appliance off the general management VLAN. Cost: appliance restart; plan it against your licence lease window so guests do not notice.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23293","https://github.com/NVIDIA/product-security/tree/main/2025/5705"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:N/I:H/A:H","cwe":["CWE-306"]},{"id":"CVE-2025-54412","cve":"CVE-2025-54412","aliases":[],"title":"skops (scikit-learn model sharing): Inconsistency in the `Operator` handling lets an untrusted model bypass the safe loader","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"skops (scikit-learn model sharing)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Inconsistency in the `Operator` handling lets an untrusted model bypass the safe loader","attack_vector":"Customer-supplied skops model file","remediation":"Upgrade past 0.11.0; skops is the \"safe alternative to pickle\" and it too has loader bypasses","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54412"],"status":"curated"},{"id":"CVE-2025-54413","cve":"CVE-2025-54413","aliases":[],"title":"skops: Method-handling inconsistency","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"skops","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"Method-handling inconsistency → safe-loader bypass","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 0.11.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54413"],"status":"curated"},{"id":"CVE-2025-61958","cve":"CVE-2025-61958","aliases":["K000154647"],"title":"F5 BIG-IP (iHealth command / tmsh restricted shell): An authenticated attacker with at least a resource-administrator role can use the iHealth command to break…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"F5 BIG-IP (iHealth command / tmsh restricted shell)","year":"2025","cvss_score":8.7,"severity":"high","kev":false,"impact":"An authenticated attacker with at least a resource-administrator role can use the iHealth command to break out of the restricted tmsh shell and get a full bash shell on the device — this specifically defeats BIG-IP's Appliance mode, the hardened mode operators use to lock admins out of the underlying OS on shared/regulated deployments.","attack_vector":"Requires an authenticated account with resource-administrator role (not full root/admin) — the attack is a privilege-escalation/shell-escape from a role that was supposed to be constrained.","remediation":"Software upgrade to the fixed BIG-IP release per F5 K000154647. Part of the same October 2025 remediation batch as CVE-2025-53521 — apply in the same maintenance window. This specifically matters for shared/managed BIG-IP deployments that rely on Appliance mode to keep administrators out of shell access.","references":["https://my.f5.com/manage/s/article/K000154647"],"status":"curated"},{"id":"CVE-2026-31837","cve":"CVE-2026-31837","aliases":[],"title":"Istio: When JWKS resolution fails, istiod falls back to hardcoded defaults, weakening JWT validation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"When JWKS resolution fails, istiod falls back to hardcoded defaults, weakening JWT validation","attack_vector":"Unauthenticated network, exploitable by first disrupting JWKS reachability","remediation":"Rolling istiod upgrade to 1.29.1/1.28.5/1.27.8+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31837"],"status":"curated"},{"id":"CVE-2026-34940","cve":"CVE-2026-34940","aliases":[],"title":"KubeAI (Ollama engine controller): Injection in `ollamaStartupProbeScript()`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"KubeAI (Ollama engine controller)","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"Injection in `ollamaStartupProbeScript()`","attack_vector":"Tenant-supplied model name in a Kubernetes AI operator","remediation":"Upgrade to 0.23.2+; operator-plane compromise from tenant input","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34940"],"status":"curated"},{"id":"CVE-2026-53230","cve":"CVE-2026-53230","aliases":["net/mlx5 slab-out-of-bounds in mlx5_query_nic_vport_mac_list"],"title":"Linux kernel mlx5_core eswitch / vport (SR-IOV): TENANT ISOLATION: mlx5_core sizes a firmware command buffer from the physical function's MAC-list capability…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch / vport (SR-IOV)","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"TENANT ISOLATION: mlx5_core sizes a firmware command buffer from the physical function's MAC-list capability, but a virtual function can be configured with a larger max. Querying that VF makes the firmware response overflow the PF's buffer - a slab out-of-bounds in the host kernel driven by a value the VF side controls. CVSS scope is Changed. This is the clean VF-to-PF memory-safety break: a tenant holding an SR-IOV VF corrupts host kernel memory belonging to the NIC that serves everyone on the node.","attack_vector":"A tenant or container with local control of an mlx5 SR-IOV VF, in combination with a VF max-MAC-list setting larger than the PF's capability. Triggered from the PF-side eswitch worker, so no host root is needed on the attacker side.","remediation":"Upgrade the host kernel to 7.1 or a stable backport (6.6.143, 6.12.94, 6.18.36, 7.0.13). Rolling reboot of every SR-IOV host - drain GPU jobs per node, this is not live-patchable in practice. Interim: audit devlink VF MAC-list max settings and keep them at or below the PF capability.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53230","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-53230.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-55732","cve":"CVE-2026-55732","aliases":["CVE-2026-12496","CVE-2026-12504","CVE-2026-55731"],"title":"Loytec L-INX, L-GATE, L-ROC, L-IOB, L-DALI, L-VIS, L-PAD and LIP-ME201C (through 8.4.18, LINX-A64): An out-of-bounds read in BACnet packet parsing lets an unauthenticated attacker crash the main…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Loytec L-INX, L-GATE, L-ROC, L-IOB, L-DALI, L-VIS, L-PAD and LIP-ME201C (through 8.4.18, LINX-A64)","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"An out-of-bounds read in BACnet packet parsing lets an unauthenticated attacker crash the main control process and reboot the device with a single malformed TimeSynchronization message - and BACnet TimeSynchronization is a broadcast service, so one packet can take out every Loytec device on the segment at once. Loytec controllers and gateways are the BACnet/LonWorks integration layer in a lot of European-designed and mixed-vendor datacenters, sitting between the IP supervisory network and the field bus that runs air handling. Rebooting them all simultaneously severs supervisory control of cooling for as long as the attacker keeps sending, which against 40-140 kW GPU racks is a straightforward path to thermal shutdown across a hall. The companion issues in the same disclosure set (unauthenticated stored XSS in the OPC XML-DA statistics page, a PAM misconfiguration allowing authentication as an unintended uid, and an SNMP agent loop-condition bug) mean the same devices also offer credential-theft and persistence paths, not just DoS.","attack_vector":"Unauthenticated, remote, over BACnet on the facility network - and via broadcast, so it does not even need to know device addresses. The SNMP and web issues are likewise reachable from anywhere on that segment. Anyone who can put a frame on the building VLAN can do this.","remediation":"Firmware update from Loytec above 8.4.18 for the device family, plus LWEB-802 5.0.8+ for the management side. This is a per-device firmware flash across every gateway and controller, done by the controls integrator, with each device offline during the flash - a maintenance window on live cooling. Because a broadcast packet is the trigger, the compensating control has to actually block broadcast BACnet from untrusted hosts, which usually means putting the BACnet segment on its own VLAN with no untrusted hosts on it at all rather than trying to filter by address. Disable the SNMP agent and the OPC XML-DA interface if you are not using them.","references":["https://www.loytec.com/support/product-security/advisories","https://nvd.nist.gov/vuln/detail/CVE-2026-55732","https://nvd.nist.gov/vuln/detail/CVE-2026-12504"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-5747","cve":"CVE-2026-5747","aliases":[],"title":"Firecracker: Out-of-bounds write in the virtio PCI transport","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2026","cvss_score":8.7,"severity":"high","kev":false,"impact":"Out-of-bounds write in the virtio PCI transport; guest root can crash or potentially compromise the VMM process","attack_vector":"Any tenant guest VM with root inside it","remediation":"Upgrade Firecracker; restart microVMs, which for a neocloud means terminating tenant instances","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5747"],"status":"curated"},{"id":"CVE-2017-3883","cve":"CVE-2017-3883","aliases":[],"title":"Cisco FXOS / NX-OS AAA: AAA implementation flaw enabling remote DoS via brute-force login attempts against the switch management plane","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco FXOS / NX-OS AAA","year":"2017","cvss_score":8.6,"severity":"high","kev":false,"impact":"AAA implementation flaw enabling remote DoS via brute-force login attempts against the switch management plane","attack_vector":"Network, unauthenticated","remediation":"NX-OS/FXOS upgrade; also a reason to keep switch management auth off any tenant-reachable path","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-3883"],"status":"curated"},{"id":"CVE-2018-0378","cve":"CVE-2018-0378","aliases":[],"title":"Cisco NX-OS PTP feature (Nexus 5500/5600/6000): FABRIC DOS: an unauthenticated remote attacker takes down a Nexus switch through the PTP subsystem. PTP is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS PTP feature (Nexus 5500/5600/6000)","year":"2018","cvss_score":8.6,"severity":"high","kev":false,"impact":"FABRIC DOS: an unauthenticated remote attacker takes down a Nexus switch through the PTP subsystem. PTP is one of the few unauthenticated protocols that switches process in the control plane by design, so it is a reliable path to the CPU on every device that has it enabled. The Cisco IOS equivalent is CVE-2018-0473.","attack_vector":"Unauthenticated, remote — PTP packets reaching the switch's PTP-enabled interfaces.","remediation":"NX-OS upgrade plus reload. Immediate config mitigation: disable PTP on interfaces that face tenant workloads and keep it only on the links that actually need it — live change, no reload.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0378","https://nvd.nist.gov/vuln/detail/CVE-2018-0473"],"status":"curated"},{"id":"CVE-2018-7093","cve":"CVE-2018-7093","aliases":[],"title":"HPE iLO3/4/5: Remote unauthenticated denial of service against the management controller — loses out-of-band access to the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO3/4/5","year":"2018","cvss_score":8.6,"severity":"high","kev":false,"impact":"Remote unauthenticated denial of service against the management controller — loses out-of-band access to the node during an incident","attack_vector":"Network, unauthenticated","remediation":"iLO firmware update; DoS on the BMC is an availability problem specifically because it removes the recovery path","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7093"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-5736","cve":"CVE-2019-5736","aliases":[],"title":"runc: Host runc binary overwritten from inside a container","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2019","cvss_score":8.6,"severity":"high","kev":false,"impact":"Host runc binary overwritten from inside a container; full host root. The canonical container-escape","attack_vector":"Any tenant workload that can exec as root in its own container, or a malicious image","remediation":"Replace the runc binary on every node; already-running containers keep the vulnerable fd, so a full drain and pod restart is required","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5736"],"status":"curated","fleet":{"ubiquity":"Universal - the original runc escape; affected Docker, containerd and CRI-O simultaneously","remediation_pain":"`node-drain` - the host runc binary itself is overwritten by the exploit, so remediation is binary replacement plus recreation of every container; the standing mitigation is making runc immutable","pain_class":"node-drain","why_fleet_wide":"A container process rewrites `/proc/self/exe` and overwrites the *host* runc binary, so every subsequent container start on that host executes attacker code as root - the canonical fleet-wide container-runtime emergency"}},{"id":"CVE-2021-1587","cve":"CVE-2021-1587","aliases":[],"title":"Cisco NX-OS (VXLAN OAM / NGOAM): FABRIC DOS: a crafted VXLAN OAM packet reloads a VTEP. In a VXLAN/EVPN GPU fabric every leaf is a VTEP, so an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (VXLAN OAM / NGOAM)","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"FABRIC DOS: a crafted VXLAN OAM packet reloads a VTEP. In a VXLAN/EVPN GPU fabric every leaf is a VTEP, so an attacker with a foothold in any tenant overlay can knock out leaves one at a time and stall collectives cluster-wide. Reachable from inside a tenant's own overlay, which is what makes it interesting — it does not need underlay access.","attack_vector":"Unauthenticated, remote — the attacker needs to be able to land a crafted VXLAN packet on the switch's VTEP address. In practice that means a compromised workload or a tenant that can source arbitrary UDP.","remediation":"NX-OS upgrade plus reload. If NGOAM is not in use, disabling the feature is a live config change with no reload and removes the exposure entirely — do that first, patch on the next window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1587"],"status":"curated"},{"id":"CVE-2021-32777","cve":"CVE-2021-32777","aliases":[],"title":"Envoy: ext-authz header handling flaw allows bypassing the external authorization service","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"ext-authz header handling flaw allows bypassing the external authorization service","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32777"],"status":"curated"},{"id":"CVE-2021-32779","cve":"CVE-2021-32779","aliases":[],"title":"Envoy: URI fragment treated as part of the path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"URI fragment treated as part of the path; authorization bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32779"],"status":"curated"},{"id":"CVE-2021-32781","cve":"CVE-2021-32781","aliases":[],"title":"Envoy: Processing continues after a local reply, causing undefined behaviour","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"Processing continues after a local reply, causing undefined behaviour","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32781"],"status":"curated"},{"id":"CVE-2021-33141","cve":"CVE-2021-33141","aliases":[],"title":"Intel Ethernet Adapter manageability firmware (NC-SI / sideband path): TENANT ISOLATION: improper input validation in the *manageability* firmware of Intel Ethernet adapters…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Adapter manageability firmware (NC-SI / sideband path)","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"TENANT ISOLATION: improper input validation in the *manageability* firmware of Intel Ethernet adapters, exploitable by an unauthenticated user. Manageability firmware is the NC-SI sideband engine — the path that carries BMC traffic over the same physical NIC as production data. A flaw there is the bridge between the data network and out-of-band management: an attacker on the fabric reaches the sideband channel, and the sideband channel reaches the BMC, which controls power and virtual media for the node. This is the highest-value shape of NIC firmware bug for a multi-tenant operator, because it crosses the boundary between 'tenant network' and 'operator management plane'.","attack_vector":"Unauthenticated attacker with network access to the adapter. No host account required.","remediation":"Flash adapter manageability firmware via the OEM firmware bundle (this is separate from the main NVM image on some platforms); cold power cycle. Architecturally, the durable control is to stop sharing the production NIC with BMC traffic — use a dedicated BMC NIC on a physically separate OOB network rather than NC-SI sideband. That is a hardware/topology decision, so it applies to your next buildout, not this one. Companion issues in the same advisory family: CVE-2021-33162, CVE-2021-33161, CVE-2021-33158, CVE-2021-33157, CVE-2022-37341.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33141","https://nvd.nist.gov/vuln/detail/CVE-2021-33162"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47136","cve":"CVE-2021-47136","aliases":[],"title":"Linux kernel mlx5_core representor TC path + net/sched tc extension: The TC_SKB_EXT skb extension is not zeroed on allocation and mlx5's representor restore path never…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core representor TC path + net/sched tc extension","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"The TC_SKB_EXT skb extension is not zeroed on allocation and mlx5's representor restore path never initialized the newer fields, so Open vSwitch reads uninitialized kernel memory as a boolean. This leaks host kernel memory contents into the OVS datapath - and it is triggered remotely, by sending packets that miss hardware offload. On a switchdev host running OVS over ConnectX, that is remote kernel-memory disclosure into the software datapath.","attack_vector":"Remote, unauthenticated: send traffic crafted to miss the hardware offload path on an mlx5 switchdev host running OVS.","remediation":"Upgrade the host kernel to 5.13 or a stable backport (5.10.42, 5.12.9). Rolling reboot. Note NVD scores this 5.5 while the kernel CNA scores it 8.6 - if you triage from NVD feeds you will under-prioritize it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47136","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2021/CVE-2021-47136.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-2601","cve":"CVE-2022-2601","aliases":[],"title":"GRUB2 (font engine, grub_font_construct_glyph): Buffer overflow when constructing a glyph from a crafted GRUB font file. Fonts are unsigned data sitting in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (font engine, grub_font_construct_glyph)","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"Buffer overflow when constructing a glyph from a crafted GRUB font file. Fonts are unsigned data sitting in the boot partition on virtually every install, which makes this one of the cheapest Secure Boot bypasses in the family.","attack_vector":"Anyone who can write a font file to the boot partition - local root, prior tenant, or a poisoned image build.","remediation":"grub2 package update + reboot. Fonts are rarely needed on a headless server image; dropping the graphical GRUB theme removes this surface outright.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2601","https://access.redhat.com/security/cve/CVE-2022-2601"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-36392","cve":"CVE-2022-36392","aliases":[],"title":"Intel AMT / Standard Manageability firmware: Improper input validation in AMT/ISM firmware, scored high because of network reach. Another entry in the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT / Standard Manageability firmware","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"Improper input validation in AMT/ISM firmware, scored high because of network reach. Another entry in the recurring pattern: the out-of-band management engine keeps producing network-reachable memory-safety bugs, and each one needs an OEM firmware cycle to fix.","attack_vector":"Network access to the AMT interface on affected firmware versions.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36392","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00783.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-3775","cve":"CVE-2022-3775","aliases":[],"title":"GRUB2 (font engine, blit_comb): Integer underflow when rendering certain unicode sequences writes out of bounds. Same class as the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (font engine, blit_comb)","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"Integer underflow when rendering certain unicode sequences writes out of bounds. Same class as the glyph-construction bug and shipped in the same advisory wave - if you patched one you probably need both.","attack_vector":"Crafted font or text rendered by GRUB, reachable by anyone who can write boot-partition content.","remediation":"grub2 package update + reboot. Verify your distro's package covers both font CVEs, not just the first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3775","https://access.redhat.com/security/cve/CVE-2022-3775"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-35941","cve":"CVE-2023-35941","aliases":[],"title":"Envoy: Malicious client constructs permanently valid credentials in the OAuth filter","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.6,"severity":"high","kev":false,"impact":"Malicious client constructs permanently valid credentials in the OAuth filter","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; rotate OAuth HMAC secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-35941"],"status":"curated"},{"id":"CVE-2024-20267","cve":"CVE-2024-20267","aliases":[],"title":"Cisco NX-OS (MPLS traffic handling / netstack): FABRIC DOS: crafted MPLS traffic restarts netstack, which stops the switch forwarding or reloads it. Applies…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (MPLS traffic handling / netstack)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"FABRIC DOS: crafted MPLS traffic restarts netstack, which stops the switch forwarding or reloads it. Applies even where you are not consciously running MPLS, because the parse path is reachable regardless. One packet stream, one leaf down.","attack_vector":"Unauthenticated, remote — an attacker able to send MPLS-labelled frames toward the device.","remediation":"NX-OS upgrade plus reload. Filter MPLS ethertype at the fabric edge as a live mitigation if you do not use MPLS, which most GPU-cluster fabrics do not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20267"],"status":"curated"},{"id":"CVE-2024-20321","cve":"CVE-2024-20321","aliases":[],"title":"Cisco NX-OS (eBGP implementation): FABRIC DOS: an unauthenticated remote attacker can wedge the switch through the eBGP implementation. In a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (eBGP implementation)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"FABRIC DOS: an unauthenticated remote attacker can wedge the switch through the eBGP implementation. In a BGP-underlay leaf/spine — the standard GPU-cluster design — the routing daemon going down means the rack loses reachability, and a coordinated attack against several leaves partitions the fabric mid-training-run.","attack_vector":"Unauthenticated, remote to the BGP process. Requires the ability to reach the switch's BGP listener.","remediation":"NX-OS upgrade plus reload. Harden with strict neighbor ACLs and CoPP in the meantime — live config. Because this affects the underlay control plane, schedule the reload per MLAG/ECMP pair so the fabric never loses both paths.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20321"],"status":"curated"},{"id":"CVE-2024-20446","cve":"CVE-2024-20446","aliases":[],"title":"Cisco NX-OS (DHCPv6 relay agent): FABRIC DOS: a crafted DHCPv6 packet takes the switch out. Relevant because DHCP relay is normally configured…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (DHCPv6 relay agent)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"FABRIC DOS: a crafted DHCPv6 packet takes the switch out. Relevant because DHCP relay is normally configured on exactly the SVIs that face tenant workloads, so this is reachable from inside a tenant network with no credentials at all.","attack_vector":"Unauthenticated, remote — a host on a VLAN where DHCPv6 relay is configured.","remediation":"NX-OS upgrade plus reload. If IPv6 is unused in your cluster (still common), disabling the DHCPv6 relay is a live config change that removes the exposure with no downtime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20446"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-21626","cve":"CVE-2024-21626","aliases":[],"title":"runc: \"Leaky Vessels\": internal file descriptor leak lets a container process start with cwd in the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"\"Leaky Vessels\": internal file descriptor leak lets a container process start with cwd in the host filesystem; full host escape","attack_vector":"Malicious image (a crafted WORKDIR is enough) or any tenant workload","remediation":"Replace runc on all nodes; running containers stay vulnerable, so drain and recreate all GPU pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21626"],"status":"curated","fleet":{"ubiquity":"Universal - runc is the default OCI runtime under Docker, containerd, CRI-O and therefore under essentially every GPU container on every neocloud","remediation_pain":"`node-drain` - the runc binary is replaced on the host, and while the swap itself is atomic, already-running containers keep the vulnerable runtime semantics; safe remediation means evacuating and recreating every container, i.e. draining paying GPU jobs off each node","pain_class":"node-drain","why_fleet_wide":"A leaked file descriptor lets any customer-supplied container image (or `docker exec`) land its working directory in the host filesystem namespace, giving host root from an untrusted tenant workload on every node in the fleet"}},{"id":"CVE-2024-23324","cve":"CVE-2024-23324","aliases":["GHSA-gq3v-vvhj-96j6"],"title":"Envoy proxy (ext_authz filter): When Envoy's ext_authz filter is configured with failure_mode_allow set to true, a downstream client can…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy proxy (ext_authz filter)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"When Envoy's ext_authz filter is configured with failure_mode_allow set to true, a downstream client can force an invalid gRPC request that circumvents the external-authorization check entirely — meaning traffic that should have been checked against an auth service (a common pattern for gating access to inference or cluster-management endpoints) sails through unauthenticated.","attack_vector":"Remote — a downstream client crafts a malformed gRPC request against an Envoy instance where ext_authz is set to fail open.","remediation":"Software upgrade to Envoy 1.29.1, 1.28.1, 1.27.3, 1.26.7, or later. No config workaround exists (the vendor advisory states none), so this is a binary/image upgrade and restart across every Envoy instance doing ext_authz-based access control in the cluster ingress path — a rolling restart avoids a hard outage if you run multiple replicas.","references":["https://github.com/envoyproxy/envoy/security/advisories/GHSA-gq3v-vvhj-96j6"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-4325","cve":"CVE-2024-4325","aliases":[],"title":"Gradio (`/queue/join`): SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio (`/queue/join`)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"SSRF","attack_vector":"Unauthenticated network to the demo","remediation":"Upgrade past 4.21.0; blocks IMDS access from the GPU node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4325"],"status":"curated"},{"id":"CVE-2024-48248","cve":"CVE-2024-48248","aliases":[],"title":"NAKIVO Backup & Replication: Unauthenticated absolute path traversal via getImageByPath","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NAKIVO Backup & Replication","year":"2024","cvss_score":8.6,"severity":"high","kev":true,"impact":"[KEV] Unauthenticated absolute path traversal via getImageByPath -> arbitrary file read incl. cleartext credentials","attack_vector":"Network (remote)","remediation":"Control-plane: patch + rotate all credentials in the backup product database","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48248"],"status":"curated"},{"id":"CVE-2024-48882","cve":"CVE-2024-48882","aliases":["CVE-2024-49572","CVE-2025-20085","CVE-2025-23417","CVE-2025-26858","CVE-2025-55221","CVE-2025-54848","TALOS-2024-2119"],"title":"Socomec DIRIS Digiware M-70 1.6.9 (Modbus TCP and Modbus RTU-over-TCP): A large cluster of unauthenticated Modbus denial-of-service and buffer-overflow issues in a device that sits…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Socomec DIRIS Digiware M-70 1.6.9 (Modbus TCP and Modbus RTU-over-TCP)","year":"2024","cvss_score":8.6,"severity":"high","kev":false,"impact":"A large cluster of unauthenticated Modbus denial-of-service and buffer-overflow issues in a device that sits on the electrical monitoring backbone - and two of them do something worse than crash: the DoS also weakens credentials such that the documented default credentials become valid on the device again. That converts a crash into an authentication bypass, which is a genuinely nasty combination on gear that monitors and in some deployments controls electrical distribution. Losing the Digiware monitoring layer blinds the operator to power conditions across the hall during whatever else the attacker is doing, and a device that has silently reverted to default credentials is a persistent foothold in the electrical segment. The Modbus angle is the general lesson here: these are unauthenticated packets to TCP 502, and the same class of embedded Modbus stack fragility exists across chiller, CDU and ATS controllers throughout the facility.","attack_vector":"Unauthenticated network packets to the device's Modbus TCP service - Talos confirms a single crafted packet suffices for several of these. No credentials, no interaction. The device lives on the facility/electrical VLAN, typically polled by the BMS or a DCIM collector, and is often reachable from any host on that segment because Modbus deployments almost never carry ACLs.","remediation":"Firmware update from Socomec for the DIRIS Digiware M-70. That is a per-device flash on live electrical monitoring gear, coordinated with the electrical contractor - not a cooling outage, but still a scheduled window and a technician per device. Because credential reversion is in scope, after patching you must re-verify that default credentials no longer work on every unit, not just assume the update handled it. The durable control is the Modbus one: default-deny on TCP 502, permit only the poller, and alert on any Modbus write function code appearing on a segment that should only see reads.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2024-2119","https://talosintelligence.com/vulnerability_reports/TALOS-2024-2118","https://nvd.nist.gov/vuln/detail/CVE-2024-48882","https://nvd.nist.gov/vuln/detail/CVE-2024-49572"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-22108","cve":"CVE-2025-22108","aliases":[],"title":"Linux bnxt_en driver (TX BD bd_cnt field masking): The 5-bit bd_cnt field in the transmit buffer descriptor is not masked, so an out-of-range value corrupts…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (TX BD bd_cnt field masking)","year":"2025","cvss_score":8.6,"severity":"high","kev":false,"impact":"The 5-bit bd_cnt field in the transmit buffer descriptor is not masked, so an out-of-range value corrupts transmit descriptors and produces TX timeouts. Practically: the NIC stops transmitting. On a training node that is a hung collective, and the failure looks like a network problem rather than a driver bug, so it burns debugging time.","attack_vector":"Triggered by transmit paths that produce a descriptor count above the representable range. Local, driven by workload traffic patterns.","remediation":"Kernel/driver upgrade plus host reboot. No firmware change required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22108"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-32008","cve":"CVE-2025-32008","aliases":["INTEL-SA-01315","CVE-2025-20080","CVE-2026-20715","INTEL-SA-01427"],"title":"Intel AMT and Intel Standard Manageability firmware (current CSME generations): Out-of-bounds write in AMT/ISM firmware reachable by an unauthenticated network adversary, with a companion…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT and Intel Standard Manageability firmware (current CSME generations)","year":"2025","cvss_score":8.6,"severity":"high","kev":false,"impact":"Out-of-bounds write in AMT/ISM firmware reachable by an unauthenticated network adversary, with a companion null-pointer dereference in the same advisory and a further unauthenticated network denial-of-service in the August 2026 batch. The immediate effect is that anyone on the management path can knock the manageability engine of a node over; memory corruption in the ME is also the standard precursor to code execution in it. The operator-facing point is that AMT is still, in 2026, an unauthenticated network attack surface on the management VLAN - and losing the ME on a node loses your out-of-band recovery path exactly when you need it.","attack_vector":"Network adversary, unauthenticated, reaching the AMT/ISM listener on the node. Same exposure as the 2017-era AMT bugs: management VLAN, or tenant space if the manageability path is not fully isolated.","remediation":"CSME firmware flash from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta) with a host reboot and job drain. The lasting control is the same one operators keep skipping: unprovision AMT on every SKU where you do not actively use it, disable it in the BIOS profile, and block 16992/16993/623/664/5900 anywhere a tenant-reachable segment could touch it. If you do use AMT, put it behind TLS with mutual auth and a dedicated segment.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32008","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01315.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01427.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-59887","cve":"CVE-2025-59887","aliases":["ETN-VA-2025-1026"],"title":"Eaton UPS Companion (EUC) software installer: The installer does not properly authenticate the library files it loads, so an attacker who can place a file…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS Companion (EUC) software installer","year":"2025","cvss_score":8.6,"severity":"high","kev":false,"impact":"The installer does not properly authenticate the library files it loads, so an attacker who can place a file alongside the installation package gets arbitrary code execution during install - at whatever privilege the installer runs with, which is administrative. The exposure window is your own deployment process.","attack_vector":"An attacker with write access to wherever the installation package is staged - a shared drive, a downloads folder, an imaging share.","remediation":"Use the fixed EUC version from Eaton's download centre and stage installers somewhere with restricted write access. Verify package hashes before running. Companion issues CVE-2025-59888 (unquoted search path) and CVE-2025-67450 (insecure library loading) have the same fix and the same mitigation.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/etn-va-2025-1026.pdf"],"status":"curated"},{"id":"CVE-2026-1603","cve":"CVE-2026-1603","aliases":[],"title":"Ivanti Endpoint Manager (EPM): Auth bypass via alternate path","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager (EPM)","year":"2026","cvss_score":8.6,"severity":"high","kev":true,"impact":"[KEV] Auth bypass via alternate path -> unauthenticated leak of stored credential data","attack_vector":"Network (remote)","remediation":"Control-plane: patch to 2024 SU5+; rotate every credential EPM stored","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1603"],"status":"curated"},{"id":"CVE-2026-22620","cve":"CVE-2026-22620","aliases":["eaton-va-2026-1005"],"title":"Eaton Tripp Lite series PADM firmware (rack PDU / ATS management): PHYSICAL. Unauthenticated authentication bypass gives privileged access to the PDU management firmware. From…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Tripp Lite series PADM firmware (rack PDU / ATS management)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"PHYSICAL. Unauthenticated authentication bypass gives privileged access to the PDU management firmware. From a privileged session on a switched PDU an attacker controls outlet state for the rack - power-cycling nodes, or holding outlets off. Note the vendor has published an end-of-life notice for this product line alongside the advisory, which means for some deployed units the fix is replacement, not a patch.","attack_vector":"Unauthenticated, remote, against the PDU's management interface on the facility or OOB network.","remediation":"Update PADM firmware where a fixed build exists. For SKUs covered by the EOL notice there is no forward-fix path and the remediation is hardware replacement - a capex line and a rack-by-rack electrical swap, not a maintenance window. Until then, isolate the PDU management network and disable remote outlet switching.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/eaton-va-2026-1005.pdf"],"status":"curated"},{"id":"CVE-2026-24222","cve":"CVE-2026-24222","aliases":[],"title":"NemoClaw: Sensitive info exposure in logs","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NemoClaw","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Sensitive info exposure in logs","attack_vector":"Anyone with log access (incl. shared logging backends)","remediation":"Upgrade the service; purge/rotate leaked secrets from log stores","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24222","https://github.com/NVIDIA/product-security/tree/main/2026/5837"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-497"]},{"id":"CVE-2026-28500","cve":"CVE-2026-28500","aliases":[],"title":"ONNX: Security-control bypass through 1.20.1","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Security-control bypass through 1.20.1","attack_vector":"Customer-supplied ONNX model","remediation":"Upgrade; ONNX external-data handling has no durable fix — sandbox model parsing","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-28500"],"status":"curated"},{"id":"CVE-2026-34445","cve":"CVE-2026-34445","aliases":[],"title":"ONNX (`ExternalDataInfo`): Security control bypass in external-data path handling","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX (`ExternalDataInfo`)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Security control bypass in external-data path handling","attack_vector":"Customer-supplied ONNX model","remediation":"Upgrade to 1.21.0+; fourth iteration of the same external-data traversal class","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34445"],"status":"curated"},{"id":"CVE-2026-59707","cve":"CVE-2026-59707","aliases":[],"title":"LocalAI (`/models/apply`): Unauthenticated SSRF fetching arbitrary internal URLs","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LocalAI (`/models/apply`)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"Unauthenticated SSRF fetching arbitrary internal URLs","attack_vector":"Unauthenticated network to the model-install endpoint","remediation":"Upgrade; block egress to internal ranges and IMDS","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-59707"],"status":"curated"},{"id":"CVE-2026-63086","cve":"CVE-2026-63086","aliases":[],"title":"Text Generation Inference (TGI): SSRF in the OpenAI-compatible multimodal chat endpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Text Generation Inference (TGI)","year":"2026","cvss_score":8.6,"severity":"high","kev":false,"impact":"SSRF in the OpenAI-compatible multimodal chat endpoint","attack_vector":"Unauthenticated request supplying an image URL","remediation":"Upgrade past 3.3.7 and block metadata/internal egress from serving pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63086"],"status":"curated"},{"id":"NCVD-2021-003-infiniband-rocev2-transport-rnic","cve":null,"aliases":["ReDMArk","ReDMArk QP/PSN predictability","RDMA packet injection by impersonation","Rothenberger et al., USENIX Security 2021"],"title":"InfiniBand / RoCEv2 transport - RNIC connection state (QP number, PSN) on Mellanox ConnectX-class and compatible RNICs: TENANT ISOLATION: RDMA has no cryptographic binding between a packet and the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand / RoCEv2 transport - RNIC connection state (QP number, PSN) on Mellanox ConnectX-class and compatible RNICs","year":"2021","cvss_score":8.6,"severity":"high","kev":false,"impact":"TENANT ISOLATION: RDMA has no cryptographic binding between a packet and the connection it claims to belong to. A Reliable Connected queue pair is identified only by destination QP number plus packet sequence number, and ReDMArk showed both are highly predictable on real RNICs - QP numbers are handed out near-sequentially by the firmware and initial PSNs are drawn from a weak generator. Anyone who can put a frame on the fabric with the victim's source GID/LID can inject a valid-looking RDMA WRITE or SEND into an established connection between two other tenants. On a shared GPU cluster that means a neighbouring tenant, or a compromised node anywhere in the same partition, can write into another tenant's registered memory - which in an AI cluster is model weights, KV cache, gradient buffers, or NCCL communication buffers - with no software on the victim host ever seeing the write.","attack_vector":"Attacker needs one host with an RNIC on the same L2/L3 RoCE domain or IB subnet as the victims (a rented bare-metal node, a container with a VF or an SR-IOV VF, or a compromised storage/management node). They enumerate the QPN space by opening their own connections to the victim host to learn the allocator's current position, then brute-force or predict PSN and emit crafted BTH/RETH headers with a spoofed source. On RoCEv2 the outer frame is ordinary UDP/4791, so spoofing is as easy as a raw socket if the switch does not enforce source MAC/IP filtering. No exploit of a software bug is required - this is the protocol working as designed.","remediation":"No patch exists; this is architectural. Config change first: put every tenant in its own InfiniBand partition (P_Key) or its own RoCE VLAN/VRF, and turn on switch-side source-address enforcement (port security / IP source guard / MAC learning locks) so a node cannot emit frames claiming another node's GID - that is a switch config push, no reload needed on most platforms. Where the NIC supports it, enable RoCEv2 link-layer encryption/authentication (IPsec or PSP offload on ConnectX-6 Dx and later, or NVIDIA BlueField DPU-terminated crypto) - this is a firmware flash plus driver upgrade on the NIC fleet and costs a rolling host reboot per node. Hard-partitioning tenants onto separate physical fabrics or separate IB subnets is the only complete answer and is a capacity/cost decision, not a patch.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://www.usenix.org/conference/usenixsecurity22/presentation/xing"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2022-002-nvme-over-fabrics-protocol-over","cve":null,"aliases":["NeVerMore NVMe-oF attacks","NQN spoofing","NVMe-oF controller hijack over RDMA"],"title":"NVMe-over-Fabrics protocol over RDMA - SPDK NVMe-oF target and Linux kernel nvmet: TENANT ISOLATION: NeVerMore implemented seven attacks against the NVMe-oF protocol itself and confirmed them…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe-over-Fabrics protocol over RDMA - SPDK NVMe-oF target and Linux kernel nvmet","year":"2022","cvss_score":8.6,"severity":"high","kev":false,"impact":"TENANT ISOLATION: NeVerMore implemented seven attacks against the NVMe-oF protocol itself and confirmed them on the two implementations that matter operationally - SPDK and Linux nvmet. The core finding is that NVMe-oF's access control leans on the RDMA connection and on the host NQN string, neither of which is authenticated by default. A host NQN is just a text identifier the initiator asserts; there is nothing stopping a tenant from claiming another tenant's NQN and being handed their namespaces. For a neocloud selling disaggregated NVMe to GPU tenants, this means dataset and checkpoint volumes belonging to one customer can be attached read-write by another.","attack_vector":"The attacker connects to the target's discovery and I/O controllers over RDMA (or TCP) and presents a forged Host NQN, or hijacks an existing connection using the RDMA injection primitives above. Because NVMe-oF allow-lists are typically written as 'NQN X may see subsystem Y', spoofing the NQN is sufficient. Discovery controllers make reconnaissance trivial by listing every subsystem NQN and transport address on the fabric to any peer that asks.","remediation":"Config change, and it is the single highest-value one in this slice: enable NVMe-oF in-band authentication (DH-HMAC-CHAP, supported in Linux nvmet since 6.0 and in SPDK) so the NQN is proven rather than asserted, and enable TLS (NVMe/TCP) or fabric-level crypto where available. No reboot; nvmet accepts this via configfs at runtime, though initiators must be reconfigured in step so plan a rolling attach/detach. Additionally restrict the discovery controller to an authenticated management network rather than the tenant fabric, and put storage traffic on its own VLAN/P_Key. For SPDK, upgrade to a current release (process restart, brief I/O pause).","references":["https://arxiv.org/abs/2202.08080","https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64320.json"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2020-11013","cve":"CVE-2020-11013","aliases":[],"title":"Helm: The `lookup` template function discloses in-cluster resources, including Secrets, to a chart author","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2020","cvss_score":8.5,"severity":"high","kev":false,"impact":"The `lookup` template function discloses in-cluster resources, including Secrets, to a chart author","attack_vector":"Anyone who can get an operator to install their chart","remediation":"Upgrade Helm; review third-party charts before install","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11013"],"status":"curated"},{"id":"CVE-2021-30465","cve":"CVE-2021-30465","aliases":[],"title":"runc: Container filesystem breakout via directory traversal in mount handling","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2021","cvss_score":8.5,"severity":"high","kev":false,"impact":"Container filesystem breakout via directory traversal in mount handling; host root","attack_vector":"Any tenant workload able to create multiple pods with specific mount config","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-30465"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-28181","cve":"CVE-2022-28181","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): MULTI-TENANT ISOLATION: A specially crafted shader causes an out-of-bounds write in the kernel mode layer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":8.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A specially crafted shader causes an out-of-bounds write in the kernel mode layer, reaching code execution with a changed scope - and NVIDIA scores this AV:Network, meaning a remotely delivered shader (WebGL, a remote render session, a shared compute service) can reach it. This is the most dangerous display-driver bug in the 2022 set for anyone running remote rendering or browser-facing GPU workloads. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28181","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-28182","cve":"CVE-2022-28182","aliases":[],"title":"NVIDIA GPU Display Driver - Windows DirectX 11 user mode driver (nvwgf2um.dll / nvwgf2umx.dll): MULTI-TENANT ISOLATION: A crafted shader causes an out-of-bounds write in the DX11 user mode driver…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows DirectX 11 user mode driver (nvwgf2um.dll / nvwgf2umx.dll)","year":"2022","cvss_score":8.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A crafted shader causes an out-of-bounds write in the DX11 user mode driver, reaching code execution with a changed scope. NVIDIA scores it AV:Network with no privileges required, which is the signature of a shader delivered over a remote rendering or browser path rather than by a local user. For a cloud-gaming, VDI or remote-workstation operator this is the bug in the 2022 set that actually crosses a customer boundary.","attack_vector":"Network, no privileges. The attacker supplies a shader that the victim's GPU compiles and runs - via a web page, a streamed application, or a shared render pipeline. On a multi-session Windows host this reaches other users' sessions.","remediation":"Install the fixed Windows driver from bulletin 5353. Cost: node reboot after driver replacement, so drain sessions. Until patched, the only real compensating control is not accepting untrusted shader input, which is not an option for a cloud-gaming or remote-workstation product.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28182","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-34671","cve":"CVE-2022-34671","aliases":[],"title":"GPU Display Driver (kernel): Local privesc to host root (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (kernel)","year":"2022","cvss_score":8.5,"severity":"high","kev":false,"impact":"Local privesc to host root (kernel buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34671","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-22736","cve":"CVE-2023-22736","aliases":[],"title":"Argo CD: Authorization bypass lets an Application be synced to a destination it is not permitted to reach","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":8.5,"severity":"high","kev":false,"impact":"Authorization bypass lets an Application be synced to a destination it is not permitted to reach","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22736"],"status":"curated"},{"id":"CVE-2025-14459","cve":"CVE-2025-14459","aliases":[],"title":"KubeVirt CDI: PVCs can be cloned from unauthorized namespaces via DataImportCron","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt CDI","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"PVCs can be cloned from unauthorized namespaces via DataImportCron; cross-tenant data theft","attack_vector":"Cluster user with namespace access","remediation":"Upgrade CDI","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14459"],"status":"curated"},{"id":"CVE-2025-23267","cve":"CVE-2025-23267","aliases":[],"title":"Container Toolkit: Container escape / host file write via symlink following","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"Container escape / host file write via symlink following","attack_vector":"Any tenant with a container","remediation":"Bump nvidia-container-toolkit + restart runtime; upgrade GPU Operator; evict tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23267","https://github.com/NVIDIA/product-security/tree/main/2025/5659"],"status":"curated","fleet":{"ubiquity":"Universal - default hook path in every install","remediation_pain":"`daemon-restart` to 1.17.8","pain_class":"daemon-restart","why_fleet_wide":"Link-following (symlink) in the privileged `update-ldcache` hook lets a crafted image tamper with host files or DoS the node, taking down every tenant scheduled on it"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:N/I:L/A:H","cwe":["CWE-59"]},{"id":"CVE-2025-53547","cve":"CVE-2025-53547","aliases":[],"title":"Helm: Crafted Chart.yaml plus a symlinked Chart.lock gives local code execution when dependencies are updated","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"Crafted Chart.yaml plus a symlinked Chart.lock gives local code execution when dependencies are updated; compromises the CD runner","attack_vector":"Malicious chart pulled by a GitOps pipeline","remediation":"Upgrade Helm to 3.18.4+ everywhere charts are rendered, including Argo CD and Flux images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53547"],"status":"curated"},{"id":"CVE-2025-61972","cve":"CVE-2025-61972","aliases":[],"title":"AMD NBIO register lock bits - System Management Network access: MULTI-TENANT ISOLATION: NBIO registers that should be locked after boot are not, so a local admin-privileged…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD NBIO register lock bits - System Management Network access","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: NBIO registers that should be locked after boot are not, so a local admin-privileged attacker gets arbitrary access to the System Management Network - the internal bus that reaches the ASP and the SMU. From SMN access the attacker executes code in the AMD Secure Processor itself, which collapses both confidentiality and integrity for every SEV-SNP guest on the machine. At 8.5 this is the most severe ASP-reachable issue in the current batch.","attack_vector":"Local, host administrator. Exactly the threat model SEV-SNP claims to defend against, which is what makes it serious rather than routine.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Treat any confidential-computing SLA you offer as void on unpatched nodes: the host operator - or anyone who compromises the host operator's tooling - can reach guest memory.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-61972","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-64324","cve":"CVE-2025-64324","aliases":[],"title":"KubeVirt: hostDisk feature mounts host files into a VM with insufficient restriction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":8.5,"severity":"high","kev":false,"impact":"hostDisk feature mounts host files into a VM with insufficient restriction; host data exposure","attack_vector":"Cluster user able to create a VM with hostDisk","remediation":"Upgrade KubeVirt to 1.6.1/1.7.0+; disable the hostDisk feature gate for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64324"],"status":"curated"},{"id":"CVE-2026-20898","cve":"CVE-2026-20898","aliases":["INTEL-SA-01439","CVE-2025-20004","CVE-2025-24305","INTEL-SA-01273","INTEL-SA-01313"],"title":"Alias Checking Trusted Module (ACTM) firmware for Intel Xeon processors, including Xeon 6: Improper access control in ACTM, the Intel-signed module that validates memory-aliasing configuration as…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Alias Checking Trusted Module (ACTM) firmware for Intel Xeon processors, including Xeon 6","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Improper access control in ACTM, the Intel-signed module that validates memory-aliasing configuration as part of the platform's trusted boot and confidential-computing plumbing on current Xeon parts. An adversary positioned in startup code or SMM can escalate through it, which undermines the platform integrity guarantee that ACTM exists to provide. On the newest Xeon generations this is the layer that a confidential-computing or attestation story is built on, so a defect here means the node's integrity claims to a tenant cannot be relied upon. Persistence obtained at this level is firmware-resident, survives host reimage, and carries into the next tenant occupying the node.","attack_vector":"A privileged adversary already executing in startup code or SMM - reached from local root plus an SMM or early-boot bug, or from a supply-chain-modified BIOS image.","remediation":"Platform firmware/BIOS update carrying the fixed ACTM, from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta, Wiwynn). Host reboot and job drain. This is current-generation silicon, so it directly affects the Xeon 6 head nodes and CPU hosts under new GPU deployments - budget for it in the same maintenance window as GPU firmware. There is no way to disable ACTM, and no host-side mitigation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-20898","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01439.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01273.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24260","cve":"CVE-2026-24260","aliases":[],"title":"Container Toolkit: Container escape to host root via TOCTOU race","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Container escape to host root via TOCTOU race","attack_vector":"Any tenant that can start a container on a GPU node","remediation":"Bump nvidia-container-toolkit + restart container runtime on every GPU node; upgrade GPU Operator chart; evict and re-admit tenant workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24260","https://github.com/NVIDIA/product-security/tree/main/2026/5850"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-367"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-25628","cve":"CVE-2026-25628","aliases":[],"title":"Qdrant (`/logger`): Append to arbitrary files via the logger endpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Qdrant (`/logger`)","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Append to arbitrary files via the logger endpoint","attack_vector":"Network user of the Qdrant API","remediation":"Upgrade to 1.16.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-25628"],"status":"curated"},{"id":"CVE-2026-44944","cve":"CVE-2026-44944","aliases":["CVE-2026-44943","CVE-2026-55995"],"title":"open-iscsi / open-isns - iscsiuio control socket authorization and iSNS record handling: Three related defects in the initiator stack that ships on every Linux host doing iSCSI. Unprivileged local…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"open-iscsi / open-isns - iscsiuio control socket authorization and iSNS record handling","year":"2026","cvss_score":8.5,"severity":"high","kev":false,"impact":"Three related defects in the initiator stack that ships on every Linux host doing iSCSI. Unprivileged local users can use the iscsiuio control socket, which controls the iSCSI offload/boot path on the NIC. A path-traversal issue lets a machine-in-the-middle attacker cause root-owned files to be created outside the iSCSI node database and inject lines into a record, which means an on-path attacker on the storage network can influence which target the host connects to on subsequent boots. And a double free in open-isns lets an unauthenticated MITM crash the daemon. For a GPU fleet where nodes boot or mount datasets over iSCSI, the middle one is the sharp end: it turns a passive network position into control over what a node believes its storage is.","attack_vector":"Local unprivileged user on the initiator host for the socket issue; machine-in-the-middle on the unauthenticated, unencrypted storage network for the path traversal and the open-isns double free.","remediation":"Package update for open-iscsi and open-isns on every initiator - i.e. every GPU node. It is a userspace daemon update, so a restart of iscsiuio/iscsid rather than a reboot, though sessions established through the offload path may need to be re-established. Disable iscsiuio entirely on hosts that do not use hardware iSCSI offload - most do not, and that removes the local-privilege surface outright. The MITM issues are an argument for CHAP mutual authentication or IPsec on the storage VLAN, since the underlying protocol offers no integrity otherwise.","references":["https://bugzilla.suse.com/show_bug.cgi?id=CVE-2026-44944","https://github.com/open-iscsi/open-iscsi/commit/668ca1df9c9a1e9bdd5c999ae1d67c9c8909237e","https://nvd.nist.gov/vuln/detail/CVE-2026-44944"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-13139","cve":"CVE-2019-13139","aliases":[],"title":"Docker / moby: Command execution via crafted remote git build path in `docker build`","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2019","cvss_score":8.4,"severity":"high","kev":false,"impact":"Command execution via crafted remote git build path in `docker build`","attack_vector":"Anyone who can submit a build to a shared builder","remediation":"Upgrade Docker Engine; isolate tenant build runners","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13139"],"status":"curated"},{"id":"CVE-2021-1051","cve":"CVE-2021-1051","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A local user gets elevated enough to rewrite display configuration through the escape handler, taking the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":8.4,"severity":"high","kev":false,"impact":"A local user gets elevated enough to rewrite display configuration through the escape handler, taking the display subsystem out. On a headless compute node the display path matters less, but the underlying escape-handler privilege gap is the same one that produces worse bugs.","attack_vector":"Any local user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1051"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-33162","cve":"CVE-2021-33162","aliases":[],"title":"Intel Ethernet Adapter manageability firmware (access control): TENANT ISOLATION: improper access control in Intel Ethernet adapter manageability firmware, allowing an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Adapter manageability firmware (access control)","year":"2021","cvss_score":8.4,"severity":"high","kev":false,"impact":"TENANT ISOLATION: improper access control in Intel Ethernet adapter manageability firmware, allowing an authenticated user to escalate privilege. Same NC-SI sideband surface as its sibling — the difference is it needs an account, which in a bare-metal rental means the tenant. A tenant escalating through the manageability engine is a tenant reaching toward the BMC of the machine they rented, and from there toward the operator's management network.","attack_vector":"Authenticated user — on bare metal, the tenant with host access.","remediation":"OEM firmware bundle update for adapter manageability firmware plus cold power cycle. Verify NC-SI is actually needed on your platform; where the server has a dedicated BMC NIC, disabling the sideband channel in BIOS/BMC config removes this path entirely and is a config change rather than a flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33162"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-21163","cve":"CVE-2022-21163","aliases":[],"title":"Crypto API Toolkit for Intel SGX: Improper access control in the SGX Crypto API Toolkit lets an authenticated user escalate privilege. The…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Crypto API Toolkit for Intel SGX","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"Improper access control in the SGX Crypto API Toolkit lets an authenticated user escalate privilege. The toolkit is what many deployments use to put HSM-style key operations inside an enclave, so a break here reaches the keys the enclave was protecting.","attack_vector":"Authenticated user of a system running the Crypto API Toolkit.","remediation":"Upgrade to Crypto API Toolkit 2.0 (commit 91ee496) or later and rotate any keys the toolkit held. Userspace/enclave update - requires re-signing and re-attesting the enclave; no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21163","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00746.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-28627","cve":"CVE-2022-28627","aliases":["HPESBHF04333"],"title":"HPE iLO 5 (local privilege escalation to code execution): An unprivileged user can execute arbitrary code in the iLO context. This is one of a batch of a dozen-plus…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 (local privilege escalation to code execution)","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"An unprivileged user can execute arbitrary code in the iLO context. This is one of a batch of a dozen-plus iLO 5 defects HPE fixed in the same firmware release, which is itself the useful signal: the iLO 5 codebase had a cluster of memory-safety and access-control problems reachable without high privilege. Any one of them lands the attacker on the service processor, with the usual consequences - power, Virtual Media boot, console, and persistence below the hypervisor.","attack_vector":"An unprivileged actor with local access to the iLO's own execution environment - in practice, someone who already has a low-privilege iLO account or a foothold reached through one of the sibling defects in the same batch. Not an unauthenticated internet-facing entry point.","remediation":"Flash iLO 5 to v2.71 or later - and note that the later CVE-2022-28639 batch requires v2.72, so go straight to the highest available iLO 5 build rather than patching to the floor. Out-of-band, per-node, no host reboot and no drain. Treat this as one flash covering a whole batch of CVEs, which makes the per-node rollout cost-effective.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04333en_us","https://nvd.nist.gov/vuln/detail/CVE-2022-28627"],"status":"curated"},{"id":"CVE-2022-41736","cve":"CVE-2022-41736","aliases":[],"title":"IBM Spectrum Scale Container Native Storage Access: A local user obtains root privileges through the Spectrum Scale container-native storage layer. Root on a…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale Container Native Storage Access","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"A local user obtains root privileges through the Spectrum Scale container-native storage layer. Root on a node that mounts the shared training filesystem means access to whatever that node can see — which on a GPFS cluster is typically a very large namespace shared across tenants. Companion issue CVE-2022-43831 is the same shape via missing security-context settings.","attack_vector":"Local user on a node running Container Native Storage Access 5.1.2.1 through 5.1.6.0.","remediation":"Upgrade past 5.1.6.0 (rolling). Independently, enforce restrictive Kubernetes security contexts on the storage-access pods — a manifest change you can apply immediately and that closes the CVE-2022-43831 variant on its own.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41736","https://nvd.nist.gov/vuln/detail/CVE-2022-43831"],"status":"curated"},{"id":"CVE-2022-42271","cve":"CVE-2022-42271","aliases":[],"title":"DGX servers (BMC firmware < 2.09.00): RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers (BMC firmware < 2.09.00)","year":"2022","cvss_score":8.4,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash BMC to 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42271","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"]},{"id":"CVE-2023-0208","cve":"CVE-2023-0208","aliases":[],"title":"NVIDIA DCGM - nv-hostengine: A heap-based buffer overflow reachable through the bound socket gives denial of service and data tampering…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DCGM - nv-hostengine","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"A heap-based buffer overflow reachable through the bound socket gives denial of service and data tampering with a changed CVSS scope, against a root-privileged daemon on every GPU node.","attack_vector":"Local or network depending on how you bound the socket. If nv-hostengine is listening on a routable interface, any host on that network can reach it.","remediation":"Update DCGM per bulletin 5453 and restart nv-hostengine. Cost: telemetry gap of seconds, no GPU job impact, no drain. Restrict the listening socket to localhost as a standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0208","https://github.com/NVIDIA/product-security/tree/main/2023/5453"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:H/A:H","cwe":["CWE-122"]},{"id":"CVE-2023-22649","cve":"CVE-2023-22649","aliases":[],"title":"Rancher: Sensitive data leaked into Rancher audit logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"Sensitive data leaked into Rancher audit logs","attack_vector":"Anyone with audit-log read access","remediation":"Upgrade Rancher; scrub and re-secure audit logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22649"],"status":"curated"},{"id":"CVE-2023-31100","cve":"CVE-2023-31100","aliases":[],"title":"Phoenix SecureCore Technology 4 (SMI handler, improper access control): An SMI handler with missing access control lets an attacker modify the SPI flash. That is the direct route to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phoenix SecureCore Technology 4 (SMI handler, improper access control)","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"An SMI handler with missing access control lets an attacker modify the SPI flash. That is the direct route to a permanent firmware implant: rewrite the boot flash and the compromise survives OS reinstall, disk replacement and node reimaging between tenants, while sitting below Secure Boot and below anything attestation can honestly measure. Highest-scored Phoenix advisory in this set.","attack_vector":"Local attacker on the host with the ability to invoke the SMI handler - in practice admin/root, or a tenant with kernel access on bare metal.","remediation":"OEM BIOS update on the fixed SecureCore Technology 4 build (affected from 4.3.0.0). Firmware flash, one reboot per node. No config workaround for the handler itself, but this is the case where platform SPI protections earn their keep: confirm BIOS Lock Enable, SMM BIOS Write Protect and protected range registers are actually set on your platform - many OEM defaults leave at least one of them off, and a properly locked flash blunts the primitive even on unpatched firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31100"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31323","cve":"CVE-2023-31323","aliases":[],"title":"AMD Secure Processor - XGMI Trusted Agent (type confusion): MULTI-TENANT ISOLATION: Type confusion in the ASP's XGMI Trusted Agent means a malformed argument is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - XGMI Trusted Agent (type confusion)","year":"2023","cvss_score":8.4,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Type confusion in the ASP's XGMI Trusted Agent means a malformed argument is interpreted as the wrong kind of object, producing a memory-safety violation inside the secure processor. On a multi-GPU Instinct node this is a path from host-privileged code to controlling the trusted agent that governs the GPU interconnect - the highest-scoring of the XGMI pair at 8.4.","attack_vector":"Local, host-privileged, on multi-GPU XGMI-connected platforms.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Same campaign as the XGMI TOCTOU issue - patch both in one OEM BIOS pass on MI-series nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31323","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-1708","cve":"CVE-2024-1708","aliases":[],"title":"ConnectWise ScreenConnect: Path traversal enabling remote code execution","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ConnectWise ScreenConnect","year":"2024","cvss_score":8.4,"severity":"high","kev":true,"impact":"[KEV] Path traversal enabling remote code execution; chained with CVE-2024-1709","attack_vector":"Network (remote)","remediation":"Control-plane: same patch; audit installed extensions","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1708"],"status":"curated"},{"id":"CVE-2024-36352","cve":"CVE-2024-36352","aliases":[],"title":"AMD Graphics Driver - crafted pointer leading to arbitrary writes: MULTI-TENANT ISOLATION: A specially crafted pointer passed to the AMD graphics driver produces arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Graphics Driver - crafted pointer leading to arbitrary writes","year":"2024","cvss_score":8.4,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A specially crafted pointer passed to the AMD graphics driver produces arbitrary writes or a crash. Arbitrary kernel writes from a GPU driver call is a privilege-escalation primitive available to anything with the device open - on a shared GPU node, that is the tenant.","attack_vector":"Local, via the graphics driver interface.","remediation":"Update the AMD graphics driver and reload or reboot. Driver-level fix, no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36352","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-12007","cve":"CVE-2025-12007","aliases":[],"title":"Supermicro BMC firmware update signature/validation logic on the X13SEM-F motherboard family: The operator loses the ability to trust or verify what firmware is actually running on the node. An…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware update signature/validation logic on the X13SEM-F motherboard family","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"The operator loses the ability to trust or verify what firmware is actually running on the node. An attacker who can push an update writes their own BMC image to flash, and from that moment the BMC reports whatever the attacker wants it to report - version strings, attestation values, health telemetry. Reimaging the host does nothing, replacing the NVMe does nothing, and a fleet-wide firmware inventory will show the node as compliant. For a neocloud reselling bare metal this is the worst class of finding, because the implant persists across tenant boundaries and the next tenant has no way to detect it. The code path that is supposed to prove a firmware image came from Supermicro before writing it to the BMC's flash.","attack_vector":"Local access to the node's firmware update path - a host-side root process reaching the BMC over the KCS/in-band interface, or an operator-adjacent process with permission to invoke the update. This is the realistic post-exploitation move after a tenant escapes to host root on bare metal, or after any compromise of the provisioning tooling that flashes firmware during node turnup.","remediation":"Firmware flash with the fixed BMC image from Supermicro's January 2026 BMC/IPMI advisory batch. Note the ordering problem: an already-implanted BMC can lie about accepting the update, so for any node you suspect was touched you need an out-of-band SPI reflash with a hardware programmer rather than a software update, which is a hands-on-metal job per node. Going forward, restrict who can call the firmware update path at all - remove host-side IPMI/KCS access from tenant-facing bare metal images, and treat firmware updates as a privileged provisioning operation rather than something any admin session can trigger.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12007","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/12xxx/CVE-2025-12007.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23356","cve":"CVE-2025-23356","aliases":[],"title":"NVIDIA Isaac Lab (Isaac Sim): SB3 configuration parsing reaches code execution with no privileges and no user interaction required. In an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac Lab (Isaac Sim)","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"SB3 configuration parsing reaches code execution with no privileges and no user interaction required. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5708 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23356","https://github.com/NVIDIA/product-security/tree/main/2025/5708"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"]},{"id":"CVE-2025-33225","cve":"CVE-2025-33225","aliases":[],"title":"NVIDIA Resiliency Extension: Predictable log-file names in the log-aggregation path let an attacker pre-create or hijack the target file…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Resiliency Extension","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Predictable log-file names in the log-aggregation path let an attacker pre-create or hijack the target file, reaching privilege escalation and code execution. Scored 8.4 with no privileges required. The Resiliency Extension is what restarts failed large training jobs, so it runs with broad access across the job's nodes.","attack_vector":"Local, no privileges required, no user interaction. Any account on a node participating in a resilient training job.","remediation":"Update the Resiliency Extension per bulletin 5746 and rebuild training images. Cost: image rebuild and job restart; no host driver or firmware change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33225","https://github.com/NVIDIA/product-security/tree/main/2025/5746"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-61"]},{"id":"CVE-2025-52565","cve":"CVE-2025-52565","aliases":[],"title":"runc: Insufficient checks when bind-mounting /dev/console allow writes to arbitrary host procfs paths","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Insufficient checks when bind-mounting /dev/console allow writes to arbitrary host procfs paths; container escape","attack_vector":"Any tenant workload / malicious image","remediation":"Replace runc on all nodes; drain required to restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52565"],"status":"curated","fleet":{"ubiquity":"Universal - same runc version range","remediation_pain":"`node-drain` - same as above; the fix ships in runc 1.2.8 / 1.3.3 / 1.4.0-rc.3 and only applies to newly created containers","pain_class":"node-drain","why_fleet_wide":"`/dev/console` bind-mount race/symlink lets runc mount an unexpected target before LSM/mount protections apply, granting write access to procfs and a breakout"}},{"id":"CVE-2025-54886","cve":"CVE-2025-54886","aliases":[],"title":"skops (`Card.get_model`): Model card loading has no trusted-types check","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"skops (`Card.get_model`)","year":"2025","cvss_score":8.4,"severity":"high","kev":false,"impact":"Model card loading has no trusted-types check → code execution","attack_vector":"Customer-supplied model card / repo","remediation":"Upgrade past 0.12.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54886"],"status":"curated"},{"id":"CVE-2026-24233","cve":"CVE-2026-24233","aliases":[],"title":"TensorRT-LLM: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild and redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24233","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-33747","cve":"CVE-2026-33747","aliases":[],"title":"BuildKit: Custom frontend can craft an API message causing daemon compromise","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"Custom frontend can craft an API message causing daemon compromise","attack_vector":"Anyone who can supply a build frontend","remediation":"Upgrade BuildKit to 0.28.1+; restrict custom frontends","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33747"],"status":"curated"},{"id":"CVE-2026-35204","cve":"CVE-2026-35204","aliases":[],"title":"Helm: Crafted plugin writes its contents to an arbitrary filesystem location on install or update","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"Crafted plugin writes its contents to an arbitrary filesystem location on install or update","attack_vector":"Malicious Helm plugin","remediation":"Upgrade Helm to 4.1.4+; restrict plugin sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35204"],"status":"curated"},{"id":"CVE-2026-35205","cve":"CVE-2026-35205","aliases":[],"title":"Helm: Helm installs plugins with no provenance file even when signature verification is required","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"Helm installs plugins with no provenance file even when signature verification is required; supply-chain gate silently fails open","attack_vector":"Malicious Helm plugin","remediation":"Upgrade Helm to 4.1.4+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35205"],"status":"curated"},{"id":"CVE-2026-53492","cve":"CVE-2026-53492","aliases":[],"title":"containerd: CRI trusts CDI annotations inside untrusted checkpoint image metadata","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":8.4,"severity":"high","kev":false,"impact":"CRI trusts CDI annotations inside untrusted checkpoint image metadata; can inject device access (directly relevant to GPU CDI devices)","attack_vector":"Malicious checkpoint image","remediation":"Rolling containerd upgrade with node drain; disable checkpoint import. High priority for GPU hosts since CDI is how GPUs are injected","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53492"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2019-9507","cve":"CVE-2019-9507","aliases":[],"title":"Vertiv Avocent UMG-4000 universal management gateway: Every command the UMG-4000's web interface runs executes as root on the underlying OS. An admin-authenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Vertiv Avocent UMG-4000 universal management gateway","year":"2019","cvss_score":8.3,"severity":"high","kev":false,"impact":"Every command the UMG-4000's web interface runs executes as root on the underlying OS. An admin-authenticated attacker who can inject shell syntax into a web form gets root on the gateway — which sits between the operator and every server it's providing KVM/serial access to.","attack_vector":"Requires an authenticated administrator session on the web interface; the app fails to neutralize shell metacharacters before executing commands.","remediation":"Software/firmware upgrade from Vertiv; download the fixed build from Vertiv's Avocent UMG support page and flash each gateway. Since the UMG-4000 is the aggregation point for KVM access to many downstream nodes, schedule the update in a maintenance window and expect KVM sessions through that gateway to drop during the flash.","references":["https://www.vertiv.com/en-us/support/software-download/it-management/avocent-universal-management-gateway-appliance--software-downloads/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-23277","cve":"CVE-2021-23277","aliases":[],"title":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - dynamic eval: Unauthenticated eval injection: user-controlled code syntax reaches a dynamic evaluation path. Second…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - dynamic eval","year":"2021","cvss_score":8.3,"severity":"high","kev":false,"impact":"Unauthenticated eval injection: user-controlled code syntax reaches a dynamic evaluation path. Second unauthenticated route to code execution on the same power-management server, from the same advisory batch - which is the point worth taking away. Patching one CVE in this batch and not the rest leaves the door open.","attack_vector":"Unauthenticated, remote, to the IPM server.","remediation":"Upgrade to IPM 1.69 or later - the whole CVE-2021-23276 through -23281 batch lands in one release, so treat it as a single action.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23277"],"status":"curated"},{"id":"CVE-2021-39155","cve":"CVE-2021-39155","aliases":[],"title":"Istio: Case-sensitivity mismatch in host matching bypasses authorization policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":8.3,"severity":"high","kev":false,"impact":"Case-sensitivity mismatch in host matching bypasses authorization policy","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39155"],"status":"curated"},{"id":"CVE-2022-25987","cve":"CVE-2022-25987","aliases":["Trojan Source"],"title":"Intel C++ Compiler Classic / oneAPI toolkits (Unicode source handling): Improper handling of Unicode bidirectional and homoglyph characters in source code means the compiler can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel C++ Compiler Classic / oneAPI toolkits (Unicode source handling)","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Improper handling of Unicode bidirectional and homoglyph characters in source code means the compiler can build something materially different from what a reviewer reads. This is the Trojan Source class - a supply-chain problem for anyone compiling third-party or contributor-supplied kernels and operators into their inference stack.","attack_vector":"Anyone who can get source into your build - an internal contributor, a vendored dependency, or a model-op repo you compile from.","remediation":"Upgrade the compiler to 2021.6 / oneAPI 2022.2 or later, and add a CI check that rejects bidirectional control characters in source. Build-toolchain change only - no node reboot, no firmware. Rebuild any artefact compiled with an affected compiler if provenance matters.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-25987","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated"},{"id":"CVE-2022-26843","cve":"CVE-2022-26843","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler (homoglyph rendering): Homoglyph characters are not visually distinguished by the toolchain, so two different identifiers can look…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI DPC++/C++ compiler (homoglyph rendering)","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Homoglyph characters are not visually distinguished by the toolchain, so two different identifiers can look identical in review. Companion to the bidirectional-override issue and part of the same supply-chain risk when compiling contributed GPU kernels.","attack_vector":"Anyone whose source reaches your compiler.","remediation":"Upgrade to oneAPI 2022.1 or later and screen source for confusable identifiers in CI. Toolchain-only, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26843","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-26872","cve":"CVE-2022-26872","aliases":[],"title":"AMI MegaRAC: Password reset interception via the API — attacker takes over an admin BMC account","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Password reset interception via the API — attacker takes over an admin BMC account","attack_vector":"Network / BMC API","remediation":"BMC firmware update; interim mitigation is disabling the self-service password reset flow","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26872"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31034","cve":"CVE-2022-31034","aliases":[],"title":"Argo CD: Predictable SSO state values allow authentication bypass during login","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Predictable SSO state values allow authentication bypass during login","attack_vector":"Unauthenticated network in a position to observe or race a login","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31034"],"status":"curated"},{"id":"CVE-2022-31105","cve":"CVE-2022-31105","aliases":[],"title":"Argo CD: Improper certificate validation lets Argo CD be tricked into trusting a hostile endpoint","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Improper certificate validation lets Argo CD be tricked into trusting a hostile endpoint","attack_vector":"Unauthenticated network in a MITM position","remediation":"Rolling Argo CD upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31105"],"status":"curated"},{"id":"CVE-2022-40259","cve":"CVE-2022-40259","aliases":[],"title":"AMI MegaRAC: Default credentials — Redfish API accessible with shipped account","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":8.3,"severity":"high","kev":false,"impact":"Default credentials — Redfish API accessible with shipped account; full BMC control","attack_vector":"Network / Redfish API","remediation":"Credential rotation at rack intake; treat any node received from an ODM as compromised-by-default until BMC accounts are reset","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40259"],"status":"curated"},{"id":"CVE-2023-0683","cve":"CVE-2023-0683","aliases":["LEN-99936"],"title":"Lenovo XClarity Controller (XCC) - API privilege escalation: A read-only XCC user gains elevated privileges through a specifically crafted API call. Same shape as the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - API privilege escalation","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"A read-only XCC user gains elevated privileges through a specifically crafted API call. Same shape as the later 2023 XCC batch and the same consequence: a monitoring-tier credential becomes control of the service processor, which means out-of-band power, Virtual Media boot of an attacker image, console into whatever the tenant is running, and firmware-level persistence that survives reimaging. Taken together with CVE-2023-4606 and CVE-2023-4607, this is a pattern rather than an isolated defect - XCC's API-side authorisation checks were repeatedly incomplete across 2023.","attack_vector":"An authenticated XCC account with read-only access, reaching the XCC API over the out-of-band management VLAN.","remediation":"Flash XCC to the version listed for your model in LEN-99936. Out-of-band, per-node, no host reboot and no job drain. Sequence it with the other 2023 XCC advisories so each node is touched once. Given three independent authorisation bypasses in the same year, the durable posture is to stop treating XCC read-only accounts as low-risk and to gate the XCC management segment tightly.","references":["https://support.lenovo.com/us/en/product_security/LEN-99936","https://nvd.nist.gov/vuln/detail/CVE-2023-0683"],"status":"curated"},{"id":"CVE-2023-25533","cve":"CVE-2023-25533","aliases":[],"title":"DGX H100 BMC: Privesc + code execution (web UI input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Privesc + code execution (web UI input validation)","attack_vector":"Authenticated BMC user on mgmt network","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:L","cwe":["CWE-20"]},{"id":"CVE-2023-28083","cve":"CVE-2023-28083","aliases":["HPESBHF04456"],"title":"HPE iLO 4 / iLO 5 / iLO 6 (remote cross-site scripting): Cross-site scripting in the iLO web interface across all three current generations. The reason a browser bug…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 4 / iLO 5 / iLO 6 (remote cross-site scripting)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Cross-site scripting in the iLO web interface across all three current generations. The reason a browser bug scores this high on a BMC is that the victim is an operator with an authenticated iLO session: script running in that session can drive the same actions the operator can - mount Virtual Media, change boot order, power-cycle, create accounts - from the operator's own browser and credentials. It converts a phishing link into out-of-band control of a node without the attacker ever needing network reachability to the iLO themselves.","attack_vector":"Requires tricking an authenticated operator into loading attacker-controlled content while they hold an iLO session. The attacker does not need to reach the management VLAN at all - the operator's browser is the bridge, which is precisely why 'the BMCs are on an isolated network' is not a complete answer.","remediation":"Flash iLO 4 to v2.82, iLO 5 to v2.78, or iLO 6 to v1.20 or later. Out-of-band, per-node, no host reboot and no job drain. Operational controls that help independently: use a dedicated browser profile or a privileged access workstation for BMC administration, and do not leave iLO sessions open in a browser that also handles general web traffic and email.","references":["https://support.hpe.com/hpesc/public/docDisplay?docLocale=en_US&docId=hpesbhf04456en_us","https://nvd.nist.gov/vuln/detail/CVE-2023-28083"],"status":"curated"},{"id":"CVE-2023-31009","cve":"CVE-2023-31009","aliases":[],"title":"DGX H100 BMC (REST): Code execution + privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Code execution + privesc","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-20"]},{"id":"CVE-2023-37294","cve":"CVE-2023-37294","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): Heap corruption in the BMC reachable without credentials. A reliable exploit gives BMC code execution and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Heap corruption in the BMC reachable without credentials. A reliable exploit gives BMC code execution and therefore persistent, below-the-OS control of the node; a sloppy one just crashes the BMC, which on a GPU node means losing remote power control and console right when you need it - the node keeps running the training job but becomes un-manageable until someone walks the row.","attack_vector":"Adjacent network, unauthenticated, but high attack complexity - the attacker needs heap grooming or a race to win, so this is a targeted-effort bug rather than a spray. Precondition is still just L2 reachability to the BMC NIC.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6, out-of-band and per node, subject to ODM rebase. Same rollout cost as the rest of the AMI-SA-2023010 batch, so treat all eight CVEs in that advisory as one flash campaign rather than eight tickets. Interim control is network segmentation of the BMC plane.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-37294"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-37296","cve":"CVE-2023-37296","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service): Stack memory corruption in the same unauthenticated BMC parsing surface. Best case for the attacker is BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC network service)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Stack memory corruption in the same unauthenticated BMC parsing surface. Best case for the attacker is BMC code execution and a firmware-level foothold on the node; worst case for the operator without an exploit is a BMC that wedges and needs a physical AC cycle to recover, which on a dense GPU rack means a hands-on trip and possibly draining neighbouring nodes.","attack_vector":"Adjacent network, no credentials, high complexity. Anything on the management VLAN - including a compromised BMC on a neighbouring node - is close enough.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6. Out-of-band, per node, ODM-gated. No config-only fix inside the BMC; the compensating control is to make the management VLAN unreachable from tenant and general corporate networks and to keep the BMC off any routable address space.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-37296"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-39266","cve":"CVE-2023-39266","aliases":[],"title":"ArubaOS-Switch web management interface: Unauthenticated stored cross-site scripting against the ArubaOS-Switch web UI. Stored XSS in a switch…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ArubaOS-Switch web management interface","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Unauthenticated stored cross-site scripting against the ArubaOS-Switch web UI. Stored XSS in a switch management interface is a credential-theft and config-change path aimed at your network operators: an attacker plants the payload without logging in, and it fires the next time an admin opens the page.","attack_vector":"Unauthenticated, remote to the switch's web management interface; the payload executes in an administrator's browser session.","remediation":"ArubaOS-Switch firmware upgrade plus reload. Immediate mitigation is a config change: disable the web management interface and manage via SSH/CLI, which most datacenter operators should be doing anyway. Related ArubaOS issue: CVE-2023-35971.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39266"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40284","cve":"CVE-2023-40284","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Stored/reflected script injection in the BMC web UI. On its own it is 'just XSS', but on a BMC the session it…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Stored/reflected script injection in the BMC web UI. On its own it is 'just XSS', but on a BMC the session it hijacks is the one that can mount virtual media, power-cycle the node and flash firmware - so it is the entry step of a full out-of-band takeover chain rather than a cosmetic web bug.","attack_vector":"Requires an operator to load an attacker-influenced BMC page. Any engineer who administers BMCs from a browser is the target, and the attacker only needs to be able to plant content the BMC will render.","remediation":"BMC firmware flash per board. Until then, treat BMC web access as a privileged action: dedicated browser profile or jump host, never the same browser session used for general web browsing, and no BMC on a routable network.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40284"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40287","cve":"CVE-2023-40287","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Script injection in the BMC management UI, scope-changing because the compromised session controls power…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Script injection in the BMC management UI, scope-changing because the compromised session controls power, console and firmware on the physical node.","attack_vector":"Network reach to the BMC web interface plus operator interaction.","remediation":"BMC firmware flash per board. Same rollout as the rest of the batch; the interim control is network isolation of the BMC plane, not browser hygiene alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40287"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40288","cve":"CVE-2023-40288","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Further injection point in the same BMC web stack","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Further injection point in the same BMC web stack; the payoff is a hijacked administrative session with out-of-band control of the server.","attack_vector":"Network reach to the BMC web interface plus operator interaction.","remediation":"BMC firmware flash per board, bundled with the batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40288"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40290","cve":"CVE-2023-40290","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS via IE11): Injection that fires specifically through Internet Explorer 11. Worth keeping on the list precisely because…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS via IE11)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Injection that fires specifically through Internet Explorer 11. Worth keeping on the list precisely because BMC web UIs are the last place in a datacenter where an ancient browser is still in the loop - legacy Java KVM clients and jump hosts frozen on old images keep IE alive long after everything else moved on.","attack_vector":"An operator administering the BMC from IE11 on Windows, with network reach to the BMC.","remediation":"BMC firmware flash per board. The cheaper immediate action is retiring IE11 from your management jump hosts entirely, which also kills a class of legacy KVM client exposure at the same time.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40290"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40547","cve":"CVE-2023-40547","aliases":[],"title":"shim (HTTP boot): Out-of-bounds write from a crafted HTTP response during network boot — full system compromise at the pre-boot…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (HTTP boot)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Out-of-bounds write from a crafted HTTP response during network boot — full system compromise at the pre-boot stage, before any OS security control exists","attack_vector":"Adjacent network, MITM on the boot server or a compromised PXE/HTTP boot server","remediation":"New signed shim rollout plus dbx revocation of the old one. Directly relevant to neoclouds because netboot provisioning is the standard bare-metal reimaging path — the provisioning network is the attack surface","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40547"],"status":"curated"},{"id":"CVE-2023-45230","cve":"CVE-2023-45230","aliases":["PixieFail","AMI-SA-2024001"],"title":"AMI AptioV UEFI BIOS (EDK II network stack, DHCPv6 client): Buffer overflow in the firmware's DHCPv6 client, triggered by an over-long server ID option. The attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (EDK II network stack, DHCPv6 client)","year":"2023","cvss_score":8.3,"severity":"high","kev":false,"impact":"Buffer overflow in the firmware's DHCPv6 client, triggered by an over-long server ID option. The attacker gets code execution in the pre-boot firmware environment - before the OS, before Secure Boot has finished mattering, with full access to the platform. For a GPU fleet the sharp edge is network provisioning: PXE and HTTP boot are how most clusters bring nodes up, and the firmware talks IPv6 DHCP during that window whether or not you intended to use IPv6. An attacker on the provisioning VLAN answers first and owns the node before it has an operating system.","attack_vector":"Adjacent network, unauthenticated, no interaction - the attacker just needs to be on the same segment as the booting node and respond to its DHCPv6 solicitation faster than your real server. Every node reboot is a fresh opportunity, so on a fleet that autoscales or recovers nodes constantly the window is effectively always open. Note that IPv6 is exercised even in IPv4-only deployments because the firmware stack solicits regardless.","remediation":"BIOS update carrying the patched EDK II network package - firmware flash plus host reboot per node, gated on your server vendor rebasing AMI's AptioV build; AMI's advisory names 'AptioV' generically rather than a BKC version, so confirm the specific BIOS release with your OEM. There is a real config-only mitigation and you should apply it regardless: disable PXE and network boot in BIOS on nodes that do not need it, and where you do need it, isolate the provisioning VLAN so nothing untrusted can answer DHCP on it. That is a BIOS setup change plus one reboot per node, far cheaper than the flash.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024001.pdf","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h","https://nvd.nist.gov/vuln/detail/CVE-2023-45230"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-22424","cve":"CVE-2024-22424","aliases":[],"title":"Argo CD: CSRF against the Argo CD API allows deploying arbitrary workloads with Argo's cluster-admin rights","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":8.3,"severity":"high","kev":false,"impact":"CSRF against the Argo CD API allows deploying arbitrary workloads with Argo's cluster-admin rights","attack_vector":"An attacker who gets a logged-in operator to visit a hostile page","remediation":"Rolling Argo CD upgrade; no GPU drain. Also scope Argo's cluster credentials down","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22424"],"status":"curated"},{"id":"CVE-2024-47084","cve":"CVE-2024-47084","aliases":[],"title":"Gradio: CORS origin validation bypass","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":8.3,"severity":"high","kev":false,"impact":"CORS origin validation bypass → cross-origin access to the demo API","attack_vector":"Malicious page visited by the demo user","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47084"],"status":"curated"},{"id":"CVE-2025-23359","cve":"CVE-2025-23359","aliases":[],"title":"Container Toolkit: Container escape to host filesystem (bypass of the CVE-2024-0132 fix)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2025","cvss_score":8.3,"severity":"high","kev":false,"impact":"Container escape to host filesystem (bypass of the CVE-2024-0132 fix)","attack_vector":"Any tenant that can run an arbitrary container image","remediation":"Bump nvidia-container-toolkit to 1.17.4+ and restart the runtime on every GPU node; upgrade GPU Operator to 24.9.2+; evict tenant workloads","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=NVIDIA%20Container%20Toolkit"],"status":"curated","fleet":{"ubiquity":"Universal - all versions <= 1.17.3, i.e. everyone who thought they had already patched 0132","remediation_pain":"`daemon-restart` to 1.17.4 / GPU Operator 24.9.2 - a second forced patch cycle across the same fleet within five months","pain_class":"daemon-restart","why_fleet_wide":"Proves the class is not one-and-done: a mount-path TOCTOU bypass re-opens host filesystem access from a crafted container, so every operator who patched in Sept 2024 had to re-patch the whole fleet in Feb 2025"},"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-367"]},{"id":"CVE-2025-67601","cve":"CVE-2025-67601","aliases":[],"title":"Rancher: CLI login with -skip-verify and no --cacert silently accepts any certificate","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2025","cvss_score":8.3,"severity":"high","kev":false,"impact":"CLI login with -skip-verify and no --cacert silently accepts any certificate; MITM on the management API","attack_vector":"Unauthenticated network in a MITM position","remediation":"Upgrade Rancher CLI; ban -skip-verify in runbooks","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-67601"],"status":"curated"},{"id":"CVE-2026-22621","cve":"CVE-2026-22621","aliases":["eaton-va-2026-1005"],"title":"Eaton Tripp Lite series PADM firmware, session management interface: An authenticated administrator can break out of the restricted shell and run arbitrary commands on the PDU.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Tripp Lite series PADM firmware, session management interface","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"An authenticated administrator can break out of the restricted shell and run arbitrary commands on the PDU. That turns a device you thought was an appliance into a persistent Linux foothold sitting on your out-of-band network, below every server it powers and outside any endpoint tooling you run.","attack_vector":"Requires administrator credentials on the PDU - which, given the companion authentication bypass, an unauthenticated attacker can obtain first. Chain the two and this is unauthenticated remote code execution on rack power infrastructure.","remediation":"PADM firmware update, or hardware replacement for EOL SKUs. Rotate PDU admin credentials, which are very commonly shared fleet-wide from the original commissioning.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/eaton-va-2026-1005.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24148","cve":"CVE-2026-24148","aliases":[],"title":"Jetson Xavier / Orin: Authentication bypass in a network service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson Xavier / Orin","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"Authentication bypass in a network service","attack_vector":"Network-adjacent unauthenticated","remediation":"Flash JetPack; edge fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24148","https://github.com/NVIDIA/product-security/tree/main/2026/5797"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:L","cwe":["CWE-1188"]},{"id":"CVE-2026-41490","cve":"CVE-2026-41490","aliases":[],"title":"Dagster: Vulnerability in Dagster Core prior to 1.13.1","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dagster","year":"2026","cvss_score":8.3,"severity":"high","kev":false,"impact":"Vulnerability in Dagster Core prior to 1.13.1","attack_vector":"Network user of the orchestrator","remediation":"Upgrade to 1.13.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41490"],"status":"curated"},{"id":"CVE-2018-3643","cve":"CVE-2018-3643","aliases":["INTEL-SA-00131"],"title":"Power Management Controller (PMC) firmware in systems using Intel CSME 11.x/12.0 or Intel SPS 4.x: An administrative attacker can reach the platform's Power Management Controller firmware. The PMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Power Management Controller (PMC) firmware in systems using Intel CSME 11.x/12.0 or Intel SPS 4.x","year":"2018","cvss_score":8.2,"severity":"high","kev":false,"impact":"An administrative attacker can reach the platform's Power Management Controller firmware. The PMC owns the platform power state machine and rails - so this is one of the few software-reachable paths with a direct physical outcome: forced or blocked power transitions on a node, and manipulation of the power/thermal control loop that the rest of the platform trusts. On GPU nodes drawing 6-10 kW, an attacker who can hold a node in the wrong power state or misreport its budget to Node Manager can trip breakers at the rack or hall level rather than merely killing one job. The PMC firmware also lives outside the host OS image, so a modification persists across reimage and tenant handoff.","attack_vector":"An attacker with administrative privileges on the platform - local root on the host reaching CSME/SPS via HECI, or an administrator on the management path.","remediation":"CSME/SPS firmware bundle including the PMC firmware, delivered by the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta) as a BIOS/ME package. Host reboot and job drain required; HPE's bundle for this advisory came out well after Intel's September 2018 date, which is the normal pattern. There is no runtime mitigation - PMC firmware cannot be disabled. Operationally, pair the rollout with independent power telemetry at the PDU/branch level so you are not relying solely on platform-reported power for capacity protection.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3643","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00131.html","https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03873en_us"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-3682","cve":"CVE-2018-3682","aliases":["INTEL-SA-00130"],"title":"BMC firmware on Intel server boards, compute modules and systems - SMBus access control: An attacker with administrative privileges on the BMC can issue unauthorized reads and writes on the platform…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BMC firmware on Intel server boards, compute modules and systems - SMBus access control","year":"2018","cvss_score":8.2,"severity":"high","kev":false,"impact":"An attacker with administrative privileges on the BMC can issue unauthorized reads and writes on the platform SMBus. SMBus is the wire that reaches the power supplies (PMBus), the voltage regulators, the DIMM SPD EEPROMs and the temperature sensors. Write access there is a physical-consequence primitive: reprogram a VR or PSU setpoint, corrupt SPD so DIMMs no longer train, or falsify thermal telemetry so the platform does not throttle. On a dense GPU node this can mean a forced power-off, a bricked-until-RMA board, or a thermal event that the DCIM layer never sees coming. It is also persistence - SMBus-attached EEPROM contents survive any host reimage and therefore cross tenant handoff.","attack_vector":"Administrative access to the BMC. That is reached from the out-of-band management network, from any credential reuse across the IPMI/Redfish fleet, or from the host itself via the KCS/host interface if you have not disabled it - which means a tenant with root on a bare-metal node is one BMC bug away from the SMBus.","remediation":"BMC firmware update from Intel or the board OEM (Intel server boards, and the ODMs building on them - Quanta, Wiwynn, Supermicro). BMC flashes generally do not require a host reboot, which makes this one of the cheaper firmware rollouts, but the update must be staged per board family. The structural controls matter more: put the BMC on a network no tenant can reach, use unique per-node BMC credentials, and disable the host-to-BMC KCS/host interface on bare-metal SKUs so a tenant with root cannot talk to the BMC at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3682","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00130.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-11248","cve":"CVE-2019-11248","aliases":[],"title":"Kubernetes (kubelet): /debug/pprof exposed on the unauthenticated kubelet healthz port","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2019","cvss_score":8.2,"severity":"high","kev":false,"impact":"/debug/pprof exposed on the unauthenticated kubelet healthz port; leaks node and workload internals","attack_vector":"Any pod on the cluster network reaching the node's healthz port","remediation":"Rolling kubelet upgrade with node drain; firewall the healthz port to the control plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11248"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2020-10713","cve":"CVE-2020-10713","aliases":["BootHole"],"title":"GRUB2: Buffer overflow in `grub.cfg` parsing allowing Secure Boot bypass and arbitrary code execution inside GRUB —…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2","year":"2020","cvss_score":8.2,"severity":"high","kev":false,"impact":"Buffer overflow in `grub.cfg` parsing allowing Secure Boot bypass and arbitrary code execution inside GRUB — a bootkit that persists across OS reinstall","attack_vector":"Local, or via a modified PXE-boot network","remediation":"dbx (Secure Boot revocation list) push plus a coordinated GRUB/shim/kernel update. The dbx push is the dangerous part: revoking the old shim before every node has the new bootloader leaves the node unbootable, and recovery is out-of-band console work per node","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-10713"],"status":"curated","fleet":{"ubiquity":"universal - GRUB2 + the Microsoft-signed shim is the boot path for nearly every Linux GPU node","remediation_pain":"node-reboot + firmware-flash-class pain: patching GRUB is easy, but the real fix is a **dbx revocation** update pushed into UEFI NVRAM on every node, and a botched dbx push makes the node unbootable - which is why operators delay it for years","pain_class":"firmware-flash","why_fleet_wide":"A config-file buffer overflow in GRUB2 lets attackers run pre-OS bootkits with Secure Boot enabled; because every old signed GRUB stays valid until revoked, the fleet remains exploitable until each node's firmware revocation list is updated."}},{"id":"CVE-2020-15705","cve":"CVE-2020-15705","aliases":["BootHole family"],"title":"GRUB2 (direct kernel boot without shim): When GRUB is booted directly by UEFI rather than chained through shim, it does not verify the kernel…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (direct kernel boot without shim)","year":"2020","cvss_score":8.2,"severity":"high","kev":false,"impact":"When GRUB is booted directly by UEFI rather than chained through shim, it does not verify the kernel signature at all. Any unsigned kernel boots with Secure Boot enabled and reporting healthy - so attestation and the operator's 'verified boot' control are simply false on those nodes. Confidential-computing claims built on measured boot become unverifiable.","attack_vector":"Applies to any node configured to load GRUB directly from the EFI System Partition. The attacker then only needs to drop a kernel, which any local root can do.","remediation":"grub2 package update + reboot, and audit the boot configuration on every node to confirm shim is actually in the chain - a surprising number of custom/netboot images skip it. This is one of the few in this family where a config check is a genuine part of the fix, not just a workaround.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15705","https://ubuntu.com/security/CVE-2020-15705"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-3165","cve":"CVE-2020-3165","aliases":[],"title":"Cisco NX-OS (BGP MD5 authentication): TENANT ISOLATION: BGP MD5 authentication can be bypassed, so an attacker can bring up a BGP session with the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (BGP MD5 authentication)","year":"2020","cvss_score":8.2,"severity":"high","kev":false,"impact":"TENANT ISOLATION: BGP MD5 authentication can be bypassed, so an attacker can bring up a BGP session with the switch without the shared key. In an EVPN fabric that means injecting routes — including type-2 and type-5 EVPN routes — which is exactly how you steer one tenant's traffic to a machine you control. The authentication you configured to prevent unauthorized peering simply does not hold.","attack_vector":"Unauthenticated, remote — an attacker that can reach TCP/179 on the switch and is permitted by the peer-group/neighbor configuration's address range.","remediation":"NX-OS upgrade plus reload. Interim mitigation is control-plane policing plus tight neighbor prefix ACLs so only known peer addresses can open a session at all — live config changes. If you run EVPN, also verify no unexpected routes were learned before you patched.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-3165"],"status":"curated"},{"id":"CVE-2021-21378","cve":"CVE-2021-21378","aliases":[],"title":"Envoy: JWT with an issuer absent from the provider list bypasses JWT authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"JWT with an issuer absent from the provider list bypasses JWT authentication","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21378"],"status":"curated"},{"id":"CVE-2021-21522","cve":"CVE-2021-21522","aliases":["CVE-2021-36285","Dell BIOS NVMe password bypass"],"title":"Dell client and server BIOS - NVMe drive password (SED credential) defeated by resetting the BIOS password via the Manageability Interface: The BIOS-managed NVMe drive password - the credential…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell client and server BIOS - NVMe drive password (SED credential) defeated by resetting the BIOS password via the…","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"The BIOS-managed NVMe drive password - the credential many operators rely on to keep a locked drive locked - can be defeated by resetting the BIOS password through the manageability interface, giving access to data on the NVMe device. The paired issue removes the limit on failed NVMe password attempts, so an administrator can brute-force the drive password instead. BREAKS TENANT HANDOFF wherever platform-level drive locking is your control: the lock lives in a BIOS that a local administrator can reset, so the drive credential inherits the security of the BIOS password rather than of the drive. This is the systems-integration failure mode of SEDs - the drive firmware may be perfectly sound while the platform that holds its credential hands it away.","attack_vector":"A local authenticated user with elevated privilege on the host - which on bare metal means the tenant you just rented the box to, if they have BIOS/manageability reach. The brute-force variant requires local administrator access.","remediation":"Apply the Dell BIOS updates named in Dell's advisory for the affected client and PowerEdge platforms; this is a host BIOS flash, so it needs a reboot and a maintenance window per node but not a drive teardown. Then fix the design, not just the bug: do not use BIOS-held NVMe passwords as your tenant-separation mechanism on bare metal, because the whole scheme assumes the tenant cannot reach platform firmware, and on a rented bare-metal node that assumption is false by definition. Lock down the manageability interface, set and monitor BIOS admin passwords, and move the actual data protection to LUKS/dm-crypt with a key your control plane holds and destroys at reclaim - a key the tenant's BIOS access cannot reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21522","https://nvd.nist.gov/vuln/detail/CVE-2021-36285","https://www.dell.com/support/kbdoc/000191495"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-33627","cve":"CVE-2021-33627","aliases":["INSYDE-SA-2022022","VU#796611"],"title":"Insyde InsydeH2O (FwBlockServiceSmm): Software SMI services reachable through EFI_SMM_COMMUNICATION_PROTOCOL never check whether the buffer address…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FwBlockServiceSmm)","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"Software SMI services reachable through EFI_SMM_COMMUNICATION_PROTOCOL never check whether the buffer address they were given points into SMRAM, MMIO or kernel memory. An OS-level attacker therefore gets SMM to write on their behalf - and FwBlockServiceSmm is the firmware-block service, so this sits directly on the path to the SPI flash. Result is a firmware implant that outlives every reimage and quietly breaks the root of trust the fleet's attestation depends on.","attack_vector":"Local admin/root on the host OS issuing a crafted SMM communication request. No physical access required.","remediation":"Fixed in InsydeH2O kernels 05.09.11 / 05.17.11 / 05.27.11 / 05.36.11 / 05.44.11 / 05.52.11 - delivered to you only as an OEM BIOS image, months downstream. Firmware flash, one reboot per node, drain the GPUs first. No configuration mitigates it. Enable and verify SPI write protection (BIOS Lock / protected range registers) as a partial hardening measure, but that does not close the SMM write primitive itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33627","https://www.insyde.com/security-pledge/SA-2022022","https://kb.cert.org/vuls/id/796611"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-4093","cve":"CVE-2021-4093","aliases":[],"title":"KVM (AMD SEV-ES): Out-of-bounds read/write in sev_es_string_io() - malicious SEV-ES guest corrupts host memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM (AMD SEV-ES)","year":"2021","cvss_score":8.2,"severity":"high","kev":false,"impact":"Out-of-bounds read/write in sev_es_string_io() - malicious SEV-ES guest corrupts host memory","attack_vector":"Tenant VM guest (SEV-ES)","remediation":"Kernel/KVM patch + reboot. Directly relevant to any confidential-VM GPU offering","references":["https://access.redhat.com/security/cve/CVE-2021-4093"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-28200","cve":"CVE-2022-28200","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The BiosCfgTool reads and writes outside its bounds in SMRAM, handing a privileged local user code execution…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"The BiosCfgTool reads and writes outside its bounds in SMRAM, handing a privileged local user code execution in System Management Mode with a changed scope. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28200","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-28735","cve":"CVE-2022-28735","aliases":[],"title":"GRUB2 (shim_lock verifier): The shim_lock verifier let non-kernel files through, so an attacker could get unsigned content loaded into…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (shim_lock verifier)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"The shim_lock verifier let non-kernel files through, so an attacker could get unsigned content loaded into the boot path while Secure Boot enforcement appeared intact. It defeats the exact control operators rely on to promise a clean handoff between bare-metal tenants.","attack_vector":"Local, with the ability to place a file GRUB will load.","remediation":"grub2 package update + reboot. This is one where the dbx revocation genuinely matters - without it the old signed GRUB remains a usable bypass tool that an attacker can simply drop onto a patched node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28735","https://access.redhat.com/security/cve/CVE-2022-28735"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-28737","cve":"CVE-2022-28737","aliases":[],"title":"shim (handle_image PE loader): Buffer overflow in shim's own image loader. Because shim is the Microsoft-signed component every Linux node…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (handle_image PE loader)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"Buffer overflow in shim's own image loader. Because shim is the Microsoft-signed component every Linux node chains through, a bug here is worse than a GRUB bug: it bypasses Secure Boot on any machine that trusts the Microsoft 3rd-party CA, regardless of which distro's GRUB sits behind it.","attack_vector":"Local, with the ability to present a crafted EFI image to shim.","remediation":"shim package update + reboot per node. Revoking the old shim means an SBAT generation bump pushed by Microsoft/vendor updates rather than a dbx entry - track SBAT levels, not just package versions, or you will believe you are fixed when you are not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28737","https://access.redhat.com/security/cve/CVE-2022-28737"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-29275","cve":"CVE-2022-29275","aliases":["INSYDE-SA-2022058"],"title":"Insyde InsydeH2O (UsbCoreDxe, untrusted pointer use): UsbCoreDxe uses pointers it was handed without establishing they point outside SMRAM, so an OS-level caller…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbCoreDxe, untrusted pointer use)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"UsbCoreDxe uses pointers it was handed without establishing they point outside SMRAM, so an OS-level caller gets SMM to tamper with either SMRAM or kernel memory. Ring -2 escalation from a driver that is resident on every node whether or not a USB device is plugged in. Not a DMA race - this one needs only host privilege, which makes it materially easier to exploit than the SA-2022042-057 set.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.21, 5.1 / 05.17.21, 5.2 / 05.27.21, 5.3 / 05.36.21, 5.4 / 05.44.21, 5.5 / 05.52.21. Disabling USB legacy/emulation support in BIOS on headless nodes shrinks the reachable surface without a flash - confirm on your platform that it unloads the SMM module rather than only hiding the setup option. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29275","https://www.insyde.com/security-pledge/SA-2022058"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-29276","cve":"CVE-2022-29276","aliases":["INSYDE-SA-2022059"],"title":"Insyde InsydeH2O (AhciBusDxe, untrusted SMI inputs): SMI functions in the AHCI/SATA driver consume untrusted inputs and corrupt SMRAM. Straight privilege…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AhciBusDxe, untrusted SMI inputs)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"SMI functions in the AHCI/SATA driver consume untrusted inputs and corrupt SMRAM. Straight privilege escalation to ring -2 for anyone with host root - firmware persistence that survives OS reinstall, plus control over the SATA path the node boots from.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.18 through 5.5 / 05.52.18.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29276","https://www.insyde.com/security-pledge/SA-2022059"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-29278","cve":"CVE-2022-29278","aliases":["INSYDE-SA-2022061"],"title":"Insyde InsydeH2O (NvmExpressDxe, incorrect pointer checks): The NVMe driver's pointer validation is wrong, allowing tampering with both SMRAM and OS memory. The driver…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressDxe, incorrect pointer checks)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"The NVMe driver's pointer validation is wrong, allowing tampering with both SMRAM and OS memory. The driver reaches the NVMe data path - datasets, checkpoints, weights on a GPU node - and the bug hands an OS-level attacker ring -2 on top of it. Distinct from the DMA race in SA-2022055 and separately fixed; a node can carry one and not the other.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.1 / 05.17.23 through 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29278","https://www.insyde.com/security-pledge/SA-2022061"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-29279","cve":"CVE-2022-29279","aliases":["INSYDE-SA-2022062"],"title":"Insyde InsydeH2O (SdHostDriver and SdMmcDevice, untrusted pointer use): One advisory covering both SD layers: untrusted pointers allow tampering with SMRAM and OS memory, giving…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdHostDriver and SdMmcDevice, untrusted pointer use)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"One advisory covering both SD layers: untrusted pointers allow tampering with SMRAM and OS memory, giving ring -2 code execution. Easy to under-prioritise because SD/eMMC looks irrelevant on a GPU box, but the driver is compiled in and the SMI is reachable regardless of whether any SD media is present.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.17 through 5.5 / 05.52.17.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29279","https://www.insyde.com/security-pledge/SA-2022062"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-30771","cve":"CVE-2022-30771","aliases":["INSYDE-SA-2022064"],"title":"Insyde InsydeH2O (PnpSmm initialization, SMRAM corruption via later PNP SMIs): An initialization-order defect: PnpSmm's init function leaves state that causes SMRAM corruption when…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm initialization, SMRAM corruption via later PNP SMIs)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"An initialization-order defect: PnpSmm's init function leaves state that causes SMRAM corruption when subsequent PNP SMI functions are called. The interesting operational property is that the damage is deferred - the node boots cleanly and the corruption only lands when something later exercises the PNP path, so this does not look like an attack when it fires.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.1 / 05.17.25, 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30771","https://www.insyde.com/security-pledge/SA-2022064"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-30772","cve":"CVE-2022-30772","aliases":["INSYDE-SA-2022065"],"title":"Insyde InsydeH2O (PnpSmm function 0x52, SMBIOS write address manipulation): PnpSmm function 0x52 takes an address and a size for data to write into the SMBIOS table and does not…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm function 0x52, SMBIOS write address manipulation)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"PnpSmm function 0x52 takes an address and a size for data to write into the SMBIOS table and does not constrain where that address points. Malware supplies its own address and overwrites SMRAM or OS kernel memory - an arbitrary write primitive handed over by a documented firmware function, no race and no exotic hardware needed. The cleanest escalation in the batch.","attack_vector":"Local admin/root on the host OS invoking the vulnerable software SMI with attacker-chosen pointers. On bare-metal GPU rental this is exactly the privilege the tenant already holds on their leased node.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.0 / 05.09.41, 5.1 / 05.17.43, 5.2 / 05.27.30, 5.3 / 05.36.30, 5.4 / 05.44.30, 5.5 / 05.52.30.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30772","https://www.insyde.com/security-pledge/SA-2022065"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31599","cve":"CVE-2022-31599","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An uninitialised pointer in the Ofbd SMM handler gives a privileged local user SMM code execution and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"An uninitialised pointer in the Ofbd SMM handler gives a privileged local user SMM code execution and privilege escalation beyond the component. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31599","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-824"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31705","cve":"CVE-2022-31705","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Heap out-of-bounds write in the USB 2.0 EHCI controller - VM escape to VMX process code execution (GeekPwn…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"Heap out-of-bounds write in the USB 2.0 EHCI controller - VM escape to VMX process code execution (GeekPwn 2022)","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with evacuation; remove USB controllers from tenant VM templates","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31705"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-35408","cve":"CVE-2022-35408","aliases":["INSYDE-SA-2022031"],"title":"Insyde InsydeH2O (UsbLegacyControlSmm): A classic SMM callout: code running inside SMM calls out to a function pointer that lives in memory the OS…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbLegacyControlSmm)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"A classic SMM callout: code running inside SMM calls out to a function pointer that lives in memory the OS can write. An attacker plants their own pointer, triggers the SMI, and their code runs at ring -2. USB legacy support is enabled by default on most server BIOS images, so the attack surface is present on nodes that have no USB device attached at all.","attack_vector":"Local admin/root on the host OS, then a software SMI into the USB legacy handler.","remediation":"OEM BIOS update carrying the fixed Insyde kernel. Firmware flash, one reboot per node. Partial config workaround that is genuinely worth doing on servers: disable USB legacy support / USB emulation in BIOS setup - it is rarely needed on a headless GPU node and can be pushed via the OEM's remote BIOS-settings tooling without a flash. Verify on your platform that the setting actually unloads the driver rather than just hiding the option.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-35408","https://www.insyde.com/security-pledge/SA-2022031"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-36337","cve":"CVE-2022-36337","aliases":["INSYDE-SA-2022039"],"title":"Insyde InsydeH2O (MebxConfiguration DXE driver): A UEFI variable that the OS can write is read back by BIOS code into a fixed-size stack buffer without a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (MebxConfiguration DXE driver)","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"A UEFI variable that the OS can write is read back by BIOS code into a fixed-size stack buffer without a length check. Set the variable from the OS, reboot, and your code runs during DXE - before Secure Boot has finished deciding what is allowed to run. The persistence mechanism is the variable store itself, which means the implant re-arms on every boot and survives disk replacement entirely.","attack_vector":"Local admin/root on the host OS with the ability to write UEFI variables (standard on Linux via efivarfs and on Windows via SetFirmwareEnvironmentVariable), then one reboot.","remediation":"OEM BIOS update built on the fixed Insyde kernel. Firmware flash, reboot per node. There is no config toggle. Detection is possible in the interim: monitor for unexpected writes to the relevant UEFI variables from the OS, and make efivarfs read-only where your workload does not need it. On a fleet, treat any node where firmware variables changed outside a maintenance window as suspect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36337","https://www.insyde.com/security-pledge/SA-2022039"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42291","cve":"CVE-2022-42291","aliases":[],"title":"GeForce Experience installer: Local privesc via untrusted search path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce Experience installer","year":"2022","cvss_score":8.2,"severity":"high","kev":false,"impact":"Local privesc via untrusted search path","attack_vector":"Local user","remediation":"Consumer-only; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42291","https://github.com/NVIDIA/product-security/tree/main/2023/5384"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:C/C:N/I:H/A:H","cwe":["CWE-1386"]},{"id":"CVE-2023-0209","cve":"CVE-2023-0209","aliases":[],"title":"NVIDIA DGX-1 - SBIOS / SMM firmware: The Uncore PEI module never authenticates the code executed by SSA, so a privileged local attacker gets…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-1 - SBIOS / SMM firmware","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"The Uncore PEI module never authenticates the code executed by SSA, so a privileged local attacker gets arbitrary firmware-phase code execution and a full Secure Boot bypass. This is the most serious DGX SBIOS issue in the set - the platform's chain of trust simply does not check. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5458. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0209","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-287"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-2163","cve":"CVE-2023-2163","aliases":[],"title":"Linux kernel (eBPF verifier): Incorrect verifier pruning marks unsafe paths as safe - arbitrary kernel read/write from BPF","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF verifier)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Incorrect verifier pruning marks unsafe paths as safe - arbitrary kernel read/write from BPF","attack_vector":"Any tenant process in a container where BPF is reachable","remediation":"Livepatchable; otherwise drain + reboot. `kernel.unprivileged_bpf_disabled=1`; audit any tenant-facing eBPF observability feature","references":["https://access.redhat.com/security/cve/CVE-2023-2163"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-26484","cve":"CVE-2023-26484","aliases":[],"title":"KubeVirt: A compromised node's virt-handler service account can be abused cluster-wide","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"A compromised node's virt-handler service account can be abused cluster-wide","attack_vector":"An attacker who owns one node","remediation":"Upgrade KubeVirt; scope down the virt-handler service account","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-26484"],"status":"curated"},{"id":"CVE-2023-27487","cve":"CVE-2023-27487","aliases":[],"title":"Envoy: Client can forge the x-envoy-original-path header and bypass JWT checks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Client can forge the x-envoy-original-path header and bypass JWT checks","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; strip x-envoy headers at the edge","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27487"],"status":"curated"},{"id":"CVE-2023-31017","cve":"CVE-2023-31017","aliases":[],"title":"GPU Display Driver (Windows): Arbitrary write to privileged locations via reparse points","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Arbitrary write to privileged locations via reparse points","attack_vector":"Local low-priv user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-59"]},{"id":"CVE-2023-31027","cve":"CVE-2023-31027","aliases":[],"title":"GPU Display Driver (Windows): Local privesc during driver update","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Local privesc during driver update","attack_vector":"Local low-priv user on the host","remediation":"Upgrade to Oct-2023 driver branch; Windows hosts only","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-427"]},{"id":"CVE-2023-34330","cve":"CVE-2023-34330","aliases":[],"title":"AMI MegaRAC SPx (Dynamic Redfish Extension): Code injection executed via the Dynamic Redfish Extension interface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (Dynamic Redfish Extension)","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Code injection executed via the Dynamic Redfish Extension interface; BMC-level code execution","attack_vector":"Network / Redfish, authenticated-adjacent","remediation":"Same BMC flash cycle as CVE-2023-34329; ODM rebase required, cannot be mitigated in the host OS","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34330"],"status":"curated"},{"id":"CVE-2023-35944","cve":"CVE-2023-35944","aliases":[],"title":"Envoy: Mixed-case HTTP/2 schemes defeat case-sensitive internal scheme checks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"Mixed-case HTTP/2 schemes defeat case-sensitive internal scheme checks","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-35944"],"status":"curated"},{"id":"CVE-2023-46805","cve":"CVE-2023-46805","aliases":[],"title":"Ivanti Connect Secure: Web-component authentication bypass reaching restricted resources","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Connect Secure","year":"2023","cvss_score":8.2,"severity":"high","kev":true,"impact":"[KEV] Web-component authentication bypass reaching restricted resources; chained for unauth RCE","attack_vector":"Network (remote)","remediation":"Control-plane: patch and rebuild the appliance - the integrity checker is not sufficient","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46805"],"status":"curated"},{"id":"CVE-2023-49938","cve":"CVE-2023-49938","aliases":[],"title":"Slurm: A user can modify their extended group list used by sbcast and open files with unauthorized permissions","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":8.2,"severity":"high","kev":false,"impact":"A user can modify their extended group list used by sbcast and open files with unauthorized permissions","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49938"],"status":"curated"},{"id":"CVE-2023-6549","cve":"CVE-2023-6549","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Buffer overflow causing denial of service when configured as Gateway or AAA vserver","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":8.2,"severity":"high","kev":true,"impact":"[KEV] Buffer overflow causing denial of service when configured as Gateway or AAA vserver","attack_vector":"Network (remote)","remediation":"Control-plane: patch during the next maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6549"],"status":"curated"},{"id":"CVE-2024-0082","cve":"CVE-2024-0082","aliases":[],"title":"ChatRTX: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0082","https://github.com/NVIDIA/product-security/tree/main/2024/5532"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-269"]},{"id":"CVE-2024-0126","cve":"CVE-2024-0126","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A privileged attacker escalates through NVIDIA GPU software with a changed CVSS…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A privileged attacker escalates through NVIDIA GPU software with a changed CVSS scope, reaching code execution, data corruption and information disclosure beyond the component. The Virtual GPU Manager is in the affected list, so on a vGPU host the scope change means the hypervisor. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. NVIDIA's description is unusually vague here; treat the scope-changed 8.2 on a vGPU host as the worst case until you have your own detail.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5586. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0126","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0179","cve":"CVE-2024-0179","aliases":[],"title":"AmdCpmDisplayFeatureSMM - SMM callout (AMD-SB-7027): MULTI-TENANT ISOLATION: An SMM callout in the AmdCpmDisplayFeatureSMM driver lets ring-0 code overwrite…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AmdCpmDisplayFeatureSMM - SMM callout (AMD-SB-7027)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An SMM callout in the AmdCpmDisplayFeatureSMM driver lets ring-0 code overwrite SMRAM. SMM callouts are the classic UEFI escalation pattern: SMM code calls outward into memory the OS controls, so the OS supplies the code SMM then runs. Reported by Quarkslab, scored 8.2 with changed scope - the attacker crosses from the OS into the platform's most privileged context.","attack_vector":"Local, ring-0 on the host.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Patch alongside the sibling AmdPspP2CmboxV2 issue in the same bulletin - they ship together and leaving one open leaves the class open.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0179","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7027.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-1220","cve":"CVE-2024-1220","aliases":["MPSA-238975"],"title":"Moxa NPort W2150A / W2250A wireless device server: A remote attacker can crash or potentially gain code execution on the device server by sending a crafted…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Moxa NPort W2150A / W2250A wireless device server","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"A remote attacker can crash or potentially gain code execution on the device server by sending a crafted payload to its web management service — the built-in web server has a stack-based buffer overflow.","attack_vector":"Remote, over the network — no authentication mentioned as a prerequisite in the vendor advisory; reachability to the web service is sufficient to trigger the overflow.","remediation":"Firmware upgrade to the version in Moxa's MPSA-238975 advisory. Flash and reboot each unit; serial sessions on that device drop briefly during the update.","references":["https://www.moxa.com/en/support/product-support/security-advisory/mpsa-238975-nport-w2150a-w2250a-series-web-server-stack-based-buffer-overflow-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21924","cve":"CVE-2024-21924","aliases":[],"title":"AmdPlatformRasSspSmm - SMM callout (AMD-SB-7028): MULTI-TENANT ISOLATION: An SMM callout in the platform RAS SMM driver lets ring-0 code modify boot service…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AmdPlatformRasSspSmm - SMM callout (AMD-SB-7028)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An SMM callout in the platform RAS SMM driver lets ring-0 code modify boot service handlers and execute at SMM. Reported by Eclypsium. Worth noting the irony for a GPU operator: the affected driver is the platform's *reliability and error-reporting* code, so the component you depend on to tell you a node is unhealthy is the one handing over the platform.","attack_vector":"Local, ring-0 on the host.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21924","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7028.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-21925","cve":"CVE-2024-21925","aliases":[],"title":"AmdPspP2CmboxV2 - SMM input validation (AMD-SB-7027): MULTI-TENANT ISOLATION: Insufficient input validation in the AmdPspP2CmboxV2 SMM driver - the PSP mailbox…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AmdPspP2CmboxV2 - SMM input validation (AMD-SB-7027)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient input validation in the AmdPspP2CmboxV2 SMM driver - the PSP mailbox interface - lets ring-0 code overwrite SMRAM and execute at SMM. This one is notable for sitting on the PSP communication path, so a single bug hands the attacker both the SMM context and a channel to the secure processor.","attack_vector":"Local, ring-0 on the host.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21925","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7027.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-3446","cve":"CVE-2024-3446","aliases":[],"title":"QEMU (virtio): DMA reentrancy leads to double free across virtio devices - guest-to-host code execution in QEMU","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (virtio)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"DMA reentrancy leads to double free across virtio devices - guest-to-host code execution in QEMU","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart or live-migration to patched hosts. GPU-passthrough VMs cannot be live-migrated, so this is a scheduled tenant-visible drain","references":["https://access.redhat.com/security/cve/CVE-2024-3446"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-35199","cve":"CVE-2024-35199","aliases":[],"title":"TorchServe (gRPC 7070/7071): gRPC ports bound to all interfaces regardless of config","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe (gRPC 7070/7071)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"gRPC ports bound to all interfaces regardless of config","attack_vector":"Unauthenticated network from a co-tenant or the internet","remediation":"Patch and enforce bind-address at the pod/network policy layer, not in app config","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35199"],"status":"curated"},{"id":"CVE-2024-36129","cve":"CVE-2024-36129","aliases":[],"title":"OpenTelemetry Collector: Unsafe decompression","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenTelemetry Collector","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"Unsafe decompression -> unauthenticated attacker crashes the collector via excessive memory consumption","attack_vector":"Network (remote)","remediation":"Data-plane: upgrade both node agents and the gateway to 0.102.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36129"],"status":"curated"},{"id":"CVE-2024-39720","cve":"CVE-2024-39720","aliases":[],"title":"Ollama (GGUF parser): Malformed 4-byte GGUF file crashes the server (two HTTP requests)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF parser)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"Malformed 4-byte GGUF file crashes the server (two HTTP requests)","attack_vector":"Unauthenticated network upload of a crafted GGUF","remediation":"Upgrade past 0.1.46; part of Oligo's six-issue Ollama disclosure","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39720"],"status":"curated"},{"id":"CVE-2024-45067","cve":"CVE-2024-45067","aliases":[],"title":"Intel Gaudi software installer: The Gaudi software installer leaves files and directories with permissions that let a non-root local user…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi software installer","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"The Gaudi software installer leaves files and directories with permissions that let a non-root local user modify components that later run as root. That is a straight local root path on any node where the Gaudi stack was installed with the affected installer - and root on a Gaudi node means every tenant's job on that node.","attack_vector":"Any local authenticated user on a node that has the Gaudi stack installed. Node images built once and cloned across the fleet propagate the bad permissions everywhere.","remediation":"Upgrade the Gaudi software installer to 1.18 or later, and re-check permissions on nodes already provisioned - upgrading the package does not always repair permissions set by an earlier install. Rebuild the golden node image rather than patching in place. No firmware or BIOS component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45067","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01271.html"],"status":"curated"},{"id":"CVE-2024-7344","cve":"CVE-2024-7344","aliases":["Howyar Reloader"],"title":"Signed third-party UEFI application (Howyar Reloader and OEM rebrands): A Microsoft-signed UEFI recovery application loads an unsigned binary from a hardcoded path using its own…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Signed third-party UEFI application (Howyar Reloader and OEM rebrands)","year":"2024","cvss_score":8.2,"severity":"high","kev":false,"impact":"A Microsoft-signed UEFI recovery application loads an unsigned binary from a hardcoded path using its own loader instead of the firmware's verified LoadImage. Anyone holding a copy of that signed application can drop it on any Secure Boot machine that trusts the Microsoft third-party CA and boot arbitrary pre-OS code - the vulnerable system does not need the vendor's product installed. That is a universal, portable Secure Boot bypass usable to plant a bootkit on a rented GPU node.","attack_vector":"Write access to the EFI System Partition - local admin/root, prior bare-metal tenant, or BMC virtual media. No relationship to whether you use the affected recovery software.","remediation":"Apply the January 2025 UEFI revocation list (dbx) update that revokes the affected binaries - this is a firmware-level revocation, not a package update, so it lands via Windows Update, fwupd/LVFS, or an OEM BIOS update depending on the platform. Verify the revocation actually took on each node; dbx pushes silently no-op on some boards. Also audit the ESP for stray signed EFI applications you never installed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7344","https://www.welivesecurity.com/en/eset-research/under-cloak-uefi-secure-boot-introducing-cve-2024-7344/","https://kb.cert.org/vuls/id/529659"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-10451","cve":"CVE-2025-10451","aliases":["INSYDE-SA-2025009"],"title":"Insyde InsydeH2O (H19Int15CallbackSmm, combined DXE/SMM driver): An unchecked output buffer in a combined DXE/SMM driver lets an attacker write into SMRAM and reach arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (H19Int15CallbackSmm, combined DXE/SMM driver)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"An unchecked output buffer in a combined DXE/SMM driver lets an attacker write into SMRAM and reach arbitrary code execution in System Management Mode. The 2025 instalment of the same pattern Binarly and Insyde have been working through since 2021 - a driver that takes an address from the caller and writes to it without confirming the address is outside SMRAM. Ring -2 compromise: survives reinstall, defeats Secure Boot and attestation, invisible from the OS.","attack_vector":"Local admin/root on the host OS issuing the vulnerable SMI with a crafted output buffer address.","remediation":"OEM BIOS update carrying Insyde patch IB05690966. Affects Intel Ice Lake and Kaby Lake and AMD Picasso platforms; Insyde scopes the advisory by OEM feature version (HP feature version before 20C1) rather than by kernel version, so map it against your own OEM's BIOS release rather than against an Insyde kernel number. Firmware flash, reboot per node. No config workaround.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10451","https://www.insyde.com/security-pledge/SA-2025009/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-20093","cve":"CVE-2025-20093","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: A missing check for an exceptional condition in the 800-series Linux driver, scored…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A missing check for an exceptional condition in the 800-series Linux driver, scored high, reachable by an authenticated user. Same rollout unit as the rest of the ice 1.17.2 batch.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20093","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-22225","cve":"CVE-2025-22225","aliases":[],"title":"VMware ESXi: Arbitrary kernel write from the VMX process - sandbox escape completing the zero-day chain","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi","year":"2025","cvss_score":8.2,"severity":"high","kev":true,"impact":"Arbitrary kernel write from the VMX process - sandbox escape completing the zero-day chain [KEV]","attack_vector":"Tenant VM guest (chained after CVE-2025-22224)","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22225"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23309","cve":"CVE-2025-23309","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An uncontrolled DLL loading path in the display driver lets a local attacker plant a DLL and get code…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"An uncontrolled DLL loading path in the display driver lets a local attacker plant a DLL and get code execution at driver-install privilege - a classic path to SYSTEM on Windows GPU hosts. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5703. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23309","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-427"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-23342","cve":"CVE-2025-23342","aliases":[],"title":"NVIDIA NVDebug tool: The NVDebug diagnostic collector lets an actor gain access to a privileged account, reaching code execution…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVDebug tool","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"The NVDebug diagnostic collector lets an actor gain access to a privileged account, reaching code execution and privilege escalation with a changed scope. NVDebug runs on DGX/HGX platform hosts and is exactly the tool an operator runs as root when something is already wrong.","attack_vector":"Local, low privileges, with user interaction - typically an operator running NVDebug on a node where an attacker already has an unprivileged foothold.","remediation":"Update NVDebug per bulletin 5696. Cost: it is a standalone tool, so updating costs nothing operationally. The real control is to stop running old NVDebug bundles as root on nodes you already suspect are compromised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23342","https://github.com/NVIDIA/product-security/tree/main/2025/5696"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-522"]},{"id":"CVE-2025-25210","cve":"CVE-2025-25210","aliases":["INTEL-SA-01325","CVE-2025-22453","CVE-2025-35999","INTEL-SA-01412","CVE-2025-24918","INTEL-SA-01400"],"title":"Intel Server Firmware Update Utility (SysFwUpdt) and Server Configuration Utility before version 16.0.12: Improper input validation in the utility operators use to flash BIOS, BMC and ME firmware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Firmware Update Utility (SysFwUpdt) and Server Configuration Utility before version 16.0.12","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"Improper input validation in the utility operators use to flash BIOS, BMC and ME firmware, letting a privileged local user escalate; the companion issues add an incorrect-permission assignment and a link-following flaw in the same tool family. The irony is the point: the tool you run to remediate firmware is itself the escalation path into firmware. A tenant or a compromised operator account on a node can subvert the update process so that the node ends up running attacker-chosen firmware while your fleet records show a successful patch. That gives below-the-OS persistence that survives reimage and crosses tenant handoff, and it corrupts the evidence you would use to detect it.","attack_vector":"A privileged local user on the host where the utility runs - which includes tenants on bare-metal nodes if the tooling is left in the host image, and any compromised operator or automation account that drives firmware rollouts.","remediation":"Update SysFwUpdt and the Server Configuration Utility to 16.0.12 or later before running any further firmware campaigns. Separately, treat firmware update tooling as privileged infrastructure: do not leave it installed in tenant-facing host images, run firmware updates from the BMC/Redfish out-of-band path rather than from the host OS wherever the platform supports it, and verify post-update firmware versions and measurements out of band instead of trusting the tool's own success report.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25210","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01325.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01412.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33045","cve":"CVE-2025-33045","aliases":["AMI-SA-2025007"],"title":"AMI AptioV UEFI BIOS (SMM): A write-what-where primitive plus an information leak in System Management Mode - the most privileged…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SMM)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"A write-what-where primitive plus an information leak in System Management Mode - the most privileged execution context on an x86 server, above the kernel and invisible to the hypervisor. An attacker who wins here can write anywhere in memory including SMRAM, disable firmware protections, and install a bootkit that persists across OS reinstall and survives every host-level detection you run. Scope is changed, so the compromise reaches beyond the BIOS into the running system. On a shared GPU host this defeats the boundary that VM isolation and confidential-computing attestation both rest on.","attack_vector":"Local, requires high privileges - root or kernel-level code on the host OS. So the precondition is that an attacker already owns the operating system on a node; this is what they use to convert a revocable OS compromise into permanent firmware residency. In a bare-metal GPU rental model, the tenant themselves have that privilege by design.","remediation":"BIOS update to AptioV_5.040 or later. That is a firmware flash plus a full host reboot per node, which on a GPU fleet means draining the node - and if it is part of a multi-node training job, draining the whole ring. Rollout is gated on your server vendor picking up the AMI BKC and publishing a rebased BIOS for your specific SKU, which typically lags AMI's advisory by months. There is no config-only mitigation for an SMM bug. If you cannot patch, the compensating control is to stop treating host root as a containable compromise: rebuild affected nodes with a verified BIOS reflash rather than an OS reimage.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025007.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-33045"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-53652","cve":"CVE-2025-53652","aliases":[],"title":"Jenkins (Git Parameter plugin): Git parameter value is not validated against the offered choices","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Jenkins (Git Parameter plugin)","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"Git parameter value is not validated against the offered choices -> injection of arbitrary values into builds","attack_vector":"Network (remote)","remediation":"Control-plane: plugin upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53652"],"status":"curated"},{"id":"CVE-2025-56547","cve":"CVE-2025-56547","aliases":["PT-2025-19"],"title":"Broadcom NetXtreme-E network adapter firmware: A high-severity flaw in the firmware of Broadcom NetXtreme-E adapters (found in firmware 231.1.162.1 and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Broadcom NetXtreme-E network adapter firmware","year":"2025","cvss_score":8.2,"severity":"high","kev":false,"impact":"A high-severity flaw in the firmware of Broadcom NetXtreme-E adapters (found in firmware 231.1.162.1 and reported through Positive Technologies' responsible-disclosure process). NetXtreme-E is Broadcom's mainstream datacenter NIC line and the base for the Thor generation used in AI server designs. Adapter-firmware bugs matter more than their CVSS suggests: NIC firmware runs below the hypervisor and below the host OS, it persists across reinstall, and on many server designs the NIC also carries the NC-SI sideband to the BMC — so a compromised NIC is a candidate pivot into out-of-band management.","attack_vector":"Reachable through the adapter's firmware interfaces. Treat any party that can drive the NIC — a host-privileged tenant on bare metal, or network-side input depending on the affected path — as in scope until the vendor detail is public.","remediation":"Flash NetXtreme-E adapter firmware to the fixed release from Broadcom (or via your server OEM's firmware bundle). NIC firmware flash plus a cold power cycle, per node. On a GPU fleet, roll it into the same drain window you use for BMC and BIOS updates — doing NIC firmware as its own campaign is how it ends up never happening.","references":["https://global.ptsecurity.com/en/about/news/pt-expert-helped-patch-vulnerabilities-broadcom-network-adapter-firmware/","https://nvd.nist.gov/vuln/detail/CVE-2025-56547"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24188","cve":"CVE-2026-24188","aliases":[],"title":"NVIDIA TensorRT: An out-of-bounds write reachable from the network reaches data tampering, scored 8.2 with no privileges…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA TensorRT","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"An out-of-bounds write reachable from the network reaches data tampering, scored 8.2 with no privileges required. TensorRT sits inside almost every optimised inference deployment, so the affected surface is wide even where TensorRT is not the thing you deployed by name.","attack_vector":"Network, unauthenticated. Reached through whatever service embeds the TensorRT runtime and passes it externally-influenced input.","remediation":"Update TensorRT per bulletin 5836 and rebuild every inference image that links it - including Triton images with the TensorRT backend. Cost: image rebuild plus rolling restart; inventory work is the expensive part because TensorRT is usually a transitive dependency.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24188","https://github.com/NVIDIA/product-security/tree/main/2026/5836"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:L","cwe":["CWE-787"],"fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-24189","cve":"CVE-2026-24189","aliases":[],"title":"NVIDIA CUDA-Q: Info disclosure / code exec (OOB read in circuit compilation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA-Q","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Info disclosure / code exec (OOB read in circuit compilation)","attack_vector":"Malicious user program","remediation":"Bump CUDA-Q; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24189","https://github.com/NVIDIA/product-security/tree/main/2026/5820"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-125"]},{"id":"CVE-2026-24253","cve":"CVE-2026-24253","aliases":[],"title":"NVIDIA Dynamo: RCE via buffer overflow in tensor shape validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"RCE via buffer overflow in tensor shape validation","attack_vector":"Any inference client","remediation":"Bump Dynamo; redeploy the serving stack","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24253","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"]},{"id":"CVE-2026-33748","cve":"CVE-2026-33748","aliases":[],"title":"BuildKit: Insufficient validation of git URL fragment subdir allows access to files outside the intended checkout","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Insufficient validation of git URL fragment subdir allows access to files outside the intended checkout","attack_vector":"Anyone who can submit a build","remediation":"Upgrade BuildKit to 0.28.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33748"],"status":"curated"},{"id":"CVE-2026-41326","cve":"CVE-2026-41326","aliases":[],"title":"Kata Containers: Oversight in the CopyFile policy from v3.4.0 to v3.28.0","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Oversight in the CopyFile policy from v3.4.0 to v3.28.0","attack_vector":"Any tenant workload under Kata","remediation":"Upgrade Kata","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41326"],"status":"curated"},{"id":"CVE-2026-47483","cve":"CVE-2026-47483","aliases":[],"title":"DCGM: DoS of the GPU telemetry/health daemon (resource exhaustion)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DCGM","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"DoS of the GPU telemetry/health daemon (resource exhaustion)","attack_vector":"Any tenant able to reach the DCGM socket/endpoint on the node","remediation":"Bump DCGM / dcgm-exporter; upgrade GPU Operator chart; restart the daemonset, no tenant eviction needed","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47483","https://github.com/NVIDIA/product-security/tree/main/2026/5857"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-770"]},{"id":"CVE-2026-47623","cve":"CVE-2026-47623","aliases":[],"title":"NVIDIA Dynamo: Deserialization of untrusted data in Dynamo reaches denial of service and data tampering from an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Deserialization of untrusted data in Dynamo reaches denial of service and data tampering from an unauthenticated network caller, scored 8.2. Dynamo is NVIDIA's disaggregated serving framework, so this sits on the request path of a production inference tier.","attack_vector":"Network, unauthenticated, no user interaction. Any client that can submit to the Dynamo endpoint.","remediation":"Upgrade Dynamo per bulletin 5842 and roll the deployment. Cost: rolling restart of the serving tier; no driver or firmware change. Verify the endpoint is not exposed beyond the mesh.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47623","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-502"],"fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-53489","cve":"CVE-2026-53489","aliases":[],"title":"containerd: CRI restores container.log from a checkpoint image without validating symlinks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"CRI restores container.log from a checkpoint image without validating symlinks; arbitrary host file write","attack_vector":"Malicious checkpoint image","remediation":"Rolling containerd upgrade with node drain; disable CRI checkpoint/restore for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53489"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-5817","cve":"CVE-2026-5817","aliases":[],"title":"Docker Model Runner (vllm-metal backend): `trust_remote_code=True` set unconditionally, no sandbox","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Docker Model Runner (vllm-metal backend)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"`trust_remote_code=True` set unconditionally, no sandbox → tokenizer code executes","attack_vector":"Customer-supplied model repo","remediation":"Rebuild/patch; no operator config can disable it","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5817"],"status":"curated"},{"id":"CVE-2026-6484","cve":"CVE-2026-6484","aliases":["INSYDE-SA-2026003"],"title":"Insyde InsydeH2O (unverified firmware volume in the boot chain): Certain firmware volumes are executed without being verified, so the verified-boot chain has a hole in it: an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (unverified firmware volume in the boot chain)","year":"2026","cvss_score":8.2,"severity":"high","kev":false,"impact":"Certain firmware volumes are executed without being verified, so the verified-boot chain has a hole in it: an attacker who can write to the unverified FV gets arbitrary code execution in firmware and the platform's own boot integrity check does not object. This is the root-of-trust failure rather than a memory-safety bug - the whole value of measured and verified boot on a GPU node is that unauthorised firmware cannot run, and here it can.","attack_vector":"An attacker able to modify the affected firmware volume - via SPI write access, a malicious capsule, or an earlier compromise with firmware-write privilege. Then any boot.","remediation":"OEM BIOS update. Insyde ships fixes per Intel platform: Arrow Lake H/U 05.56.17.0022, Arrow Lake S/HX 05.56.17.0037, Raptor Lake 05.47.24.0058 (mobile) and 05.47.24.0057 (server/embedded), Alder Lake 05.47.24.2057, Meteor Lake 05.56.07.0022, Elkhart Lake 05.48.17.0030 - so check your exact silicon generation, since several platforms are listed unaffected. Advisory dated 2026-08-12, meaning OEM images are only just starting to appear; expect the Dell/HPE/Lenovo/Supermicro rebase to trail by months. Firmware flash, reboot per node. No config workaround. Meanwhile, verify SPI flash write protection is actually enforced (BIOS Lock Enable, protected range registers) so the unverified FV is not writable in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-6484","https://www.insyde.com/security-pledge/SA-2026003/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-15372","cve":"CVE-2018-15372","aliases":[],"title":"Cisco IOS XE MACsec Key Agreement (MKA over EAP-TLS): TENANT ISOLATION: a logic error in MKA over EAP-TLS lets an unauthenticated adjacent attacker bypass…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco IOS XE MACsec Key Agreement (MKA over EAP-TLS)","year":"2018","cvss_score":8.1,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a logic error in MKA over EAP-TLS lets an unauthenticated adjacent attacker bypass authentication and pass traffic through a Layer 3 interface. MACsec here is doing double duty as link encryption and as port admission control, and both fail — an unauthenticated device on the wire gets its traffic forwarded as though it had authenticated.","attack_vector":"Unauthenticated attacker with adjacent (same-link) access to an interface configured for MKA with EAP-TLS.","remediation":"Software upgrade plus device reload. Do not treat MACsec/MKA as your only port-admission control — pair it with per-port VLAN pinning and MAC allowlisting, which are live config changes that keep an unauthenticated device from reaching anything useful even when the MKA check fails.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15372"],"status":"curated"},{"id":"CVE-2019-11247","cve":"CVE-2019-11247","aliases":[],"title":"Kubernetes (kube-apiserver): Cluster-scoped custom resources reachable through namespaced requests, so namespace-scoped RBAC grants…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2019","cvss_score":8.1,"severity":"high","kev":false,"impact":"Cluster-scoped custom resources reachable through namespaced requests, so namespace-scoped RBAC grants cluster-wide CR access","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain. Audit CRD RBAC afterwards","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11247"],"status":"curated"},{"id":"CVE-2020-12693","cve":"CVE-2020-12693","aliases":[],"title":"Slurm: Race condition in message aggregation allows launching a process as another user","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2020","cvss_score":8.1,"severity":"high","kev":false,"impact":"Race condition in message aggregation allows launching a process as another user","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12693"],"status":"curated"},{"id":"CVE-2021-21540","cve":"CVE-2021-21540","aliases":[],"title":"Dell iDRAC9: Stack overflow overwriting iDRAC configuration via oversized payloads","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Stack overflow overwriting iDRAC configuration via oversized payloads","attack_vector":"Network, authenticated","remediation":"iDRAC firmware update","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21540"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-23214","cve":"CVE-2021-23214","aliases":[],"title":"PostgreSQL: With cert/trust+clientcert auth, a MITM can inject arbitrary SQL at connection setup","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"With cert/trust+clientcert auth, a MITM can inject arbitrary SQL at connection setup","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; enforce full-verify TLS between control-plane services","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23214"],"status":"curated"},{"id":"CVE-2021-29492","cve":"CVE-2021-29492","aliases":[],"title":"Envoy: Escaped slash sequences %2F and %5C not decoded","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Escaped slash sequences %2F and %5C not decoded; path-based authorization bypass","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; enable path normalization and reject encoded slashes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-29492"],"status":"curated"},{"id":"CVE-2021-3139","cve":"CVE-2021-3139","aliases":[],"title":"tcmu-runner 1.3.x - 1.5.2 (userspace backstore handler for the Linux LIO target, used by Ceph iSCSI gateways and other software targets): xcopy_locate_udev() does not enforce transport-layer…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"tcmu-runner 1.3.x - 1.5.2","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"xcopy_locate_udev() does not enforce transport-layer restrictions, so an XCOPY (extended copy) request can name a source or destination by path traversal. An attacker who has been legitimately given one iSCSI LUN can therefore read or write files outside it - including other tenants' LUN backing files on the same target. This is the cleanest cross-tenant storage break in this set: no memory corruption, no crash, just a normal SCSI command that the target happily executes against the wrong tenant's data. It is the same mistake as CVE-2020-28374 in a different code path, so a target that patched only the earlier one is still exposed.","attack_vector":"An authenticated tenant with a single provisioned LUN on the affected target, issuing a crafted XCOPY over the normal iSCSI data path. No privilege escalation on the target and no access to the management network required.","remediation":"Upgrade tcmu-runner past 1.5.2 - distro package update plus a restart of tcmu-runner, which briefly stalls I/O on the LUNs it backs but does not require a kernel reboot. Check what actually ships tcmu-runner in your stack: Ceph iSCSI gateway deployments and several appliance images vendor it, so the fixed version may need to come from the appliance vendor rather than the distro. Verify the fix covers both this and CVE-2020-28374; patching one code path was the original mistake. Until patched, disable XCOPY/ODX support on the target if your backstore allows it.","references":["https://www.openwall.com/lists/oss-security/2021/01/12/12","https://bugzilla.suse.com/show_bug.cgi?id=1178372","https://nvd.nist.gov/vuln/detail/CVE-2021-3139"],"status":"curated"},{"id":"CVE-2021-39156","cve":"CVE-2021-39156","aliases":[],"title":"Istio: Host header with a port bypasses AuthorizationPolicy host matching","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Host header with a port bypasses AuthorizationPolicy host matching","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39156"],"status":"curated"},{"id":"CVE-2021-42387","cve":"CVE-2021-42387","aliases":[],"title":"ClickHouse: Attacker-controlled offset in the LZ4 codec","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ClickHouse","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Attacker-controlled offset in the LZ4 codec -> heap out-of-bounds read via a malicious query","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the telemetry/usage-metering cluster","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-42387"],"status":"curated"},{"id":"CVE-2021-42388","cve":"CVE-2021-42388","aliases":[],"title":"ClickHouse: Second heap out-of-bounds read in LZ4::decompressImpl reachable from a client query","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ClickHouse","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"Second heap out-of-bounds read in LZ4::decompressImpl reachable from a client query","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; require auth on the native port","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-42388"],"status":"curated"},{"id":"CVE-2022-28733","cve":"CVE-2022-28733","aliases":[],"title":"GRUB2 (net/ip IPv4 reassembly): Integer underflow in grub_net_recv_ip4_packets from a crafted IP packet. This one matters far more than the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (net/ip IPv4 reassembly)","year":"2022","cvss_score":8.1,"severity":"high","kev":false,"impact":"Integer underflow in grub_net_recv_ip4_packets from a crafted IP packet. This one matters far more than the filesystem bugs for a GPU cloud, because it is reachable over the network during PXE boot - an attacker who can answer on the provisioning VLAN owns the node before any OS, tenant, or agent exists.","attack_vector":"Anyone who can put packets on the provisioning/PXE network while a node is netbooting. No credentials, no prior access to the node.","remediation":"grub2 package update + reboot, and update the netboot GRUB image you actually serve - patching running nodes does nothing if the TFTP/HTTP-served binary is stale. Compensating control: put provisioning on an isolated L2 segment with DHCP snooping, and do not let tenant workloads share it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28733","https://access.redhat.com/security/cve/CVE-2022-28733"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-42272","cve":"CVE-2022-42272","aliases":[],"title":"DGX servers BMC: RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":8.1,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42272","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"]},{"id":"CVE-2022-42273","cve":"CVE-2022-42273","aliases":[],"title":"DGX servers BMC: RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":8.1,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42273","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-120"]},{"id":"CVE-2023-25409","cve":"CVE-2023-25409","aliases":[],"title":"ATEN PE8108 switched PDU: TENANT ISOLATION: a restricted (non-admin) user account on the PDU's web interface can control outlets…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATEN PE8108 switched PDU","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a restricted (non-admin) user account on the PDU's web interface can control outlets belonging to other users — meaning one tenant sharing this PDU with others can power-cycle or power-off outlets feeding another tenant's equipment, not just their own.","attack_vector":"Requires only a low-privileged, restricted user account on the PDU — no admin credentials needed to reach outlets outside the account's assigned scope.","remediation":"Firmware upgrade from ATEN to a version that enforces per-outlet authorization correctly. Flash each PDU; since this is the power path for the racks it feeds, coordinate the maintenance window with anyone whose equipment is on that PDU.","references":["https://www.pentagrid.ch/en/blog/multiple-vulnerabilities-in-aten-PE8108-power-distribution-unit"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25552","cve":"CVE-2023-25552","aliases":["SEVD-2023-101-04"],"title":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - Device File Transfer settings: Missing authorisation on the Device File Transfer settings lets an attacker view, change or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - Device File Transfer settings","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Missing authorisation on the Device File Transfer settings lets an attacker view, change or delete content and invoke functions they should not have. Device File Transfer is how DCE pushes firmware and config to managed power and cooling devices - so control of it is control of what firmware lands on your UPS and PDU fleet.","attack_vector":"Remote access to the DCE endpoints, without the authorisation the function should require.","remediation":"Upgrade past V7.9.2. Independently, gate firmware distribution: a DCIM system that can silently push firmware to power devices is a single point of catastrophic failure, and it deserves change control rather than trust.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25552"],"status":"curated"},{"id":"CVE-2023-27493","cve":"CVE-2023-27493","aliases":[],"title":"Envoy: Request properties are not escaped when generating request headers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Request properties are not escaped when generating request headers; header injection into upstreams","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27493"],"status":"curated"},{"id":"CVE-2023-31424","cve":"CVE-2023-31424","aliases":[],"title":"Brocade SANnav Management Portal web interface, before v2.3.0 and v2.2.2a: Remote unauthenticated users can bypass web authentication and authorization on the SANnav portal. That is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade SANnav Management Portal web interface, before v2.3.0 and v2.2.2a","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Remote unauthenticated users can bypass web authentication and authorization on the SANnav portal. That is the front door to the whole FC management estate - fabric inventory, zoning pushes, switch credentials, firmware distribution. Chained with the zone-management SQL injection above it turns a fully unauthenticated network position into control of tenant isolation across every managed fabric.","attack_vector":"Any host with network reachability to the SANnav web interface. No credentials at all.","remediation":"Upgrade SANnav to 2.3.0 or 2.2.2a. Management-plane upgrade only - no fabric or array disruption. Because the pre-fix window allowed unauthenticated access, also rotate stored switch credentials and diff the live zonesets against your intended configuration rather than assuming the upgrade closes the incident.","references":["https://support.broadcom.com/web/ecx/support-content-notification/-/external/content/SecurityAdvisories/0/22507","https://security.netapp.com/advisory/ntap-20240229-0004/","https://nvd.nist.gov/vuln/detail/CVE-2023-31424"],"status":"curated"},{"id":"CVE-2023-34336","cve":"CVE-2023-34336","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (IPMI handler): Buffer overflow in the BMC's IPMI message handler leading to code execution or privilege escalation inside…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (IPMI handler)","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Buffer overflow in the BMC's IPMI message handler leading to code execution or privilege escalation inside the BMC. This is the pre-Redfish legacy protocol that almost every fleet still leaves enabled for ipmitool-based power control and sensor scraping, so the exposed surface is usually larger than operators assume. A win here means the attacker controls power, console and firmware update paths for the node.","attack_vector":"Network-reachable IPMI service, no credentials required per AMI's own CVSS vector, but high attack complexity. Any host that can send IPMI RMCP+ traffic to UDP/623 on the BMC is in range - which is every host on the management VLAN, and in badly-built estates any host that can route there.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, gated on ODM rebase. Unlike the network-stack bugs in this cluster there IS a meaningful config-only mitigation here: disable IPMI-over-LAN entirely and drive power/sensors through Redfish instead. That is a config change on the BMC, no reboot, no flash - but it breaks any ipmitool-based tooling in your provisioning and monitoring stack, so cost it as a tooling migration.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34336"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-4606","cve":"CVE-2023-4606","aliases":["LEN-140960"],"title":"Lenovo XClarity Controller (XCC) - user account API: A read-only XCC user can change any other user's password through a crafted API call. That is a direct path…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - user account API","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"A read-only XCC user can change any other user's password through a crafted API call. That is a direct path from the least-privileged BMC account you hand out to full administrative control of the service processor: change the admin's password, log in as admin, and you have power control, remote media, console and firmware on the node. It also locks out the legitimate administrator, which turns a quiet compromise into a visible outage. Affects ThinkSystem V2 and V3 servers - the generations that carry the SR670 V2 / SR675 V3 / SR685a GPU platforms. V1 servers are not affected.","attack_vector":"An authenticated XCC account holding only read-only permission - typically a monitoring collector, a DCIM integration, or an account issued to remote hands. Reachable over the out-of-band management VLAN.","remediation":"Flash XCC to the per-model version listed in Lenovo's advisory - out-of-band, per-node, no host reboot and no drain of running jobs. Model-specific version floors mean you cannot use one target build across a mixed fleet; pull the table from LEN-140960 and drive the campaign per SKU. Config-only mitigation in the meantime: audit and prune read-only XCC accounts, since 'read-only' provides no protection against this bug.","references":["https://support.lenovo.com/us/en/product_security/LEN-140960","https://nvd.nist.gov/vuln/detail/CVE-2023-4606"],"status":"curated"},{"id":"CVE-2023-6572","cve":"CVE-2023-6572","aliases":[],"title":"Gradio: Command injection","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2023","cvss_score":8.1,"severity":"high","kev":false,"impact":"Command injection","attack_vector":"Network user of the demo","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6572"],"status":"curated"},{"id":"CVE-2024-0114","cve":"CVE-2024-0114","aliases":[],"title":"Hopper HGX 8-GPU (HMC/firmware): Improper input validation in baseboard firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Hopper HGX 8-GPU (HMC/firmware)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Improper input validation in baseboard firmware -> code exec / tampering across the GPU baseboard","attack_vector":"Local privileged host access / mgmt path","remediation":"Flash HGX baseboard firmware out-of-band; full node drain and power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0114","https://github.com/NVIDIA/product-security/tree/main/2025/5561"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:L/I:H/A:H","cwe":["CWE-1244"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-10220","cve":"CVE-2024-10220","aliases":[],"title":"Kubernetes (kubelet): Arbitrary command execution on the node via a gitRepo volume","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Arbitrary command execution on the node via a gitRepo volume","attack_vector":"Cluster user able to create a pod with a gitRepo volume","remediation":"Rolling kubelet upgrade with node drain; block the deprecated gitRepo volume type by admission policy","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-28088","cve":"CVE-2024-28088","aliases":[],"title":"LangChain: Directory traversal via the template path parameter","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Directory traversal via the template path parameter","attack_vector":"Attacker-controlled final path segment","remediation":"Upgrade past 0.1.10","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28088"],"status":"curated"},{"id":"CVE-2024-28233","cve":"CVE-2024-28233","aliases":[],"title":"JupyterHub: Malicious subdomain tricks a user","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterHub","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Malicious subdomain tricks a user → session takeover","attack_vector":"User visiting an attacker-controlled subdomain","remediation":"Upgrade; per-user subdomain isolation is the mitigation and the vuln","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28233"],"status":"curated"},{"id":"CVE-2024-36623","cve":"CVE-2024-36623","aliases":[],"title":"Docker / moby: Race condition in the streamformatter package causing data corruption or daemon crash","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Race condition in the streamformatter package causing data corruption or daemon crash","attack_vector":"Any tenant workload driving concurrent daemon streams","remediation":"Upgrade moby; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36623"],"status":"curated"},{"id":"CVE-2024-4888","cve":"CVE-2024-4888","aliases":[],"title":"LiteLLM: Arbitrary file deletion via `/audio/transcriptions`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Arbitrary file deletion via `/audio/transcriptions`","attack_vector":"Network user of the proxy","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4888"],"status":"curated"},{"id":"CVE-2024-5154","cve":"CVE-2024-5154","aliases":[],"title":"CRI-O: Malicious container creates a symlink via directory traversal and gets arbitrary host read/write","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"Malicious container creates a symlink via directory traversal and gets arbitrary host read/write","attack_vector":"Any tenant workload / malicious image","remediation":"Upgrade CRI-O on all nodes; drain and recreate pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5154"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-6387","cve":"CVE-2024-6387","aliases":[],"title":"OpenSSH (sshd): regreSSHion: signal-handler race in sshd giving unauthenticated remote root on glibc Linux","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (sshd)","year":"2024","cvss_score":8.1,"severity":"high","kev":false,"impact":"regreSSHion: signal-handler race in sshd giving unauthenticated remote root on glibc Linux","attack_vector":"Unauthenticated network","remediation":"Package update + sshd restart; no reboot. Interim mitigation `LoginGraceTime 0` (costs DoS resilience). Highest-priority item for any tenant-reachable bastion or management SSH endpoint","references":["https://access.redhat.com/security/cve/CVE-2024-6387"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-1094","cve":"CVE-2025-1094","aliases":[],"title":"PostgreSQL (libpq): Improper quoting in PQescape*","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"PostgreSQL (libpq)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Improper quoting in PQescape* -> SQL injection, chainable to shell via psql \\!","attack_vector":"Network (remote)","remediation":"Control-plane: patch metadata/billing DB servers and all libpq clients","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1094"],"status":"curated"},{"id":"CVE-2025-14279","cve":"CVE-2025-14279","aliases":[],"title":"MLflow (REST API): DNS rebinding — no Origin header validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (REST API)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"DNS rebinding — no Origin header validation","attack_vector":"Operator's browser visiting a malicious page while on the cluster network","remediation":"Upgrade past 3.4.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14279"],"status":"curated"},{"id":"CVE-2025-23318","cve":"CVE-2025-23318","aliases":[],"title":"NVIDIA Triton (Python backend): Out-of-bounds write in the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (Python backend)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Out-of-bounds write in the Python backend","attack_vector":"Unauthenticated network","remediation":"Patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23318"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-805"]},{"id":"CVE-2025-23319","cve":"CVE-2025-23319","aliases":[],"title":"NVIDIA Triton (Python backend shared memory): Out-of-bounds write in the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (Python backend shared memory)","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Out-of-bounds write in the Python backend","attack_vector":"Unauthenticated network","remediation":"Patch; the Python backend's shared-memory region is the pivot in the Wiz chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23319"],"status":"curated","fleet":{"ubiquity":"Very common - Triton is the default multi-model serving layer inside NIM and in many neocloud managed-inference products","remediation_pain":"`daemon-restart` (upgrade to 25.07, restart the serving process) - no host reboot, but it is a customer-visible endpoint outage","pain_class":"node-reboot","why_fleet_wide":"Unauthenticated remote chain: an oversized request leaks the internal shared-memory key, then an OOB write in the Python backend yields full RCE as the Triton process - i.e. theft of every model and every inference request on that server"},"cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-805"]},{"id":"CVE-2025-3935","cve":"CVE-2025-3935","aliases":[],"title":"ConnectWise ScreenConnect: ViewState code injection","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"ConnectWise ScreenConnect","year":"2025","cvss_score":8.1,"severity":"high","kev":true,"impact":"[KEV] ViewState code injection -> RCE once ASP.NET machine keys are obtained","attack_vector":"Network (remote)","remediation":"Control-plane: patch + rotate ASP.NET machine keys","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3935"],"status":"curated"},{"id":"CVE-2025-71340","cve":"CVE-2025-71340","aliases":[],"title":"picklescan: Misses `idlelib.pyshell.ModifiedInterpreter.runcode` gadget","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Misses `idlelib.pyshell.ModifiedInterpreter.runcode` gadget","attack_vector":"Customer-supplied pickle","remediation":"Upgrade; the denylist approach is structurally incomplete","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71340"],"status":"curated"},{"id":"CVE-2025-71342","cve":"CVE-2025-71342","aliases":[],"title":"picklescan: Misses `idlelib.run.Executive.runcode` gadget","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":8.1,"severity":"high","kev":false,"impact":"Misses `idlelib.run.Executive.runcode` gadget","attack_vector":"Customer-supplied pickle","remediation":"Upgrade to 0.0.30+; treat gadget denylists as best-effort only","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71342"],"status":"curated"},{"id":"CVE-2026-11816","cve":"CVE-2026-11816","aliases":[],"title":"Keras (archive extraction utils): Path traversal in `keras/src/utils/file_utils.py`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (archive extraction utils)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Path traversal in `keras/src/utils/file_utils.py`","attack_vector":"Customer-supplied archive","remediation":"Upgrade to 3.14.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-11816"],"status":"curated"},{"id":"CVE-2026-18577","cve":"CVE-2026-18577","aliases":[],"title":"N-able N-central: Incomplete patch for CVE-2026-18556","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2026","cvss_score":8.1,"severity":"high","kev":true,"impact":"[KEV] Incomplete patch for CVE-2026-18556 -> auth bypass and full account takeover","attack_vector":"Network (remote)","remediation":"Control-plane: apply the follow-up patch; audit N-central accounts created since","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-18577"],"status":"curated"},{"id":"CVE-2026-24218","cve":"CVE-2026-24218","aliases":[],"title":"DGX Spark: Hardcoded credentials in config","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX Spark","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Hardcoded credentials in config -> unauthorized access","attack_vector":"Local or network attacker with config access","remediation":"Update DGX Spark software; rotate all affected credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24218","https://github.com/NVIDIA/product-security/tree/main/2026/5835"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-321"]},{"id":"CVE-2026-43629","cve":"CVE-2026-43629","aliases":[],"title":"llama-server (KV cache state restore): Heap buffer overflow in `state_read_data`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-server (KV cache state restore)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Heap buffer overflow in `state_read_data`","attack_vector":"Attacker-supplied KV/session state file restored by the server","remediation":"Rebuild; session-state restore is an untrusted-input parser","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43629"],"status":"curated"},{"id":"CVE-2026-43632","cve":"CVE-2026-43632","aliases":[],"title":"llama-server (tokenization endpoints): Use-after-free across six tokenization endpoints","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-server (tokenization endpoints)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"Use-after-free across six tokenization endpoints","attack_vector":"Unauthenticated network to llama-server","remediation":"Rebuild past b9060","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43632"],"status":"curated"},{"id":"CVE-2026-49121","cve":"CVE-2026-49121","aliases":["AITER pickle RCE"],"title":"AI Tensor Engine for ROCm (AITER) - MessageQueue.recv() in shm_broadcast.py: MULTI-TENANT ISOLATION: AITER's MessageQueue.recv() deserialises whatever arrives on a ZeroMQ SUB socket with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"AI Tensor Engine for ROCm (AITER) - MessageQueue.recv() in shm_broadcast.py","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: AITER's MessageQueue.recv() deserialises whatever arrives on a ZeroMQ SUB socket with Python pickle, so anyone who can reach that socket gets **unauthenticated remote code execution** as the process running the ROCm inference or training job. This is the highest-value AMD-stack finding for an AI operator in the whole set: it needs no local access, no privilege and no GPU device handle - just network reach to a port that distributed ROCm jobs open between workers. If your tensor-parallel or pipeline-parallel workers talk over an unauthenticated ZMQ fabric on a shared cluster network, another tenant on that network owns your jobs.","attack_vector":"Network. Unauthenticated. The attacker needs only IP reachability to the ZMQ SUB socket that AITER opens for shared-memory broadcast between distributed workers. On a flat cluster network - which is the norm for RDMA/RoCE training fabrics - that means any other tenant on the fabric. Affects AITER through 0.1.14.","remediation":"Upgrade AITER past 0.1.14. Independently of the patch, fix the exposure: bind the ZMQ sockets to localhost or the job's private network namespace rather than 0.0.0.0, put distributed-training traffic on a per-job network segment, and enforce that with NetworkPolicy or equivalent so worker-to-worker ports are not reachable across tenants. No driver reload, no reboot, no firmware - this is an application and network-segmentation fix, which also means your firmware and kernel patch tooling will never surface it. Audit the rest of your stack for the same pattern: pickle-over-socket is endemic in distributed ML frameworks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49121","https://github.com/ROCm/aiter"],"status":"curated"},{"id":"CVE-2026-63925","cve":"CVE-2026-63925","aliases":[],"title":"Linux MACsec (replay protection at XPN lower-PN wrap): TENANT ISOLATION: MACsec replay protection fails at the extended-packet-number lower-PN wrap. When the packet…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux MACsec (replay protection at XPN lower-PN wrap)","year":"2026","cvss_score":8.1,"severity":"high","kev":false,"impact":"TENANT ISOLATION: MACsec replay protection fails at the extended-packet-number lower-PN wrap. When the packet number is U32_MAX the increment overflows to zero and neither replay branch fires, so `next_pn_halves` is never advanced — an attacker who captured legitimate ciphertext can replay it and have it accepted. Replay protection is the property that stops a passive observer from becoming an active injector on an encrypted link; losing it turns a tap into a traffic-injection capability on links that carry multiple tenants.","attack_vector":"An attacker who can capture and re-transmit frames on a MACsec-protected link, timed to the PN wrap. Passive tap plus injection capability, no keys required.","remediation":"Kernel upgrade plus host reboot on any node terminating software MACsec (switch NOSes based on Linux included, where it arrives as a NOS image update plus reload). No config workaround — you cannot turn replay protection back on if the check itself is broken. Companion Linux MACsec issues in the same window: CVE-2026-72019, CVE-2022-48720.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63925","https://nvd.nist.gov/vuln/detail/CVE-2026-72019"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2021-004-infiniband-roce-memory-protectio","cve":null,"aliases":["ReDMArk","RDMA rkey brute-forcing","unauthorized RDMA memory access","memory window abuse"],"title":"InfiniBand / RoCE memory protection - memory region rkey/lkey namespace and protection domains: TENANT ISOLATION: The only thing standing between a remote peer and a registered memory region is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand / RoCE memory protection - memory region rkey/lkey namespace and protection domains","year":"2021","cvss_score":8.1,"severity":"high","kev":false,"impact":"TENANT ISOLATION: The only thing standing between a remote peer and a registered memory region is the 32-bit rkey, and ReDMArk found that real RNIC firmware generates rkeys in a small, largely sequential, and therefore guessable space. Combined with the connection-injection weakness, an attacker can issue RDMA READ or WRITE against another tenant's memory region without ever having been granted access to it. Applications make this dramatically worse in practice by registering one huge region with IBV_ACCESS_REMOTE_WRITE covering far more than the buffers actually meant to be shared - a common shortcut in RDMA key-value stores, parameter servers, and disaggregated-memory layers. On a GPU node the registered region frequently includes GPU memory reachable through GPUDirect, so a guessed rkey reads model state directly out of HBM.","attack_vector":"From any node able to establish or inject into a queue pair on the target RNIC, the attacker sweeps rkey values in RDMA READ requests and watches for a completion instead of an error. Because the RNIC services the read entirely in hardware, the victim's CPU never runs and nothing is logged on the victim host - the sweep is silent and can run at line rate. Successful reads return the contents of whatever the victim registered; successful writes corrupt it. Memory windows (type 1/2) bind more narrowly but are seldom used, and applications that reuse a single protection domain across tenant contexts collapse the boundary entirely.","remediation":"No vendor patch. Application and config work: register the smallest possible regions, never grant REMOTE_WRITE where REMOTE_READ suffices, use a separate protection domain per tenant/connection, and prefer type-2 memory windows with short lifetimes so a guessed key expires. Where the RNIC firmware supports stronger key generation, apply the vendor firmware update (firmware flash, rolling per-node, ~5-15 min per host plus a reboot on most ConnectX generations). Fabric-level containment - one P_Key or VLAN per tenant - limits who can reach the RNIC to try keys at all and is a switch config change. Budget this as an application-rearchitecture item, not a maintenance window.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf","https://arxiv.org/abs/1903.09355"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2021-21505","cve":"CVE-2021-21505","aliases":["DSA-2021-020"],"title":"Dell EMC Integrated System for Microsoft Azure Stack Hub (undocumented iDRAC account): Dell shipped these integrated racks with an undocumented iDRAC account whose credentials are the same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell EMC Integrated System for Microsoft Azure Stack Hub (undocumented iDRAC account)","year":"2021","cvss_score":8,"severity":"high","kev":false,"impact":"Dell shipped these integrated racks with an undocumented iDRAC account whose credentials are the same everywhere. Anyone who learns them owns the BMC on every node in the system - power control, Virtual Media boot of an attacker image, KVM into the console, and a persistent foothold under the hypervisor. A shared default credential is the worst shape of this class of bug because it does not need an exploit, scales across the whole install base at once, and is invisible to a vulnerability scan that only checks firmware versions. Affects builds 1906 through 2011.","attack_vector":"Anything routable to the iDRAC addresses on the out-of-band management VLAN, using credentials that are effectively public once disclosed. No exploit, no prior foothold, no privilege escalation step.","remediation":"Apply the Dell update package that brings the system to build 2102 or later. Because the root cause is an account rather than a code defect, also verify by hand afterward that the account is gone from every node's iDRAC user list, and rotate any other shared BMC credentials while you are in there. Firmware/update-package rollout is out-of-band; the account audit is config-only and should be done immediately regardless of the patch schedule.","references":["https://www.dell.com/support/kbdoc/en-us/000186008/dsa-2021-020-dell-emc-integrated-system-for-microsoft-azure-stack-hub-security-update-for-an-idrac-undocumented-account-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2021-21505"],"status":"curated"},{"id":"CVE-2021-23279","cve":"CVE-2021-23279","aliases":[],"title":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - meta_driver_srv.js: Unauthenticated arbitrary file deletion on the IPM server. Less glamorous than the RCEs but operationally…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton Intelligent Power Manager (IPM) prior to 1.69 - meta_driver_srv.js","year":"2021","cvss_score":8,"severity":"high","kev":false,"impact":"Unauthenticated arbitrary file deletion on the IPM server. Less glamorous than the RCEs but operationally pointed: an attacker can delete the configuration and driver files that let IPM talk to your UPS estate, silently disabling the power-response layer without triggering anything that looks like an attack.","attack_vector":"Unauthenticated, remote, to the IPM server.","remediation":"Upgrade to IPM 1.69 or later. Also verify you have restorable backups of IPM configuration - the recovery path for this bug is restore, and most operators have never tested it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23279"],"status":"curated"},{"id":"CVE-2021-43816","cve":"CVE-2021-43816","aliases":[],"title":"containerd: On SELinux hosts, an unprivileged pod with a hostPath volume can gain full read/write to the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":8,"severity":"high","kev":false,"impact":"On SELinux hosts, an unprivileged pod with a hostPath volume can gain full read/write to the host filesystem","attack_vector":"Cluster user able to create a pod with a hostPath volume","remediation":"Rolling containerd upgrade with node drain; block hostPath in tenant namespaces","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43816"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-21225","cve":"CVE-2022-21225","aliases":[],"title":"Intel Data Center Manager: Improper neutralisation (injection) in Data Center Manager lets an authenticated user with adjacent access…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center Manager","year":"2022","cvss_score":8,"severity":"high","kev":false,"impact":"Improper neutralisation (injection) in Data Center Manager lets an authenticated user with adjacent access escalate privilege on the management plane.","attack_vector":"Authenticated user with network adjacency to the DCM server.","remediation":"Upgrade the Intel Data Center Manager software. This is a management-plane application, so the update is an application upgrade and service restart - no node drain, no firmware, no reboot of managed hosts. The real work is deciding what DCM is allowed to reach: it holds credentials for platform power and telemetry across the fleet, so its network exposure matters more than its version. Upgrade to 4.1 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21225","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00662.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2022-32519","cve":"CVE-2022-32519","aliases":["SEVD-2023-010-06"],"title":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - credential storage: DCE stores device passwords in a recoverable format. Every UPS, PDU and cooling controller credential the…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - credential storage","year":"2022","cvss_score":8,"severity":"high","kev":false,"impact":"DCE stores device passwords in a recoverable format. Every UPS, PDU and cooling controller credential the DCIM system polls with can be recovered by an attacker who reaches the appliance. This is the amplifier that turns any of the other DCE bugs from 'one appliance' into 'the facility'.","attack_vector":"Network access to the DCE instance; recoverable-format storage means no cracking effort is needed once a foothold exists.","remediation":"Upgrade to v7.9.0 or later, then rotate every stored credential - the upgrade re-protects new secrets, it does not un-leak old ones. Budget the rotation properly: it means touching every managed device, and on a large site that is the expensive part, not the upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32519"],"status":"curated"},{"id":"CVE-2023-25529","cve":"CVE-2023-25529","aliases":[],"title":"DGX H100 BMC (KVM daemon): Session token theft via timing side channel","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (KVM daemon)","year":"2023","cvss_score":8,"severity":"high","kev":false,"impact":"Session token theft via timing side channel","attack_vector":"Network-adjacent unauthenticated","remediation":"Flash BMC 23.08.18; rotate BMC credentials","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-208"]},{"id":"CVE-2023-25530","cve":"CVE-2023-25530","aliases":[],"title":"DGX H100 BMC (KVM): Code execution + privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (KVM)","year":"2023","cvss_score":8,"severity":"high","kev":false,"impact":"Code execution + privesc","attack_vector":"Network-adjacent BMC user","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"]},{"id":"CVE-2024-24590","cve":"CVE-2024-24590","aliases":[],"title":"ClearML client SDK: Deserialization of untrusted data — a malicious uploaded artifact executes code on the consuming host","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML client SDK","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"Deserialization of untrusted data — a malicious uploaded artifact executes code on the consuming host","attack_vector":"Customer-supplied artifact pulled by another user's ClearML job","remediation":"Upgrade past 1.14.2. Cross-tenant if a shared ClearML server serves multiple teams","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24590"],"status":"curated","fleet":{"ubiquity":"Common - ClearML is a widely used MLOps/experiment platform beside GPU clusters; affects SDK 0.17.0-1.14.2","remediation_pain":"**Image rebuild** of every image containing the ClearML SDK, plus revocation of every artifact in the store","pain_class":"other","why_fleet_wide":"`Artifact.get()` pickle-deserializes a maliciously uploaded artifact, so one poisoned artifact runs code on every user or job that reads it - a supply-chain fan-out inside the platform, with 10+ public PoCs"}},{"id":"CVE-2024-24591","cve":"CVE-2024-24591","aliases":[],"title":"ClearML client SDK: Path traversal — a malicious dataset writes arbitrary files on the consumer","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML client SDK","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"Path traversal — a malicious dataset writes arbitrary files on the consumer","attack_vector":"Customer-supplied dataset","remediation":"Upgrade past 1.14.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24591"],"status":"curated"},{"id":"CVE-2024-28860","cve":"CVE-2024-28860","aliases":[],"title":"Cilium: IPsec transparent encryption is cryptographically ineffective","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"IPsec transparent encryption is cryptographically ineffective; inter-node traffic can be decrypted or forged","attack_vector":"Anyone with access to the underlay network between nodes","remediation":"Upgrade Cilium and rotate IPsec keys; assume prior inter-node traffic was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28860"],"status":"curated"},{"id":"CVE-2024-37774","cve":"CVE-2024-37774","aliases":[],"title":"Sunbird DCIM dcTrack v9.1.2: CSRF in admin screens lets an authenticated attacker escalate privileges by getting an administrator to load…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Sunbird DCIM dcTrack v9.1.2","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"CSRF in admin screens lets an authenticated attacker escalate privileges by getting an administrator to load a crafted page. dcTrack is the system of record for where every asset, circuit and outlet lives - so administrator access is both a map of the facility and, through its integrations, a route into the devices themselves.","attack_vector":"Requires an authenticated dcTrack administrator to visit an attacker-controlled page while logged in.","remediation":"Upgrade dcTrack past 9.1.2. Software upgrade on one host. Also worth doing: check what credentials dcTrack holds for integrated PDU and power gear, since DCIM asset databases quietly accumulate device logins.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37774"],"status":"curated"},{"id":"CVE-2024-39368","cve":"CVE-2024-39368","aliases":[],"title":"Intel Neural Compressor (SQL injection): SQL injection reachable by an authenticated user of Neural Compressor. Gets an attacker the backing database…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Neural Compressor (SQL injection)","year":"2024","cvss_score":8,"severity":"high","kev":false,"impact":"SQL injection reachable by an authenticated user of Neural Compressor. Gets an attacker the backing database of the optimisation service - job metadata, model references, and whatever credentials the deployment stored there.","attack_vector":"Any authenticated user of the Neural Compressor service.","remediation":"Upgrade to Neural Compressor v3.0 or later. Application-level update, restart the service.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39368","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2025-12638","cve":"CVE-2025-12638","aliases":[],"title":"Keras (`utils.get_file`): Path traversal in tar extraction in 3.11.3 (incomplete fix)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`utils.get_file`)","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Path traversal in tar extraction in 3.11.3 (incomplete fix)","attack_vector":"Customer-supplied archive","remediation":"Upgrade past 3.11.3","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12638"],"status":"curated"},{"id":"CVE-2025-23268","cve":"CVE-2025-23268","aliases":[],"title":"Triton Inference Server: RCE / privesc via input-validation failure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"RCE / privesc via input-validation failure","attack_vector":"Unauthenticated client of the inference endpoint","remediation":"Upgrade Triton; rebuild and redeploy all serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23268","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"]},{"id":"CVE-2025-30165","cve":"CVE-2025-30165","aliases":[],"title":"vLLM (multi-node ZeroMQ): Secondary vLLM host trusts unauthenticated ZeroMQ messages","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (multi-node ZeroMQ)","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Secondary vLLM host trusts unauthenticated ZeroMQ messages","attack_vector":"Co-tenant or compromised worker on the multi-node deployment","remediation":"Upgrade; enforce mTLS or a private VRF for intra-job traffic","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-30165"],"status":"curated"},{"id":"CVE-2025-33179","cve":"CVE-2025-33179","aliases":[],"title":"Cumulus Linux / NVOS: Privesc to switch admin","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux / NVOS","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Privesc to switch admin","attack_vector":"Authenticated switch user","remediation":"Upgrade Cumulus Linux / NVOS; rolling switch upgrade across the fabric","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33179","https://github.com/NVIDIA/product-security/tree/main/2026/5722"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-266"]},{"id":"CVE-2025-33180","cve":"CVE-2025-33180","aliases":[],"title":"Cumulus Linux / NVOS: Command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux / NVOS","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Command injection -> switch takeover","attack_vector":"Authenticated switch user","remediation":"Upgrade Cumulus Linux / NVOS; rolling switch upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33180","https://github.com/NVIDIA/product-security/tree/main/2026/5722"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"]},{"id":"CVE-2025-33188","cve":"CVE-2025-33188","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An attacker tampers with hardware controls directly, reaching data tampering and denial of service at the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"An attacker tampers with hardware controls directly, reaching data tampering and denial of service at the platform level. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33188","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-269"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33245","cve":"CVE-2025-33245","aliases":[],"title":"NeMo Framework: Remote RCE via insecure deserialization over the network","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Remote RCE via insecure deserialization over the network","attack_vector":"Network attacker feeding a serialized payload","remediation":"Bump NeMo; rebuild and redeploy; restrict network exposure of NeMo services","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33245","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-7766","cve":"CVE-2025-7766","aliases":[],"title":"Lantronix Provisioning Manager: Provisioning Manager reads configuration files supplied by the network devices it manages. Because it doesn't…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lantronix Provisioning Manager","year":"2025","cvss_score":8,"severity":"high","kev":false,"impact":"Provisioning Manager reads configuration files supplied by the network devices it manages. Because it doesn't lock down XML external entity resolution, a device (or something spoofing one) can hand it a poisoned config file that leads to unauthenticated remote code execution on the host running Provisioning Manager — which typically has admin reach into the whole fleet of Lantronix devices it provisions.","attack_vector":"Attacker needs to get a malicious XML config file processed by Provisioning Manager — either by compromising/spoofing a managed device on the network, or by feeding it a crafted import file if the workflow allows manual uploads.","remediation":"Software upgrade of Provisioning Manager to the patched release. This runs on a management workstation/server rather than the appliances themselves, so it's a single upgrade rather than a per-device fleet rollout — but treat it as high priority since it's the box with admin credentials to the whole console-server fleet.","references":["https://www.lantronix.com/"],"status":"curated"},{"id":"CVE-2026-24213","cve":"CVE-2026-24213","aliases":[],"title":"Triton Inference Server: Info disclosure / RCE (OOB read in tensor processing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"Info disclosure / RCE (OOB read in tensor processing)","attack_vector":"Network client sending crafted tensors","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24213","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"]},{"id":"CVE-2026-24214","cve":"CVE-2026-24214","aliases":[],"title":"Triton Inference Server: RCE (integer overflow in model config parsing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"RCE (integer overflow in model config parsing)","attack_vector":"Malicious model config","remediation":"Upgrade Triton; redeploy; restrict model repo writes","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24214","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-190"]},{"id":"CVE-2026-47237","cve":"CVE-2026-47237","aliases":[],"title":"Kubeflow Community Distribution: Insecure default in the platform install","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow Community Distribution","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"Insecure default in the platform install","attack_vector":"Tenant user of a Kubeflow-based platform","remediation":"Upgrade to 26.03-rc.1+; provider-owned if the provider ships managed Kubeflow","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47237"],"status":"curated","fleet":{"ubiquity":"Common - Kubeflow is a standard multi-tenant ML platform layer over GPU K8s; official manifests before 1.10 and most packaged distros affected","remediation_pain":"`daemon-restart` - Istio policy and manifest upgrade; the expensive part is invalidating every user token issued before the fix","pain_class":"daemon-restart","why_fleet_wide":"Overly permissive Istio permissions let any user with `kubeflow-edit` in *any* namespace create a notebook that steals other users' authorization tokens - a straight cross-tenant takeover on a shared AI platform"}},{"id":"CVE-2026-59690","cve":"CVE-2026-59690","aliases":[],"title":"Progress Kemp LoadMaster Multi Tenant: TENANT ISOLATION: the Multi Tenant product line's REST API doesn't check whether a caller's permission level…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Progress Kemp LoadMaster Multi Tenant","year":"2026","cvss_score":8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: the Multi Tenant product line's REST API doesn't check whether a caller's permission level actually allows the administrative operation they're requesting. A low-privileged tenant account can invoke privileged administrative operations meant only for the LoadMaster operator, reaching functionality that should be walled off from other tenants sharing the same appliance.","attack_vector":"Requires an authenticated low-privilege account on the Multi Tenant platform — no privilege escalation exploit needed, just calling the REST API endpoints the UI hides but the backend doesn't actually gate.","remediation":"Software upgrade to the fixed release per Progress's July 2026 LoadMaster Critical Security Bulletin (issued alongside four related CVEs). Prioritize this on any shared/multi-tenant LoadMaster deployment, since the whole point of the missing check is that tenant boundaries aren't being enforced.","references":["https://community.progress.com/s/article/LoadMaster-Critical-Security-Bulletin-July-2026-CVE-2026-59686-CVE-2026-59687-CVE-2026-59688-CVE-2026-59689-CVE-2026-59690"],"status":"curated"},{"id":"CVE-2022-41739","cve":"CVE-2022-41739","aliases":[],"title":"IBM Spectrum Scale / Storage Scale Container Native Storage Access: TENANT ISOLATION: programs running inside a container can overcome the isolation mechanism of IBM Spectrum…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Spectrum Scale / Storage Scale Container Native Storage Access","year":"2022","cvss_score":7.9,"severity":"high","kev":false,"impact":"TENANT ISOLATION: programs running inside a container can overcome the isolation mechanism of IBM Spectrum Scale Container Native Storage Access. Spectrum Scale (GPFS) is one of the two or three filesystems that actually keep up with large training clusters, and the container-native access layer is how Kubernetes-scheduled GPU jobs mount it. An isolation escape here means one tenant's pod reaching outside its intended storage boundary on shared cluster storage.","attack_vector":"A process inside a container that has Spectrum Scale container-native storage access — i.e. any tenant workload with a mounted volume.","remediation":"Upgrade Container Native Storage Access past 5.1.6.0. This is a rolling upgrade of the storage-access DaemonSet/operator; pods remount as it rolls, so drain latency-sensitive jobs. No filesystem downtime, but plan for I/O stalls during the roll.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41739"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-45745","cve":"CVE-2023-45745","aliases":[],"title":"Intel TDX module: MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2023","cvss_score":7.9,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: improper input validation reachable by a privileged host user.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45745","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01036.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21980","cve":"CVE-2024-21980","aliases":["SNP firmware write restriction bypass","UMC seed overwrite"],"title":"AMD SEV-SNP firmware (EPYC Milan, Genoa, Bergamo, Siena): SNP firmware fails to restrict where a hypervisor-driven write can land, letting a malicious host overwrite a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP firmware (EPYC Milan, Genoa, Bergamo, Siena)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"SNP firmware fails to restrict where a hypervisor-driven write can land, letting a malicious host overwrite a confidential guest's memory or the UMC seed that derives that guest's memory encryption key. Overwriting guest memory defeats the integrity half of SNP outright; overwriting the UMC seed lets the operator choose the key material protecting a tenant. This is the highest-severity item in the AMD-SB-3011 batch and it is squarely an operator-versus-tenant break - the exact threat model a customer pays the SNP premium for.","attack_vector":"Malicious hypervisor / host root issuing SNP firmware commands against a running confidential guest. Local, high privilege.","remediation":"Two options per AMD-SB-3011: hot-loadable SEV firmware (1.37.14 hex on Milan, 1.37.24 on Genoa) which does NOT require a reboot, or a full Platform Initialization / BIOS flash (MilanPI 1.0.0.D, GenoaPI 1.0.0.C) which does. Take the SEV firmware path first to close the window without draining training jobs, then fold the PI update into your next maintenance cycle. Fixing it moves the SNP TCB version above 0x16 (Milan) / 0x15 (Genoa), which changes every attestation report the host issues - tenants' verifiers must be updated to accept the new TCB or their attestation will start failing.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3011.html","https://nvd.nist.gov/vuln/detail/CVE-2024-21980"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-23599","cve":"CVE-2024-23599","aliases":[],"title":"Intel reference platforms (Seamless Firmware Updates): A race condition in the seamless firmware update mechanism lets a privileged local user cause denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel reference platforms (Seamless Firmware Updates)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"A race condition in the seamless firmware update mechanism lets a privileged local user cause denial of service. Seamless update is the feature that is supposed to let you patch platform firmware without a reboot - a bug that turns it into an outage undermines the whole reason operators enable it.","attack_vector":"Privileged local access on the host.","remediation":"Platform firmware update from the OEM. Ironically this one does need a conventional drain and reboot to install, since the seamless path is what is broken.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23599","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01071.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-37307","cve":"CVE-2024-37307","aliases":[],"title":"Cilium: cilium-bugtool output contains sensitive data","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"cilium-bugtool output contains sensitive data","attack_vector":"Anyone who receives a support bundle","remediation":"Upgrade Cilium; treat existing bugtool archives as secret-bearing","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37307"],"status":"curated"},{"id":"CVE-2024-39585","cve":"CVE-2024-39585","aliases":[],"title":"Dell SmartFabric OS10 (hard-coded password): A hard-coded password in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x, usable by a low-privileged…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (hard-coded password)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"A hard-coded password in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x, usable by a low-privileged attacker with remote access. A shipped credential in a switch NOS is identical on every unit in the fleet and on every other customer's fleet, so once it is known there is no per-device secrecy left. CVE-2024-48831 and CVE-2025-36609 are further hard-coded-password findings in the same product, which makes this a pattern rather than an incident.","attack_vector":"Low-privileged attacker with remote access to the switch. In practice any credential that gets you onto the box at all.","remediation":"OS10 upgrade plus switch reload. There is no config workaround — you cannot change a hard-coded credential. Until patched, the only real control is making the management interface unreachable from anything but a bastion. Rollout: one reload per switch across the leaf/spine.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39585","https://nvd.nist.gov/vuln/detail/CVE-2024-48831","https://nvd.nist.gov/vuln/detail/CVE-2025-36609"],"status":"curated"},{"id":"CVE-2024-52880","cve":"CVE-2024-52880","aliases":["INSYDE-SA-2024016"],"title":"Insyde InsydeH2O (VariableRuntimeDxe, SecureBootHandler): The Secure Boot variable handler bounds-checks incoming data using length fields that the caller supplies, so…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (VariableRuntimeDxe, SecureBootHandler)","year":"2024","cvss_score":7.9,"severity":"high","kev":false,"impact":"The Secure Boot variable handler bounds-checks incoming data using length fields that the caller supplies, so an attacker who lies about the sizes gets the handler to read and write outside the buffer. The affected code is the gatekeeper for the Secure Boot key databases (PK/KEK/db/dbx), which means the compromise targets the mechanism that decides what firmware and bootloaders are allowed to run. Highest-scored member of the four-CVE SA-2024016 VariableRuntimeDxe batch.","attack_vector":"Local admin/root on the host OS making crafted SetVariable / SMM variable-service calls.","remediation":"OEM BIOS update on Insyde kernel 5.2 / 05.29.50, 5.3 / 05.38.50, 5.4 / 05.46.50, 5.5 / 05.54.50, 5.6 / 05.61.50, 5.7 / 05.70.50 or later. Firmware flash, reboot per node. No config workaround. Patch the whole SA-2024016 set together - CVE-2024-52877, -52878 and -52879 are separate defects in the same driver and a partial fix leaves the driver reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-52880","https://www.insyde.com/security-pledge/sa-2024016/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-22889","cve":"CVE-2025-22889","aliases":[],"title":"Intel Xeon 6 with TDX (protected memory range handling): MULTI-TENANT ISOLATION: Improper handling of overlap between protected memory ranges on Xeon 6 with TDX lets…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 with TDX (protected memory range handling)","year":"2025","cvss_score":7.9,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper handling of overlap between protected memory ranges on Xeon 6 with TDX lets a privileged user escalate. Protected memory range enforcement is how TDX keeps one trust domain's pages away from the host and from other TDs, so overlap handling failing is the isolation primitive itself failing.","attack_vector":"Privileged host user on a Xeon 6 TDX platform.","remediation":"OEM platform firmware/BIOS update, not just a TDX module update - which means waiting on your server vendor, a per-node drain and a reboot. Re-attest all trust domains after.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22889","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01311.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-41520","cve":"CVE-2026-41520","aliases":[],"title":"Cilium: cilium-bugtool leaks sensitive data (recurrence of the 2024 issue)","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":7.9,"severity":"high","kev":false,"impact":"cilium-bugtool leaks sensitive data (recurrence of the 2024 issue)","attack_vector":"Anyone who receives a support bundle","remediation":"Upgrade Cilium; treat bugtool archives as secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41520"],"status":"curated"},{"id":"CVE-2026-6540","cve":"CVE-2026-6540","aliases":[],"title":"Calico: Application Layer Policy (Dikastes) does not normalise URL paths, so path-traversal and encoded slashes…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":7.9,"severity":"high","kev":false,"impact":"Application Layer Policy (Dikastes) does not normalise URL paths, so path-traversal and encoded slashes bypass HTTP rules","attack_vector":"Unauthenticated network reaching a policy-protected service","remediation":"Rolling Calico upgrade; do not rely on ALP HTTP rules as the only authorization","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-6540"],"status":"curated"},{"id":"CVE-2026-72312","cve":"CVE-2026-72312","aliases":[],"title":"Linux octeontx2-af (VF rx-mode affecting PF promiscuous state): TENANT ISOLATION: a VF setting its receive mode causes the *physical function's* promiscuous and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux octeontx2-af (VF rx-mode affecting PF promiscuous state)","year":"2026","cvss_score":7.9,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a VF setting its receive mode causes the *physical function's* promiscuous and all-multicast MCAM rules to be deleted, because the enable/disable APIs operate on the PF even when the request arrives over a VF's mailbox. One tenant's VF can therefore change what the host's own interface receives — either blinding the operator's PF, or, in the inverse direction, the coupling means VF-driven rx-mode changes have effects outside the VF's own scope. On a shared OCTEON adapter that is one tenant reaching across the SR-IOV boundary into the host's receive path.","attack_vector":"A tenant with an assigned OCTEON VF issuing a normal `nix_set_rx_mode` mailbox request — no exploit primitive needed, just the ordinary API.","remediation":"Kernel upgrade plus host reboot across OCTEON-equipped nodes. No config workaround; the coupling is in the mailbox handler. Rolling drain per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72312"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2015-2291","cve":"CVE-2015-2291","aliases":["INTEL-SA-00051","iqvw64e.sys","Intel Ethernet Diagnostics Driver BYOVD"],"title":"Intel Ethernet diagnostics driver for Windows (iqvw64e.sys / iqvw32.sys), shipped with Intel network adapter tooling: A signed Intel kernel driver that exposes IOCTLs allowing arbitrary kernel…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet diagnostics driver for Windows (iqvw64e.sys / iqvw32.sys), shipped with Intel network adapter tooling","year":"2015","cvss_score":7.8,"severity":"high","kev":true,"impact":"A signed Intel kernel driver that exposes IOCTLs allowing arbitrary kernel memory read/write. This is the classic bring-your-own-vulnerable-driver primitive: an attacker who already has admin on a Windows host drops this legitimately signed Intel driver, loads it, and gets ring-0 - which they use to disable EDR, tamper with the boot chain, and install persistence. It is actively exploited in the wild (ransomware and intrusion crews). For a GPU operator running any Windows nodes, Windows-based fleet management, or Windows jump hosts on the management plane, this is a live escalation path from a compromised admin account to kernel and then to firmware tooling.","attack_vector":"Local administrator on a Windows host. The driver does not need to have been installed by you - the attacker brings the file, so the fleet's own driver inventory tells you nothing about exposure.","remediation":"There is no patch to deploy in the useful sense, because the attack supplies its own copy of the driver. The fix is blocklisting: enable the Microsoft vulnerable-driver blocklist (HVCI / Windows Defender Application Control), which lists iqvw64e.sys, and enforce driver-signature and WDAC policy on every Windows node and management jump host. Also remove Intel's diagnostic/adapter tooling from golden images where it is not operationally needed. No reboot-and-drain cost on the Linux GPU fleet, but real policy work on the Windows management surface.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-2291","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00051.html","https://www.cisa.gov/known-exploited-vulnerabilities-catalog"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2016-5195","cve":"CVE-2016-5195","aliases":[],"title":"Linux kernel (mm, COW): Dirty COW: privilege escalation via MAP_PRIVATE COW breakage","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (mm, COW)","year":"2016","cvss_score":7.8,"severity":"high","kev":true,"impact":"Dirty COW: privilege escalation via MAP_PRIVATE COW breakage; classic container-escape-to-root primitive [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable (kpatch/Ksplice/KernelCare shipped livepatches); otherwise drain + reboot. Any node still vulnerable is EOL and should be reimaged","references":["https://access.redhat.com/security/cve/CVE-2016-5195"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2017-1000253","cve":"CVE-2017-1000253","aliases":[],"title":"Linux kernel (ELF loader): PIE stack buffer corruption, local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ELF loader)","year":"2017","cvss_score":7.8,"severity":"high","kev":true,"impact":"PIE stack buffer corruption, local root [KEV]","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot. Only affects pre-2018 kernels - treat presence as a fleet-hygiene failure","references":["https://access.redhat.com/security/cve/CVE-2017-1000253"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2017-5709","cve":"CVE-2017-5709","aliases":["INTEL-SA-00086","CVE-2017-5706","CVE-2017-5705"],"title":"Intel Server Platform Services (SPS) firmware 4.0 kernel - the server-chipset variant of ME, Lewisburg PCH / Xeon Scalable: Multiple buffer overflows and privilege escalations in the SPS kernel…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Platform Services (SPS) firmware 4.0 kernel - the server-chipset variant of ME, Lewisburg PCH / Xeon…","year":"2017","cvss_score":7.8,"severity":"high","kev":false,"impact":"Multiple buffer overflows and privilege escalations in the SPS kernel let an unauthorized process on the host run code inside the SPS firmware itself. SPS is the management-engine flavour that ships on Xeon server chipsets - it owns Node Manager power capping, PECI thermal telemetry, and the sideband path to the BMC. Code execution there is below-the-OS persistence that survives a tenant reimage and is invisible to any host-based agent, and because SPS drives power and thermal management it carries direct physical consequence: an attacker in SPS can lie about power/thermal telemetry to the BMC and DCIM, or manipulate power limits on a node.","attack_vector":"A local unprivileged-to-privileged process on the host reaching the SPS firmware through the HECI/MEI interface. No network exposure required, no physical access. On a bare-metal GPU node this is reachable by any tenant who has root.","remediation":"SPS firmware flash, delivered only as an OEM BIOS/firmware bundle - Dell (iDRAC-driven DUP), HPE (SPP), Supermicro, Lenovo, Gigabyte, Quanta each ship their own, and for SA-00086 the OEM releases trailed Intel's November 2017 advisory by one to six months on server boards. Needs a full host reboot, so it drains every running training job on the node. There is no host-side mitigation: SPS cannot be disabled on a server chipset the way AMT can be unprovisioned. Verify the post-flash SPS version out of band via the BMC, not from the host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5709","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00086.html","https://cert-portal.siemens.com/productcert/pdf/ssa-892715.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-20669","cve":"CVE-2018-20669","aliases":[],"title":"Linux i915 GPU kernel driver (execbuffer2 ioctl): MULTI-TENANT ISOLATION: The execbuffer2 ioctl accepted a userspace-supplied address without an access_ok()…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (execbuffer2 ioctl)","year":"2018","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The execbuffer2 ioctl accepted a userspace-supplied address without an access_ok() check, so a local user submitting GPU work could get the kernel to touch an arbitrary address. Execbuffer is the hot path every GPU workload uses, which makes this trivially reachable from any GPU-enabled container.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-20669","http://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.git/log/drivers/gpu/drm/i915/i915_gem_execbuffer.c"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6251","cve":"CVE-2018-6251","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX 10 user-mode driver: A crafted pixel shader writes into unallocated memory in the user-mode driver. Because shaders are…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX 10 user-mode driver","year":"2018","cvss_score":7.8,"severity":"high","kev":false,"impact":"A crafted pixel shader writes into unallocated memory in the user-mode driver. Because shaders are attacker-authored content in any remote-rendering, VDI or cloud-gaming path, this is a code-execution primitive reachable from whatever feeds you shaders.","attack_vector":"Anyone who can submit a shader to the host - a tenant VM in a vGPU/vSGA setup, a remote desktop session, or a local user.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://www.talosintelligence.com/vulnerability_reports/TALOS-2018-0514","https://nvd.nist.gov/vuln/detail/CVE-2018-6251"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-0123","cve":"CVE-2019-0123","aliases":[],"title":"Intel processors supporting SGX (memory protection): Insufficient memory protection on SGX-capable processors gives a privileged local user a privilege-escalation…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors supporting SGX (memory protection)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient memory protection on SGX-capable processors gives a privileged local user a privilege-escalation path. Practically, another entry on the list of reasons an SGX platform needs both microcode currency and enforced re-attestation.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update and TCB recovery. Late-loadable microcode where the OS ships it; reboot required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0123","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00220.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2019-0155","cve":"CVE-2019-0155","aliases":["iGPU Leak"],"title":"Intel processor graphics blitter command streamer: MULTI-TENANT ISOLATION: The graphics blitter accepted commands that could reference memory outside the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processor graphics blitter command streamer","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The graphics blitter accepted commands that could reference memory outside the submitting context, letting a local user with GPU access read or write memory belonging to another context. This is the archetype of the GPU isolation failure operators care about: one tenant's GPU commands reaching another tenant's GPU memory. Mitigated by kernel command parsing, which costs measurable submission throughput.","attack_vector":"Any local user or container able to submit GPU work through the DRM interface.","remediation":"Take the kernel fix (which re-enables blitter command parsing) plus the Intel graphics firmware/driver update, then reboot. Expect a performance regression on blitter-heavy workloads - that is the mitigation working, not a bug. Kernel + graphics driver, no BIOS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0155","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00242.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-11085","cve":"CVE-2019-11085","aliases":[],"title":"Intel i915 graphics kernel-mode driver for Linux (< 5.0): Insufficient input validation in the i915 kernel-mode driver gives a local authenticated user a path to…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel i915 graphics kernel-mode driver for Linux (< 5.0)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient input validation in the i915 kernel-mode driver gives a local authenticated user a path to privilege escalation. Old but still present on long-lived nodes running frozen enterprise kernels.","attack_vector":"Local user with access to the i915 DRM device - including any container granted /dev/dri.","remediation":"Move to a fixed kernel (5.0+ upstream or a distro backport) and reboot the node. Kernel-only fix, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11085","https://www.intel.com/content/www/us/en/security-center/advisory/INTEL-SA-00249.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-13272","cve":"CVE-2019-13272","aliases":[],"title":"Linux kernel (ptrace): Broken permission and object lifetime handling for PTRACE_TRACEME, local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ptrace)","year":"2019","cvss_score":7.8,"severity":"high","kev":true,"impact":"Broken permission and object lifetime handling for PTRACE_TRACEME, local root [KEV]","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2019-13272"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-14565","cve":"CVE-2019-14565","aliases":[],"title":"Intel SGX SDK: Insufficient initialisation in the SGX SDK means enclaves built with the affected SDK can leak uninitialised…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient initialisation in the SGX SDK means enclaves built with the affected SDK can leak uninitialised memory or be pushed into privilege escalation. The fix has to be applied by whoever builds the enclave, which for an operator means chasing your confidential-compute vendors rather than patching your own fleet.","attack_vector":"Local authenticated user interacting with an enclave built against a vulnerable SDK.","remediation":"Rebuild enclaves against SGX SDK 2.5 (Windows) / 2.7 (Linux) or later. Not an operator-side patch: it requires a new enclave binary from the software vendor, a re-signed enclave, and re-attestation. No node reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14565","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00293.html"],"status":"curated"},{"id":"CVE-2019-14566","cve":"CVE-2019-14566","aliases":[],"title":"Intel SGX SDK: Insufficient input validation in the SGX SDK's generated edge routines, letting a local user reach…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insufficient input validation in the SGX SDK's generated edge routines, letting a local user reach information disclosure, privilege escalation, or denial of service against an enclave built with the affected SDK.","attack_vector":"Local authenticated user calling into a vulnerable enclave.","remediation":"Rebuild and re-sign enclaves with a fixed SDK, then re-attest. Vendor-side fix; no operator reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14566","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00293.html"],"status":"curated"},{"id":"CVE-2019-15752","cve":"CVE-2019-15752","aliases":[],"title":"Docker Desktop: Trojan docker-credential-wincred.exe in a world-writable path gives local privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker Desktop","year":"2019","cvss_score":7.8,"severity":"high","kev":true,"impact":"Trojan docker-credential-wincred.exe in a world-writable path gives local privilege escalation. [KEV]","attack_vector":"Local user on a developer/build Windows host","remediation":"Not a cluster-node issue; patch any Windows build/dev machines that run Docker Desktop","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-15752"],"status":"curated"},{"id":"CVE-2019-18181","cve":"CVE-2019-18181","aliases":[],"title":"Arista CloudVision Portal (Configlet Builder API): A read-only CloudVision user escapes their permissions through Configlet Builder API calls and can execute…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista CloudVision Portal (Configlet Builder API)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"A read-only CloudVision user escapes their permissions through Configlet Builder API calls and can execute restricted functionality. Read-only accounts are the ones handed out most freely — dashboards, auditors, tenant liaisons — so this is a large-blast-radius privilege escalation on the fabric controller.","attack_vector":"Any authenticated read-only CVP user with API access.","remediation":"CloudVision Portal upgrade. Controller-side, no switch impact. Review Configlet Builder execution history for anything run by a read-only principal.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18181"],"status":"curated"},{"id":"CVE-2019-5665","cve":"CVE-2019-5665","aliases":[],"title":"NVIDIA Windows GPU Display Driver, 3D Vision stereo service: A privileged NVIDIA service opens files without checking for hard links, so an unprivileged user can point it…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, 3D Vision stereo service","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"A privileged NVIDIA service opens files without checking for hard links, so an unprivileged user can point it at a file they should not be able to write and get arbitrary overwrite - the standard route to SYSTEM on Windows.","attack_vector":"Any local unprivileged user on a Windows host with the driver's stereo service running.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5665"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5666","cve":"CVE-2019-5666","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The context-creation DDI uses an untrusted array index without validating it. Unprivileged local code gets an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"The context-creation DDI uses an untrusted array index without validating it. Unprivileged local code gets an out-of-bounds kernel access, which NVIDIA rates as escalation-capable.","attack_vector":"Any local user with GPU device access on the host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5666"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5667","cve":"CVE-2019-5667","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the page-table DDI handler, reachable locally","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the page-table DDI handler, reachable locally; crash the host, possibly worse.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5667"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5668","cve":"CVE-2019-5668","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the virtual command submission handler. Any local process that can submit work can…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the virtual command submission handler. Any local process that can submit work can bugcheck the node.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5668"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5669","cve":"CVE-2019-5669","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Out-of-bounds kernel buffer access through the escape handler from an unprivileged caller - denial of…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds kernel buffer access through the escape handler from an unprivileged caller - denial of service, with escalation possible.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5669"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5670","cve":"CVE-2019-5670","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Same escape-handler length bug but with information disclosure and code execution called out: kernel memory…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Same escape-handler length bug but with information disclosure and code execution called out: kernel memory contents can leak back to an unprivileged caller.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5670"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5675","cve":"CVE-2019-5675","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Unsynchronized shared state (static variables across threads) in the escape handler. Two threads racing…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unsynchronized shared state (static variables across threads) in the escape handler. Two threads racing produce undefined kernel behaviour - crash, escalation or information disclosure depending on what lands.","attack_vector":"Any local user able to issue concurrent escape calls, which is any process with GPU access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5675"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5683","cve":"CVE-2019-5683","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Hard-link attack on the user-mode video driver's trace logger: a privileged component writes where an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Hard-link attack on the user-mode video driver's trace logger: a privileged component writes where an unprivileged user points it. Arbitrary file overwrite as SYSTEM.","attack_vector":"Any local unprivileged user on the Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://nvd.nist.gov/vuln/detail/CVE-2019-5683"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5690","cve":"CVE-2019-5690","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The escape handler does not validate input buffer size. Unprivileged local caller gets a kernel memory-safety…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"The escape handler does not validate input buffer size. Unprivileged local caller gets a kernel memory-safety bug, rated escalation-capable.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5690"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5691","cve":"CVE-2019-5691","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the escape handler - unprivileged local crash of the GPU node, escalation not excluded","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the escape handler - unprivileged local crash of the GPU node, escalation not excluded.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5691"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5692","cve":"CVE-2019-5692","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Untrusted array index used in the escape handler: an unprivileged process indexes kernel memory of its…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.8,"severity":"high","kev":false,"impact":"Untrusted array index used in the escape handler: an unprivileged process indexes kernel memory of its choosing.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5692"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-0561","cve":"CVE-2020-0561","aliases":[],"title":"Intel SGX SDK (< 2.6.100.1): Improper initialisation in the SGX SDK gives an authenticated local user a privilege escalation path against…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (< 2.6.100.1)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper initialisation in the SGX SDK gives an authenticated local user a privilege escalation path against enclaves built with it.","attack_vector":"Local authenticated user against a vulnerable enclave.","remediation":"Rebuild enclaves with SGX SDK 2.6.100.1 or later and re-attest. Vendor-side.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0561","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00336.html"],"status":"curated"},{"id":"CVE-2020-10699","cve":"CVE-2020-10699","aliases":["CVE-2020-13867","CVE-2020-14019"],"title":"targetcli-fb 2.1.50/2.1.51 and rtslib-fb through 2.1.72 (configuration tooling for the Linux LIO iSCSI/NVMe-oF target): The targetclid socket is world-writable, so any local user on the storage…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"targetcli-fb 2.1.50/2.1.51 and rtslib-fb through 2.1.72 (configuration tooling for the Linux LIO iSCSI/NVMe-oF target)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"The targetclid socket is world-writable, so any local user on the storage node can drive the tool that configures LIO - creating backstores, changing LUN mappings and ACLs, and thereby escalating to root. The related defects leave /etc/target, its backup directory and saveconfig.json with weak permissions, exposing the entire target configuration including initiator ACLs and CHAP secrets to any local reader. Concretely: a non-root foothold on a storage node becomes 'map any tenant's LUN to an initiator I control', which is a cross-tenant data compromise achieved entirely through the target's own supported configuration path, leaving no memory-corruption artefacts to find afterwards.","attack_vector":"Any unprivileged local account on the LIO target host - a container escape, a compromised exporter or backup agent, a shared ops account. Requires that the targetclid socket unit is enabled, which several distributions do by default.","remediation":"Distro package update for targetcli-fb and rtslib-fb; no reboot and no I/O interruption. Independently, mask targetclid.socket unless you actually use the daemon mode - most operators drive targetcli interactively or from configuration management and never need it. Fix permissions on /etc/target, its backups and saveconfig.json, and rotate any CHAP secrets that were stored there, since the file was readable before you fixed it.","references":["https://github.com/open-iscsi/targetcli-fb/issues/162","https://bugzilla.redhat.com/show_bug.cgi?id=CVE-2020-10699","https://nvd.nist.gov/vuln/detail/CVE-2020-10699"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2020-12929","cve":"CVE-2020-12929","aliases":[],"title":"AMD PSP trusted applications shipped in the AMD Graphics Driver: MULTI-TENANT ISOLATION: Trusted applications bundled with the AMD graphics driver and running on the PSP do…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP trusted applications shipped in the AMD Graphics Driver","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Trusted applications bundled with the AMD graphics driver and running on the PSP do not validate their parameters, letting a local attacker bypass security restrictions and get arbitrary code execution in the secure processor. Notable because the entry point is the GPU driver stack rather than platform firmware - a GPU-adjacent compromise reaching the platform root of trust.","attack_vector":"Local, via the graphics driver's PSP interface.","remediation":"Fixed by updating the AMD graphics driver package (which carries the PSP trusted applications), then reloading the driver or rebooting the node. Unlike the pure-firmware ASP issues this does not wait on an OEM BIOS cycle - it moves at driver-release speed, so it is one of the cheaper ASP-class fixes to deploy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12929","https://www.amd.com/en/resources/product-security.html"],"status":"curated"},{"id":"CVE-2020-12930","cve":"CVE-2020-12930","aliases":[],"title":"AMD Secure Processor (ASP) drivers: Improper parameter handling in the ASP driver layer lets an already-privileged attacker escalate further and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) drivers","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper parameter handling in the ASP driver layer lets an already-privileged attacker escalate further and corrupt state the platform relies on for integrity. The realistic gain is turning root on the host into control over firmware-level components, i.e. persistence that survives reimaging.","attack_vector":"Local, privileged. Requires root or equivalent on the host first.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12930","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-12931","cve":"CVE-2020-12931","aliases":[],"title":"AMD Secure Processor (ASP) kernel: Improper parameter handling in the ASP's own kernel gives a privileged attacker a path to elevate inside the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) kernel","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper parameter handling in the ASP's own kernel gives a privileged attacker a path to elevate inside the secure processor and damage platform integrity. Same practical outcome as the ASP driver flaw: root on the box becomes control of the firmware trust anchor.","attack_vector":"Local, privileged.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12931","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-12961","cve":"CVE-2020-12961","aliases":[],"title":"AMD PSP - System Management Network privileged register zeroing: MULTI-TENANT ISOLATION: An attacker can zero any privileged register on the System Management Network via the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP - System Management Network privileged register zeroing","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An attacker can zero any privileged register on the System Management Network via the PSP. Zeroing SMN registers is a general-purpose way to disable platform protections - lock bits, access-control gates, security configuration - which turns into a bypass of whatever those registers were enforcing. It is a primitive rather than a single bug: whatever protection you were relying on at the SMN level can be switched off.","attack_vector":"Local, privileged, through the PSP interface.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12961","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-12964","cve":"CVE-2020-12964","aliases":[],"title":"AMD Radeon Kernel Mode driver - Escape 0x2000c00 call handler: MULTI-TENANT ISOLATION: A low-privileged attacker can drive the Radeon kernel-mode driver's Escape 0x2000c00…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD Radeon Kernel Mode driver - Escape 0x2000c00 call handler","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A low-privileged attacker can drive the Radeon kernel-mode driver's Escape 0x2000c00 handler into privilege escalation or denial of service. Escape call handlers are the driver's catch-all ioctl surface and historically its weakest - a low-privilege caller reaching kernel-mode escalation is the pattern that makes GPU device nodes dangerous to hand out.","attack_vector":"Local, low privilege - reachable by an ordinary user with the GPU device open.","remediation":"Update the AMD graphics driver and reload or reboot. Note this is the Windows kernel-mode driver surface; Linux ROCm fleets are not affected by this specific handler, though the lesson about escape/ioctl surfaces carries over.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12964","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-36787","cve":"CVE-2020-36787","aliases":[],"title":"ASPEED video engine driver clock/reset sequencing (drivers/media/platform/aspeed): The driver brings the video engine out of reset in the wrong order relative to its two clocks, and the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED video engine driver clock/reset sequencing (drivers/media/platform/aspeed)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver brings the video engine out of reset in the wrong order relative to its two clocks, and the hardware responds by issuing DMA writes to effectively random BMC memory. This is a DMA engine scribbling on the management processor's RAM with no software mediating the target address - the failure mode is silent BMC state corruption and unexplained BMC hangs or reboots. Operators usually log these as flaky hardware and RMA the board; the actual cost is a management processor whose memory integrity you cannot reason about, on every node running an affected image with video capture enabled.","attack_vector":"Triggers on video engine initialization, i.e. whenever iKVM/video capture is started or restarted on an affected BMC image. No attacker required for the corruption itself, but a host-side tenant who can force display-mode changes can force the reinit repeatedly.","remediation":"Fixed in the kernel driver and backported to stable branches; delivery is a BMC firmware flash, per node, out-of-band, ODM-gated. Config-only mitigation is to leave the video capture service stopped on nodes that do not use graphical console. Worth pairing with a check of your BMC crash/reboot telemetry - if you have unexplained BMC resets on ASPEED nodes with iKVM enabled, this is a candidate cause rather than bad silicon.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-36787","https://git.kernel.org/stable/c/1dc1d30ac101bb8335d9852de2107af60c2580e7"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-5963","cve":"CVE-2020-5963","aliases":[],"title":"NVIDIA GPU Display Driver, inter-process communication APIs: Improper access control on the driver's IPC surface gives a local attacker code execution, a crash, or a read…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, inter-process communication APIs","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper access control on the driver's IPC surface gives a local attacker code execution, a crash, or a read of data crossing that IPC. This one also shipped as an Ubuntu security update, so Linux compute nodes are in scope, not just Windows workstations.","attack_vector":"Any local user or container on the host that can reach the driver's IPC endpoints.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://usn.ubuntu.com/4404-1/","https://usn.ubuntu.com/4404-2/","https://nvd.nist.gov/vuln/detail/CVE-2020-5963"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5964","cve":"CVE-2020-5964","aliases":[],"title":"NVIDIA GPU Display Driver service host component: The service host can skip its integrity check on application resources, so a local attacker who swaps a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver service host component","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"The service host can skip its integrity check on application resources, so a local attacker who swaps a resource gets code execution in a privileged NVIDIA service.","attack_vector":"Local user able to write the resources the service loads.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5964"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5966","cve":"CVE-2020-5966","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the escape handler","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"NULL dereference in the escape handler; unprivileged local caller crashes the node, escalation not excluded.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5966"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5968","cve":"CVE-2020-5968","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the vGPU plugin fails to bound an indexed or pointer-based access, and NVIDIA calls…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the vGPU plugin fails to bound an indexed or pointer-based access, and NVIDIA calls out code execution, escalation and information disclosure. This is the guest-to-host escape shape - a tenant VM running code in the hypervisor's vGPU plugin owns every other tenant on that GPU. Affects vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU assigned.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5968"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5971","cve":"CVE-2020-5971","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: out-of-bounds read in the vGPU plugin that NVIDIA rates as code-execution capable. A…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: out-of-bounds read in the vGPU plugin that NVIDIA rates as code-execution capable. A tenant VM reads past a host buffer - the direct route to leaking another tenant's data or to a host escape. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5971"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5980","cve":"CVE-2020-5980","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A securely loaded system DLL then loads its own dependencies insecurely - the fix for one preloading bug…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"A securely loaded system DLL then loads its own dependencies insecurely - the fix for one preloading bug leaving a second-order one behind. Local user gets code execution in a privileged driver component.","attack_vector":"Local user with write access on the dependency search path.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5980"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5981","cve":"CVE-2020-5981","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll): Crafted shader triggers an out-of-bounds access in the DX11 user-mode driver, this time rated code-execution…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted shader triggers an out-of-bounds access in the DX11 user-mode driver, this time rated code-execution capable. Anywhere untrusted shaders reach the GPU - VDI, cloud gaming, remote rendering - treat it as RCE at the session's privilege level.","attack_vector":"Anyone who can submit shaders to the host, including from a guest VM or remote session.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5981"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5984","cve":"CVE-2020-5984","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: use-after-free while releasing resources in the vGPU plugin, rated code-execution…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: use-after-free while releasing resources in the vGPU plugin, rated code-execution capable. Guest-triggered UAF in the host plugin is the classic guest-to-host escape. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5984"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5987","cve":"CVE-2020-5987","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: parameters stay writable by the guest after the plugin has validated them - a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: parameters stay writable by the guest after the plugin has validated them - a textbook double-fetch. The guest passes validation with benign values then swaps them, feeding invalid parameters into host handlers. Escalation on the hypervisor host. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU who can race the host's validation window.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5987"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5991","cve":"CVE-2020-5991","aliases":[],"title":"NVIDIA CUDA Toolkit, nvJPEG library: Out-of-bounds read/write in nvJPEG while decoding an image, rated code-execution capable. nvJPEG sits in…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit, nvJPEG library","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds read/write in nvJPEG while decoding an image, rated code-execution capable. nvJPEG sits in exactly the place you feed untrusted data: GPU-accelerated image decode in training and inference pipelines, DALI preprocessing, and media services. If your ingest path decodes user-uploaded JPEGs on the GPU, this is remote code execution reachable by whoever can upload an image.","attack_vector":"Anyone who can get a malformed JPEG into a pipeline that decodes with nvJPEG - typically an unauthenticated uploader, several hops upstream of the GPU.","remediation":"Upgrade the CUDA Toolkit to 11.1.1 or later. The real cost is not the host install: every container image, wheel, conda package and vendored build that statically carries the affected library has to be rebuilt and re-pushed, and running jobs restarted onto the new image. No driver reload, no node drain, no firmware flash - but a full image-fleet rebuild.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5991"],"status":"curated"},{"id":"CVE-2020-7053","cve":"CVE-2020-7053","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2020","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the per-process GTT close path freeing structures the VM still references.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-7053","https://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.git/commit/?id=7dc40713618c884bf07c030d1ab1f47a9dc1f310"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-0084","cve":"CVE-2021-0084","aliases":[],"title":"Intel RDMA driver for Ethernet X722 and 800 series (Linux): MULTI-TENANT ISOLATION: Improper input validation in the Intel RDMA Linux driver gives an authenticated user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel RDMA driver for Ethernet X722 and 800 series (Linux)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper input validation in the Intel RDMA Linux driver gives an authenticated user privilege escalation. Earlier member of the same family as the irdma access-control issue and the same reason to care: RDMA queue pairs are mapped into userspace by design.","attack_vector":"Authenticated local user with access to the RDMA verbs interface.","remediation":"Update the Intel RDMA driver to 1.3.19 or later. Module reload drops RDMA links - drain first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0084","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00515.html"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-1052","cve":"CVE-2021-1052","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): User-mode clients can reach legacy privileged APIs in the kernel-mode layer through DxgkDdiEscape or ioctl.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"User-mode clients can reach legacy privileged APIs in the kernel-mode layer through DxgkDdiEscape or ioctl. This is a whole class of privileged driver functionality left exposed to unprivileged callers - NVIDIA rates it escalation and disclosure capable on both Windows and Linux. On a Linux GPU node, any container with /dev/nvidiactl can reach it.","attack_vector":"Any local user or GPU container on the host. Container isolation does not help: the device node is the attack surface and the standard container toolkit maps it in.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1052"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1057","cve":"CVE-2021-1057","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the vGPU plugin lets guests allocate resources they are not authorised to hold, which…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the vGPU plugin lets guests allocate resources they are not authorised to hold, which NVIDIA describes as integrity and confidentiality loss. A tenant VM grabbing unauthorised host GPU resources is the isolation boundary failing outright. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1057"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1059","cve":"CVE-2021-1059","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated guest index leads to integer overflow in the host plugin, then tampering…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated guest index leads to integer overflow in the host plugin, then tampering or disclosure. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1059"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1063","cve":"CVE-2021-1063","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated offset produces a buffer overread in the host plugin. A tenant reads host…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated offset produces a buffer overread in the host plugin. A tenant reads host memory past the intended buffer - the direct path to another tenant's data. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1063"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1080","cve":"CVE-2021-1080","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated guest input in the vGPU Manager plugin giving host information…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated guest input in the vGPU Manager plugin giving host information disclosure, data tampering or a shared-GPU outage. Affects vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1080"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1081","cve":"CVE-2021-1081","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): MULTI-TENANT ISOLATION: unvalidated length across the guest kernel-mode driver and vGPU plugin boundary.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated length across the guest kernel-mode driver and vGPU plugin boundary. Tenant-to-host disclosure or tampering. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1081"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1082","cve":"CVE-2021-1082","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated input length in the vGPU Manager plugin, yielding host disclosure…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated input length in the vGPU Manager plugin, yielding host disclosure, tampering or denial of service on the shared GPU. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1082"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1083","cve":"CVE-2021-1083","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): MULTI-TENANT ISOLATION: unvalidated length across the guest driver and vGPU Manager boundary. vGPU 12.x…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated length across the guest driver and vGPU Manager boundary. vGPU 12.x before 12.2, 11.x before 11.4.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1083"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1084","cve":"CVE-2021-1084","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): MULTI-TENANT ISOLATION: another unvalidated-length path between guest driver and vGPU Manager, with the same…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: another unvalidated-length path between guest driver and vGPU Manager, with the same tenant-to-host disclosure and tampering outcome. vGPU 12.x before 12.2, 11.x before 11.4.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1084"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1089","cve":"CVE-2021-1089","aliases":[],"title":"NVIDIA GPU Display Driver for Windows, nvidia-smi: nvidia-smi loads DLLs from an uncontrolled path. nvidia-smi is run by monitoring agents, schedulers and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver for Windows, nvidia-smi","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"nvidia-smi loads DLLs from an uncontrolled path. nvidia-smi is run by monitoring agents, schedulers and health checks, usually as a privileged service and often on a timer - so an unprivileged user who can plant a DLL gets scheduled code execution as that service on every GPU node running the same monitoring stack.","attack_vector":"Any local user who can write to a directory on nvidia-smi's DLL search path on a Windows GPU host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1089"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1097","cve":"CVE-2021-1097","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the vGPU Manager trusts a length field in a guest request that does not match the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the vGPU Manager trusts a length field in a guest request that does not match the actual input. A malicious tenant lies about the size and gets host-side disclosure, tampering or a shared-GPU outage. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1097"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1098","cve":"CVE-2021-1098","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the vGPU Manager fails to release resources on guest driver unload, and the guest can…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the vGPU Manager fails to release resources on guest driver unload, and the guest can then reuse them. A tenant that unloads and reloads its driver operates on stale host resources - potentially ones now belonging to another tenant. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any user inside a guest VM who can unload and reload the vGPU guest driver, which requires only admin inside their own VM.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1098"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1118","cve":"CVE-2021-1118","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the guest OS can execute privileged operations through the vGPU Manager, which NVIDIA…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the guest OS can execute privileged operations through the vGPU Manager, which NVIDIA rates as escalation, tampering and disclosure. Privileged host operations driven from inside a tenant VM is the escape scenario, not a degraded-service scenario.","attack_vector":"Any user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1118"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-22555","cve":"CVE-2021-22555","aliases":[],"title":"Linux kernel (netfilter x_tables): Heap out-of-bounds write in xt_compat_target_from_user()","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter x_tables)","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"Heap out-of-bounds write in xt_compat_target_from_user(); reliable container escape via unprivileged userns + CAP_NET_ADMIN [KEV]","attack_vector":"Any tenant process in a container with a user namespace","remediation":"Livepatchable; otherwise drain + reboot. Compensating control: disable unprivileged user namespaces (`kernel.unprivileged_userns_clone=0`) - breaks rootless Podman/Apptainer, which many HPC tenants rely on","references":["https://access.redhat.com/security/cve/CVE-2021-22555"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-25124","cve":"CVE-2021-25124","aliases":[],"title":"BMC firmware on the HPE Cloudline whitebox line: An attacker directs the BMC's video-deletion routine at arbitrary filesystem paths, deleting files inside the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BMC firmware on the HPE Cloudline whitebox line","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"An attacker directs the BMC's video-deletion routine at arbitrary filesystem paths, deleting files inside the controller. Destroying the right files on a BMC means destroying its configuration, its credential store, or its ability to boot - a denial of service against the out-of-band plane on nodes you may not be able to reach any other way. It also destroys the BMC-side record of what happened, which makes it a useful anti-forensics step after a more serious compromise. The broader operator lesson is inventory: Cloudline nodes look like HPE in your asset database but share a BMC codebase with the ODM whiteboxes, and this is one of sixteen CVEs published against that BMC's REST service in a single disclosure. CL5800 Gen9, CL5200 Gen9, CL4100 Gen10, CL3100 Gen10 and CL5800 Gen10. Path traversal in the deletevideo_func handler of spx_restservice. Cloudline is HPE's ODM-manufactured whitebox range and runs a MegaRAC-derived BMC, not iLO, so iLO advisories and iLO tooling do not cover it.","attack_vector":"Access to the BMC's spx_restservice REST interface on the management network. The advisory characterises the path as local to the BMC, meaning it needs a session against the controller rather than pure network anonymity.","remediation":"BMC firmware update per HPE's Cloudline advisory - HPE does publish a readable advisory document for this line, which puts it ahead of most whitebox vendors. Flash per node, out of band. The inventory action matters as much as the flash: audit your fleet for Cloudline nodes specifically and confirm they are being tracked against Cloudline BMC advisories rather than iLO ones, because tooling that assumes 'HPE server means iLO' will silently report them as covered when they are not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25124","https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04073en_us"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-25749","cve":"CVE-2021-25749","aliases":[],"title":"Kubernetes (kubelet): Windows workloads run as ContainerAdministrator despite runAsNonRoot","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Windows workloads run as ContainerAdministrator despite runAsNonRoot","attack_vector":"Any tenant workload on a Windows node","remediation":"Rolling kubelet upgrade with Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25749"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-26315","cve":"CVE-2021-26315","aliases":[],"title":"AMD PSP boot ROM - integrity of decrypted firmware image: MULTI-TENANT ISOLATION: The PSP boot ROM authenticates and decrypts firmware but does not sufficiently verify…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP boot ROM - integrity of decrypted firmware image","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The PSP boot ROM authenticates and decrypts firmware but does not sufficiently verify the integrity of the *decrypted* image before using it. An attacker who can influence the encrypted blob can therefore get the boot ROM to execute content it never really validated - code execution in the earliest, most privileged stage of the platform, in mask ROM territory where no patch can reach the flawed check itself.","attack_vector":"Local, requires the ability to modify the firmware image in SPI ROM - root plus flash write, a compromised BMC, or supply-chain access.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. Since the flawed logic is in boot ROM, the mitigation is in the firmware AMD ships around it rather than a fix to the ROM. Practical compensating controls: enforce SPI write protection, require signed BIOS update packages, and restrict which management paths can drive host flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26315","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26324","cve":"CVE-2021-26324","aliases":[],"title":"AMD SEV-ES Trusted Memory Region - SNP guest memory integrity: MULTI-TENANT ISOLATION: A bug in the SEV-ES Trusted Memory Region handling costs memory integrity for…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES Trusted Memory Region - SNP guest memory integrity","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A bug in the SEV-ES Trusted Memory Region handling costs memory integrity for SNP-active VMs. The TMR is the region the SEV firmware itself works in; a defect there means the component enforcing confidential-VM isolation can have its own working memory disturbed, and SNP guests lose the integrity guarantee they were sold.","attack_vector":"Requires host/hypervisor privilege on a machine running SNP guests.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26324","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26331","cve":"CVE-2021-26331","aliases":["SMU mailbox manipulation"],"title":"AMD System Management Unit (SMU) mailbox interface: A malicious user can manipulate SMU mailbox entries and reach arbitrary code execution in the System…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD System Management Unit (SMU) mailbox interface","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A malicious user can manipulate SMU mailbox entries and reach arbitrary code execution in the System Management Unit. The SMU is the microcontroller that owns voltage, clock and thermal control for the package. Code execution there is not just another privilege boundary - it is PHYSICAL control of the part: undervolting to induce computational faults (the technique behind fault-injection attacks on secure enclaves), forcing thermal or power states that throttle or hard-shut-down a node, and doing it in a way the OS reports as a normal thermal event. On a dense GPU rack that is a denial-of-service lever against neighbours and potentially a hardware-damage lever. Sibling issues CVE-2021-26329 and CVE-2021-26330 are the overflow variants in the same interface.","attack_vector":"Local attacker able to reach the SMU mailbox - typically ring 0 on the host, or a bare-metal tenant.","remediation":"AGESA / BIOS update per AMD-SB-1021 from the OEM, flash plus reboot. Independently, restrict tenant access to power and thermal management interfaces (no raw MSR access, no vendor overclocking or power-tuning drivers in tenant images) - that mitigation is under your control and does not wait on a BIOS drop. Add out-of-band power and thermal telemetry so an SMU-driven event is distinguishable from a genuine cooling fault.","references":["https://www.amd.com/en/corporate/product-security/bulletin/amd-sb-1021","https://nvd.nist.gov/vuln/detail/CVE-2021-26331"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26335","cve":"CVE-2021-26335","aliases":[],"title":"AMD Secure Processor (ASP) bootloader - image header parsing: MULTI-TENANT ISOLATION: The ASP bootloader reads and acts on fields from a firmware image header *before* it…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) bootloader - image header parsing","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The ASP bootloader reads and acts on fields from a firmware image header *before* it verifies that image's signature. Attacker-controlled values therefore reach range checks and pointer arithmetic in the pre-verification window, giving code execution in the secure processor's own bootloader. This is upstream of every signature check on the platform: an attacker who lands here owns the root of trust and can persist beneath any OS reinstall or disk wipe.","attack_vector":"Local. Requires the ability to place a crafted image where the ASP bootloader will parse it - in practice SPI ROM write access or a compromised firmware update path, so root plus flash access, or a supply-chain/refurbishment scenario.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this is bootloader code inside the ASP, there is no software workaround and no kernel-side mitigation - you either get the OEM BIOS or you do not. In the meantime, the compensating control is guarding SPI write access: enable the platform's SPI ROM protection and BIOS write-protect, and treat any node that has been through third-party hands as untrusted.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26335","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26360","cve":"CVE-2021-26360","aliases":[],"title":"AMD Secure Processor - SoC security-configuration registers: MULTI-TENANT ISOLATION: A local attacker can make unauthorised changes to the SoC's security configuration…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - SoC security-configuration registers","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A local attacker can make unauthorised changes to the SoC's security configuration registers and from there corrupt the AMD Secure Processor's encrypted memory. Corrupting ASP memory is not just a crash - it is a route to influencing what the secure processor computes, which is the same engine that gates memory encryption and attestation for every confidential guest on the node.","attack_vector":"Local, with access to the SoC register interface - root on the host.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26360","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26409","cve":"CVE-2021-26409","aliases":[],"title":"AMD SEV-ES - bounds checking on Reverse Map table memory: MULTI-TENANT ISOLATION: Insufficient bounds checking in SEV-ES lets an attacker corrupt Reverse Map table…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES - bounds checking on Reverse Map table memory","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient bounds checking in SEV-ES lets an attacker corrupt Reverse Map table memory, breaking SEV-SNP memory integrity. The RMP is the single data structure that decides which physical page belongs to which guest; corrupting it is the most direct possible attack on confidential-VM isolation, because after that the hardware itself believes the wrong owner.","attack_vector":"Host/hypervisor-privileged attacker.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands. Treat RMP-corruption issues as the top tier of your SEV patch queue - everything else in SNP rests on the RMP being correct.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26409","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-27365","cve":"CVE-2021-27365","aliases":[],"title":"Linux iSCSI: iSCSI netlink structures lack length checks","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iSCSI","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"iSCSI netlink structures lack length checks -> heap overflow, unprivileged local root","attack_vector":"Local","remediation":"Data-plane: kernel patch + reboot on every node using iSCSI LUNs","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27365"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-28500","cve":"CVE-2021-28500","aliases":[],"title":"Arista EOS (AAA API): Incorrect AAA API usage enables unrestricted local device access — an operator with limited role gets full…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (AAA API)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Incorrect AAA API usage enables unrestricted local device access — an operator with limited role gets full switch control","attack_vector":"Local","remediation":"EOS upgrade fabric-wide","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28500"],"status":"curated"},{"id":"CVE-2021-28501","cve":"CVE-2021-28501","aliases":[],"title":"Arista EOS (TerminAttr AAA): TerminAttr streaming-telemetry agent bypasses AAA, giving unauthorized local device access — TerminAttr is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (TerminAttr AAA)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"TerminAttr streaming-telemetry agent bypasses AAA, giving unauthorized local device access — TerminAttr is the CloudVision telemetry agent running on essentially every Arista switch in a managed fabric","attack_vector":"Local","remediation":"EOS + TerminAttr upgrade across the fabric","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28501"],"status":"curated"},{"id":"CVE-2021-3156","cve":"CVE-2021-3156","aliases":[],"title":"sudo: Baron Samedit: heap overflow in sudo argument parsing, root from any local account","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"sudo","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"Baron Samedit: heap overflow in sudo argument parsing, root from any local account [KEV]","attack_vector":"Local user","remediation":"Package update only; no reboot","references":["https://access.redhat.com/security/cve/CVE-2021-3156"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2021-33123","cve":"CVE-2021-33123","aliases":["INTEL-SA-00601","CVE-2021-0159","CVE-2021-33124","CVE-2021-33103"],"title":"BIOS Authenticated Code Module (ACM) for a broad set of Intel processors, including Xeon Scalable: Improper access control (plus, in the sibling CVEs, an out-of-bounds write and an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BIOS Authenticated Code Module (ACM) for a broad set of Intel processors, including Xeon Scalable","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper access control (plus, in the sibling CVEs, an out-of-bounds write and an input-validation flaw) in the BIOS ACM. ACMs are Intel-signed code that runs in the authenticated-code execution mode used to establish the root of trust for Boot Guard and TXT - it is more privileged than SMM and more privileged than the hypervisor. Code execution or state corruption at ACM level lets an attacker subvert the measurement chain from underneath, which means a platform can present a valid measured-boot report while running attacker-controlled firmware. Persistence at this level is below-the-OS, survives reimage, and defeats attestation-based tenant-handoff checks.","attack_vector":"A privileged local user - local root or SMM-capable code on the host. On bare-metal GPU nodes handed to tenants with root, that is the tenant.","remediation":"The ACM ships inside the BIOS image, so the fix is a BIOS/platform-firmware update from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta, Wiwynn) - reboot and job drain, and for a May 2022 Intel advisory the OEM server BIOS releases spread across the rest of 2022. There is no configuration workaround: you cannot disable the ACM. If your fleet's trust story depends on Boot Guard or TXT measurements, treat un-updated nodes as not attestable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33123","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00601.html","https://security.netapp.com/advisory/ntap-20220818-0003/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-33909","cve":"CVE-2021-33909","aliases":[],"title":"Linux kernel (seq_file / fs layer): Sequoia: size_t-to-int conversion in the filesystem layer, local root on default configs","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (seq_file / fs layer)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Sequoia: size_t-to-int conversion in the filesystem layer, local root on default configs","attack_vector":"Local user / tenant process able to mount or traverse deep paths","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-33909"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-34398","cve":"CVE-2021-34398","aliases":[],"title":"NVIDIA DCGM (nv-hostengine, DIAG module): Any local user can inject a shared library into the DCGM server, which normally runs as root - so DCGM turns…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DCGM (nv-hostengine, DIAG module)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Any local user can inject a shared library into the DCGM server, which normally runs as root - so DCGM turns every unprivileged account on a GPU node into root. This is the single most operationally relevant NVIDIA bug of 2021 for a cluster operator, because DCGM is not optional infrastructure: it is what feeds dcgm-exporter, Prometheus, node health checks and scheduler telemetry, so it is installed and running as root on essentially every GPU node in a modern fleet. Combined with a container that has host access to the DCGM socket, it is a container-to-host root escape. Versions before 2.2.9.","attack_vector":"Any local user on a node running nv-hostengine, including anything running inside a container that can reach the DCGM socket or port - which is the normal configuration for dcgm-exporter deployments.","remediation":"Upgrade DCGM to 2.2.9 or later on every node running nv-hostengine, then restart the service (systemctl restart nvidia-dcgm) and restart dcgm-exporter containers so they link the new library. No GPU drain, no reboot, no firmware flash. Until you patch, stop running nv-hostengine as root or block local access to its socket/port - the DIAG path is reachable by any local user.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34398"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2021-3490","cve":"CVE-2021-3490","aliases":[],"title":"Linux kernel (eBPF verifier): eBPF ALU32 bitwise-op bounds tracking flaw - arbitrary kernel read/write from unprivileged BPF","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF verifier)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"eBPF ALU32 bitwise-op bounds tracking flaw - arbitrary kernel read/write from unprivileged BPF","attack_vector":"Any tenant process in a container where unprivileged BPF is enabled","remediation":"Livepatchable; otherwise drain + reboot. Durable control: `kernel.unprivileged_bpf_disabled=1` fleet-wide","references":["https://access.redhat.com/security/cve/CVE-2021-3490"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-3560","cve":"CVE-2021-3560","aliases":[],"title":"polkit: Local privilege escalation via polkit_system_bus_name_get_creds_sync() race","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"polkit","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"Local privilege escalation via polkit_system_bus_name_get_creds_sync() race [KEV]","attack_vector":"Local user","remediation":"Package update + restart polkitd; no reboot","references":["https://access.redhat.com/security/cve/CVE-2021-3560"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2021-4034","cve":"CVE-2021-4034","aliases":[],"title":"polkit (pkexec): PwnKit: local privilege escalation to root via argv handling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"polkit (pkexec)","year":"2021","cvss_score":7.8,"severity":"high","kev":true,"impact":"PwnKit: local privilege escalation to root via argv handling; no exploit prerequisites [KEV]","attack_vector":"Local user, incl. any shell inside a privileged/host-namespace container","remediation":"Package update only; no reboot. Interim mitigation: `chmod 0755 /usr/bin/pkexec` (strip setuid)","references":["https://access.redhat.com/security/cve/CVE-2021-4034"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2021-41103","cve":"CVE-2021-41103","aliases":[],"title":"containerd: Container root dirs and plugin dirs created world-traversable","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Container root dirs and plugin dirs created world-traversable; unprivileged host user can read/modify container state","attack_vector":"Any process on the node, including a container that has partial host filesystem access","remediation":"Rolling containerd upgrade plus permission fix; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41103"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-42252","cve":"CVE-2021-42252","aliases":[],"title":"ASPEED LPC control driver (drivers/soc/aspeed/aspeed-lpc-ctrl.c) in the OpenBMC kernel: A process on the BMC that can open the Aspeed LPC control device gets to mmap past the region it is supposed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED LPC control driver (drivers/soc/aspeed/aspeed-lpc-ctrl.c) in the OpenBMC kernel","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A process on the BMC that can open the Aspeed LPC control device gets to mmap past the region it is supposed to own and write into BMC kernel memory. The size check compares the wrong quantities, so it is not a bounds check at all. The payoff is BMC root/kernel from a lower-privileged BMC daemon - which matters because OpenBMC's whole isolation story is that bmcweb, ipmid and the host-mailbox daemons run as separate constrained users. Chain it behind any of the bmcweb memory-corruption bugs and you go from a crashed web server to full control of the management processor, which is the position you need to write flash and persist.","attack_vector":"Local on the BMC itself. Requires code execution as a user with access to the LPC control character device - i.e. an attacker who already landed on the BMC via a network daemon bug, or a malicious/compromised OpenBMC package.","remediation":"Kernel fix, landed in Linux 5.14.6 and backported. For an operator this is not a package update - the BMC kernel is baked into the firmware image, so it means a full BMC firmware flash per node, out-of-band, gated on your ODM rebasing their OpenBMC tree. Many ODM images sit years behind upstream. No config-only mitigation; the device node has to exist for host-BMC mailbox and flash-sharing features to work.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-42252","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=b49a0e69a7b1a68c8d3f64097d06dabb770fec96","https://security.netapp.com/advisory/ntap-20211112-0006/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-46771","cve":"CVE-2021-46771","aliases":[],"title":"AMD Secure Processor (ASP) firmware system-call interface: MULTI-TENANT ISOLATION: The ASP firmware does not validate addresses passed across its system-call boundary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) firmware system-call interface","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The ASP firmware does not validate addresses passed across its system-call boundary, so a compromised user application that can issue ASP syscalls can steer the secure processor into reading or writing memory of the caller's choosing, ending in code execution inside the ASP. That converts a userspace compromise into control of the platform's security engine.","attack_vector":"Local. Requires an already-compromised application with ASP syscall access - typically a privileged agent or a trusted application, not a plain tenant container.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46771","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-47046","cve":"CVE-2021-47046","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix off by one in hdmi_14_process_transaction()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47046","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47081","cve":"CVE-2021-47081","aliases":[],"title":"habanalabs kernel driver (gaudi_memset_device_memory): Use-after-free in the Gaudi device-memory memset path: the command buffer is released on the error path and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (gaudi_memset_device_memory)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in the Gaudi device-memory memset path: the command buffer is released on the error path and then dereferenced again. Gives a local accelerator user a kernel UAF - crash at minimum, and the usual UAF privilege-escalation potential with enough heap grooming.","attack_vector":"Local user holding the habanalabs device node, reached by driving the memset ioctl down an error path.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47081","https://git.kernel.org/stable/c/115726c5d312b462c9d9931ea42becdfa838a076"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47142","cve":"CVE-2021-47142","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix a use-after-free","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47142","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47421","cve":"CVE-2021-47421","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A race condition or locking defect in the amdgpu kernel driver core. Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu kernel driver core. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: handle the case of pci_channel_io_frozen only in amdgpu_pci_resume","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47421","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47489","cve":"CVE-2021-47489","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix even more out of bound writes from debugfs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47489","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47551","cve":"CVE-2021-47551","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amd/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amd/amdkfd)","year":"2021","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/amdkfd: Fix kernel panic when reset failed and been triggered again","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47551","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-0185","cve":"CVE-2022-0185","aliases":[],"title":"Linux kernel (fs_context): Heap overflow in legacy filesystem parameter handling","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (fs_context)","year":"2022","cvss_score":7.8,"severity":"high","kev":true,"impact":"Heap overflow in legacy filesystem parameter handling; escapes unprivileged containers to host root [KEV]","attack_vector":"Any tenant process in a container with a user namespace","remediation":"Livepatchable; otherwise drain + reboot. Mitigate by disabling unprivileged user namespaces","references":["https://access.redhat.com/security/cve/CVE-2022-0185"],"status":"curated","fleet":{"ubiquity":"Universal - kernel 5.1 through 5.16.1; exploitable wherever unprivileged user namespaces are on, which is the default on Ubuntu GPU images","remediation_pain":"`node-reboot` - kernel upgrade; the only no-reboot mitigation is disabling unprivileged user namespaces, which breaks rootless/Podman-style tenant workflows","pain_class":"node-reboot","why_fleet_wide":"Heap overflow in `fs_context` gives a container-confined attacker full host root, demonstrated as a Kubernetes container escape on GKE/EKS/AKS-class engines - one tenant image compromises the whole node and its co-tenants"}},{"id":"CVE-2022-0330","cve":"CVE-2022-0330","aliases":[],"title":"Linux i915 GPU kernel driver (GTT TLB handling): MULTI-TENANT ISOLATION: Stale GPU TLB entries let the GPU keep reading physical pages after they were…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (GTT TLB handling)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Stale GPU TLB entries let the GPU keep reading physical pages after they were unmapped and handed to somebody else. A tenant running crafted GPU code reads whatever the host recycled those pages into - other tenants' data, or kernel memory. This is a true cross-tenant memory disclosure on Intel GPU nodes, not a crash bug.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Update the kernel and reboot; the fix forces a full TLB flush on unbind, which costs GPU unbind throughput on memory-churning workloads. Drain the node - the driver cannot be swapped under live GPU jobs. Kernel-only, no firmware or microcode.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0330","https://access.redhat.com/security/cve/CVE-2022-0330"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-0847","cve":"CVE-2022-0847","aliases":[],"title":"Linux kernel (pipe): Dirty Pipe: uninitialised pipe_buffer flags allow overwriting read-only files, incl. host binaries from…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (pipe)","year":"2022","cvss_score":7.8,"severity":"high","kev":true,"impact":"Dirty Pipe: uninitialised pipe_buffer flags allow overwriting read-only files, incl. host binaries from inside a container [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable (all major vendors shipped livepatches); otherwise drain + reboot. Highest-priority historical container-escape primitive","references":["https://access.redhat.com/security/cve/CVE-2022-0847"],"status":"curated","fleet":{"ubiquity":"Universal - every kernel 5.8+ before 5.16.11/5.15.25/5.10.102, which was the mainstream range on GPU hosts at the time","remediation_pain":"`node-reboot` - a kernel fix means a reboot of every GPU host unless the operator runs livepatch/kpatch; either way jobs must be drained first","pain_class":"node-reboot","why_fleet_wide":"An unprivileged process in a container overwrites read-only files, and the modification lands on the *host* file (page cache is shared), so a tenant overwrites host SUID binaries and takes the node"}},{"id":"CVE-2022-21821","cve":"CVE-2022-21821","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: An integer overflow reached by disassembling a corrupted fatbin gives remote code execution in the context of…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An integer overflow reached by disassembling a corrupted fatbin gives remote code execution in the context of the user running the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5334). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21821","https://github.com/NVIDIA/product-security/tree/main/2022/5334"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-190"]},{"id":"CVE-2022-2588","cve":"CVE-2022-2588","aliases":[],"title":"Linux kernel (net/sched cls_route): Use-after-free in the cls_route filter - local privilege escalation, publicly exploited in container escapes","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched cls_route)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in the cls_route filter - local privilege escalation, publicly exploited in container escapes","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot. Blacklist `cls_route` module as a stopgap","references":["https://access.redhat.com/security/cve/CVE-2022-2588"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-26358","cve":"CVE-2022-26358","aliases":[],"title":"Xen on AMD-Vi - unity map handling on device reassignment: MULTI-TENANT ISOLATION: AMD-Vi unity mappings are not correctly torn down or re-established when a device…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling on device reassignment","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: AMD-Vi unity mappings are not correctly torn down or re-established when a device moves between guests, so mappings from a previous owner can persist. On a GPU cloud this is the device-handoff problem in its purest form: the accelerator you just reassigned to a new tenant may still carry DMA reach into the previous tenant's memory.","attack_vector":"Requires device reassignment between guests - i.e. exactly what happens when you recycle a passed-through GPU from one customer to the next.","remediation":"Fixed in Xen via XSA-400. Update the hypervisor and reboot. Operationally, treat GPU reassignment as a security transition: reset the device, verify IOMMU mappings are torn down, and prefer a host reboot between tenants where your margins allow it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26358","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-26359","cve":"CVE-2022-26359","aliases":[],"title":"Xen on AMD-Vi - unity map handling: MULTI-TENANT ISOLATION: Second XSA-400 AMD-Vi unity-map issue. Stale or incorrect IOMMU mappings across…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Second XSA-400 AMD-Vi unity-map issue. Stale or incorrect IOMMU mappings across device assignment leave DMA windows open into memory the current device owner should not reach.","attack_vector":"Guest with an assigned device, particularly across reassignment.","remediation":"Fixed in Xen (XSA-400). Hypervisor update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26359","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-26360","cve":"CVE-2022-26360","aliases":[],"title":"Xen on AMD-Vi - unity map handling: MULTI-TENANT ISOLATION: Third XSA-400 AMD-Vi issue. Same class - IOMMU mappings that outlive their…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Third XSA-400 AMD-Vi issue. Same class - IOMMU mappings that outlive their justification give an assigned device cross-guest DMA reach.","attack_vector":"Guest with an assigned device.","remediation":"Fixed in Xen (XSA-400). Hypervisor update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26360","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-26361","cve":"CVE-2022-26361","aliases":[],"title":"Xen on AMD-Vi - unity map handling: MULTI-TENANT ISOLATION: Fourth XSA-400 AMD-Vi issue. Patch the set together","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - unity map handling","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Fourth XSA-400 AMD-Vi issue. Patch the set together; individually they are variants of the same broken teardown of device DMA permissions.","attack_vector":"Guest with an assigned device.","remediation":"Fixed in Xen (XSA-400). Hypervisor update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26361","https://xenbits.xen.org/xsa/advisory-400.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-27666","cve":"CVE-2022-27666","aliases":[],"title":"Linux kernel (IPsec ESP): Buffer overflow in the IPsec ESP transformation code - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (IPsec ESP)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Buffer overflow in the IPsec ESP transformation code - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-27666"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-29919","cve":"CVE-2022-29919","aliases":["INTEL-SA-00692","CVE-2022-45112","INTEL-SA-00846","CVE-2023-31271","INTEL-SA-00953","CVE-2024-23489"],"title":"Intel Virtual RAID on CPU (VROC) software before 7.7.6.1003, with follow-on issues through 8.6.0.1191: Use-after-free in the VROC software giving an authenticated local user privilege escalation…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Virtual RAID on CPU (VROC) software before 7.7.6.1003, with follow-on issues through 8.6.0.1191","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in the VROC software giving an authenticated local user privilege escalation, followed by a run of access-control, default-permission and path-traversal escalations in later VROC releases. VROC is the software RAID layer sitting directly on NVMe on Xeon platforms - it is in the storage path for boot volumes and local scratch on many server designs. A local escalation here is a tenant-to-root path on any node where VROC is installed, and root on the node is the gateway to the whole firmware stack below it. The recurring pattern across four advisories is the useful signal: this component has a weak security history and should not be left installed where it is not needed.","attack_vector":"Authenticated local user on the host with the VROC software installed.","remediation":"Update VROC to 8.6.0.1191 or later (or the latest available for your platform) via the OEM's storage software package - Dell, HPE, Lenovo and Supermicro all redistribute it. Software update, so no firmware flash, but a reboot is typical. The better move for most GPU fleets: if you are not actually using VROC RAID, uninstall it rather than patch it - it is frequently present in golden images purely because it came with the platform driver bundle, and each release has brought a new local escalation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29919","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00692.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00953.html"],"status":"curated"},{"id":"CVE-2022-31606","cve":"CVE-2022-31606","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Missing data validation lets a basic user cause an out-of-bounds access in kernel mode through DxgkDdiEscape…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing data validation lets a basic user cause an out-of-bounds access in kernel mode through DxgkDdiEscape, reaching privilege escalation and kernel information disclosure. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31606","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31607","cve":"CVE-2022-31607","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Missing input validation in nvidia.ko lets a basic local user escalate privileges, tamper with data, or leak…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing input validation in nvidia.ko lets a basic local user escalate privileges, tamper with data, or leak a limited amount of kernel information - a genuine local root path from inside a GPU container. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5383. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31607","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31608","cve":"CVE-2022-31608","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An optional D-Bus configuration file shipped with the Linux driver leaves protected D-Bus endpoints reachable…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An optional D-Bus configuration file shipped with the Linux driver leaves protected D-Bus endpoints reachable by a basic local user, which chains to code execution and privilege escalation on the host. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5383. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31608","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-281"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31609","cve":"CVE-2022-31609","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: The vGPU plugin lets a guest VM allocate resources it is not authorised to hold…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The vGPU plugin lets a guest VM allocate resources it is not authorised to hold, breaking the resource partition between tenants and reaching information disclosure and data tampering. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. This is a direct authorisation failure on the guest/host boundary, not a memory-safety accident - the guest simply asks for something it should not get and the plugin agrees.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5383. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31609","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-285"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31610","cve":"CVE-2022-31610","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds write in the kernel mode layer gives a basic local user code execution in kernel context…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds write in the kernel mode layer gives a basic local user code execution in kernel context - full compromise of the Windows GPU node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31610","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31617","cve":"CVE-2022-31617","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read in the kernel mode layer chains to code execution and privilege escalation on the node.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds read in the kernel mode layer chains to code execution and privilege escalation on the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31617","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-32250","cve":"CVE-2022-32250","aliases":[],"title":"Linux kernel (netfilter): Use-after-free write in the netfilter subsystem - privilege escalation to root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free write in the netfilter subsystem - privilege escalation to root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-32250"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-34325","cve":"CVE-2022-34325","aliases":["INSYDE-SA-2022057"],"title":"Insyde InsydeH2O (StorageSecurityCommandDxe SMI input buffer, DMA TOCTOU): Highest-scored DMA entry in the 2022 batch. StorageSecurityCommandDxe issues TCG/Opal security commands to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (StorageSecurityCommandDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Highest-scored DMA entry in the 2022 batch. StorageSecurityCommandDxe issues TCG/Opal security commands to self-encrypting drives, so this driver reaches SED authentication material. An attacker winning the race gets SMRAM corruption plus a position inside the code path that unlocks encrypted drives - which is precisely the control a GPU cloud relies on to claim tenant data is protected at rest between leases.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in the kernel releases named in the advisory (Insyde does not enumerate per-kernel versions for this one).  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34325","https://www.insyde.com/security-pledge/SA-2022057"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-34670","cve":"CVE-2022-34670","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Truncation when casting to a smaller primitive loses data inside the kernel handler, giving an unprivileged…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Truncation when casting to a smaller primitive loses data inside the kernel handler, giving an unprivileged user a denial of service or an information leak out of kernel memory. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34670","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-197"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34918","cve":"CVE-2022-34918","aliases":[],"title":"Linux kernel (nf_tables): Heap overflow in nft_set_elem_init() - local root, weaponised inside containers","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Heap overflow in nft_set_elem_init() - local root, weaponised inside containers","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-34918"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-3650","cve":"CVE-2022-3650","aliases":[],"title":"Ceph: ceph-crash.service local privilege escalation to root plus privileged crash-dump disclosure","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"ceph-crash.service local privilege escalation to root plus privileged crash-dump disclosure","attack_vector":"Local","remediation":"Data-plane: package update on every OSD/MON host","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3650"],"status":"curated"},{"id":"CVE-2022-36763","cve":"CVE-2022-36763","aliases":["GHSA-xvv8-66cq-prwr"],"title":"EDK II SecurityPkg (Tcg2Dxe, Tcg2MeasureGptTable): A crafted GPT partition table overflows the heap inside the very code that is supposed to measure the disk…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II SecurityPkg (Tcg2Dxe, Tcg2MeasureGptTable)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A crafted GPT partition table overflows the heap inside the very code that is supposed to measure the disk layout into the TPM. Two consequences that matter for a fleet: the attacker gets code execution during boot, and the compromise happens inside the measured-boot machinery itself, so the PCR values that downstream attestation trusts are produced by code the attacker already controls. Remote attestation of the node becomes meaningless while telling you everything is fine.","attack_vector":"Anyone who can present a disk with an attacker-controlled GPT to the node - a tenant who had the box before you and wrote to a local drive, a removable device, or an iSCSI/SAN LUN whose contents the attacker influences. Requires the node to boot with that disk attached.","remediation":"OEM BIOS update; the upstream edk2 fix predates public disclosure by over a year, so most current server BIOS lines already carry it - confirm against the OEM release notes for your exact platform generation rather than assuming. Flash + reboot per node. No config workaround inside firmware; operationally, wiping and re-partitioning tenant disks between leases reduces exposure but does not close the bug.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36763","https://github.com/tianocore/edk2/security/advisories/GHSA-xvv8-66cq-prwr"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-36765","cve":"CVE-2022-36765","aliases":["GHSA-ch4w-v7m3-g8wx"],"title":"EDK II MdePkg (CreateHob, HOB list construction): An integer overflow in the routine that allocates Hand-Off Blocks lets an attacker read and write outside the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II MdePkg (CreateHob, HOB list construction)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An integer overflow in the routine that allocates Hand-Off Blocks lets an attacker read and write outside the HOB list. HOBs are the structure PEI uses to describe memory, security state and platform configuration to DXE, so corrupting them is a way to lie to the rest of the boot about what is trusted - including the memory ranges Secure Boot and SMM protections are supposed to cover. Result is early-boot code execution with the firmware's own privileges.","attack_vector":"Requires influence over PEI-phase input on the node - typically a local attacker with the ability to feed the early boot path (attacker-controlled firmware volume content, a malicious capsule, or an earlier compromise that persists into PEI). Not remotely reachable.","remediation":"OEM BIOS update - the fix is a core MdePkg change, so every IBV downstream of edk2 had to rebase it, and coverage in shipped server images varies by platform generation. Flash + reboot per node. No configuration mitigates it; the only operational lever is keeping firmware-write paths (capsule update, SPI programming) locked down so an attacker cannot get to PEI in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36765","https://github.com/tianocore/edk2/security/advisories/GHSA-ch4w-v7m3-g8wx"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-37459","cve":"CVE-2022-37459","aliases":["AMP-SB-0004","Retbleed on Arm","Arm KA005138"],"title":"Ampere Altra before 1.08g and Altra Max before 2.05a - return address prediction: An attacker can control return-address predictions and steer speculative execution into a chosen gadget, then…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra before 1.08g and Altra Max before 2.05a - return address prediction","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An attacker can control return-address predictions and steer speculative execution into a chosen gadget, then read the result out of the cache. Same practical outcome as Spectre-v2 on these parts: cross-privilege and, where cores are shared, cross-tenant reads of memory the attacker has no rights to. Relevant to anyone running Altra as a CPU head node in front of GPU workers, since that node typically holds cluster credentials and customer data in flight.","attack_vector":"Unprivileged local code on an affected Altra / Altra Max part, targeting a victim on the same core or the same branch predictor structures. Local only.","remediation":"Update Altra firmware to 1.08g / Altra Max 2.05a or later from the board OEM, plus the corresponding kernel mitigations. Flash + reboot + drain. Speculation mitigations on Arm cost real throughput on syscall- and context-switch-heavy paths, so benchmark your actual serving stack rather than accepting the vendor's number. As with every predictor-sharing bug, not co-scheduling untrusted tenants on the same physical core is the mitigation that does not degrade over time.","references":["https://amperecomputing.com/products/security-bulletins/retbleed.html","https://developer.arm.com/documentation/ka005138/1-0/","https://nvd.nist.gov/vuln/detail/CVE-2022-37459"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-4139","cve":"CVE-2022-4139","aliases":[],"title":"Linux i915 GPU kernel driver (TLB invalidation): MULTI-TENANT ISOLATION: An incorrect TLB flush in i915 leaves the GPU able to reach memory it should no…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (TLB invalidation)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An incorrect TLB flush in i915 leaves the GPU able to reach memory it should no longer see, producing random memory corruption or leakage across contexts. Same family as the earlier GTT TLB bug and with the same consequence for a shared GPU node: one tenant's GPU reads or corrupts pages that now belong to someone else.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Kernel update and reboot. Drain the node first. No firmware or microcode component; the fix is entirely in the driver's invalidation logic.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4139","https://access.redhat.com/security/cve/CVE-2022-4139"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-42260","cve":"CVE-2022-42260","aliases":[],"title":"NVIDIA vGPU software - guest driver (inside tenant VM): A D-Bus configuration file shipped with the Linux vGPU guest driver leaves protected endpoints reachable, so…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - guest driver (inside tenant VM)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A D-Bus configuration file shipped with the Linux vGPU guest driver leaves protected endpoints reachable, so an unauthorised user inside the guest VM reaches code execution and privilege escalation within that VM. Scope is confined to the guest VM, so this is a tenant-internal privilege problem rather than a break of your isolation boundary - but it is the first half of a chain if a vGPU Manager bug is also unpatched.","attack_vector":"An unprivileged user inside a guest VM that has a vGPU attached. You may not control these VMs at all if tenants bring their own images.","remediation":"Ship the fixed guest driver (bulletin 5415) to tenant VMs. Cost: low on your side, but you often cannot force it - if tenants own their guest images, the realistic control is a supported-driver-version policy plus refusing host attach below a floor version. No host drain required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42260","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-281"]},{"id":"CVE-2022-42261","cve":"CVE-2022-42261","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: An unvalidated input index in the vGPU plugin produces a host-side buffer overrun…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An unvalidated input index in the vGPU plugin produces a host-side buffer overrun, reaching data tampering, information disclosure or denial of service from inside a guest. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. A buffer overrun in host context driven by a guest-controlled index is the shape of a hypervisor escape; treat it as high priority even though NVIDIA scores it local.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5415. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42261","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42267","cve":"CVE-2022-42267","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read in the Windows driver escalates to code execution and full node compromise from an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds read in the Windows driver escalates to code execution and full node compromise from an ordinary user account. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42267","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-345"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42268","cve":"CVE-2022-42268","aliases":[],"title":"Isaac Sim / Omniverse: Local privesc (improper input validation in config)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Isaac Sim / Omniverse","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (improper input validation in config)","attack_vector":"Local user on the workstation/node","remediation":"Upgrade Omniverse packages; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42268","https://github.com/NVIDIA/product-security/tree/main/2023/5418"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-74"]},{"id":"CVE-2022-42270","cve":"CVE-2022-42270","aliases":[],"title":"Jetson AGX Xavier / Orin bootloader: Local privesc (stack overflow in bootloader)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson AGX Xavier / Orin bootloader","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (stack overflow in bootloader)","attack_vector":"Local operator / physical","remediation":"Flash bootloader firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42270","https://github.com/NVIDIA/product-security/tree/main/2023/5442"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2022-42274","cve":"CVE-2022-42274","aliases":[],"title":"DGX-2 BMC: RCE on BMC (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE on BMC (buffer overflow)","attack_vector":"Network-adjacent mgmt-LAN attacker","remediation":"Flash DGX-2 BMC firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42274","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42332","cve":"CVE-2022-42332","aliases":["XSA-427"],"title":"Xen (shadow paging): x86 shadow plus log-dirty mode use-after-free - guest to host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (shadow paging)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"x86 shadow plus log-dirty mode use-after-free - guest to host","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation. Shadow paging is used during live migration, so this is reachable in normal operations","references":["https://xenbits.xen.org/xsa/advisory-427.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-42973","cve":"CVE-2022-42973","aliases":["SEVD-2022-256-01"],"title":"APC Easy UPS Online Monitoring Software - embedded database credentials: Hardcoded credentials let any local user connect to the software's database and escalate. The database holds…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Easy UPS Online Monitoring Software - embedded database credentials","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Hardcoded credentials let any local user connect to the software's database and escalate. The database holds the site's UPS inventory and the credentials used to command shutdowns, so this converts a low-privilege foothold on one Windows host into control of the power-shutdown path.","attack_vector":"Local access to the machine running the monitoring software.","remediation":"Software upgrade. Hardcoded credentials mean the value is public once the advisory ships, so also confirm the database is not listening beyond localhost. Cheap to fix, and worth doing in the same window as the two RCEs above.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42973"],"status":"curated"},{"id":"CVE-2022-4318","cve":"CVE-2022-4318","aliases":[],"title":"CRI-O: Crafted environment variable injects arbitrary lines into /etc/passwd","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted environment variable injects arbitrary lines into /etc/passwd","attack_vector":"Malicious image or tenant-controlled pod env","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4318"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-48632","cve":"CVE-2022-48632","aliases":[],"title":"Linux kernel i2c-mlxbf (BlueField DPU I2C/SMBus controller): memcpy() is called in a loop with no upper bound on the operation length while the index keeps incrementing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel i2c-mlxbf (BlueField DPU I2C/SMBus controller)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"memcpy() is called in a loop with no upper bound on the operation length while the index keeps incrementing - a kernel stack overflow on the BlueField DPU itself. The DPU Arm cores run your infrastructure services (telemetry, storage emulation, security agents) alongside tenant-facing datapaths, so stack corruption there is compromise of the control point you deployed the DPU to be.","attack_vector":"Local on the DPU Arm side - a process with access to the SMBus/I2C interface. Relevant when tenants or third-party agents get any foothold on the DPU OS.","remediation":"Upgrade the DPU Arm-side kernel: 6.0 or a stable backport (5.10.146, 5.15.71, 5.19.12). In practice this arrives as a BFB bundle re-image or a DOCA/BFOS package upgrade on the DPU, followed by a DPU reset - which drops the host's network and storage while it reboots. Plan it as a per-node maintenance window, not a live update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48632","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48632.json"],"status":"curated"},{"id":"CVE-2022-48662","cve":"CVE-2022-48662","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the GEM context link being manipulated outside reference protection, which i915_perf then follows into freed memory.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48662","https://git.kernel.org/stable/c/713fa3e4591f65f804bdc88e8648e219fabc9ee1"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-48932","cve":"CVE-2022-48932","aliases":["net/mlx5 DR slab-out-of-bounds in mlx5_cmd_dr_create_fte"],"title":"Linux kernel mlx5_core software steering (fs_dr): Adding a flow rule with 32 destinations overflows an undersized action buffer in the software steering path…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core software steering (fs_dr)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"Adding a flow rule with 32 destinations overflows an undersized action buffer in the software steering path - there was no check on the action count. Multi-destination rules are how you build mirroring and multicast replication in the NIC, so an operator or automation system building a large fan-out rule corrupts kernel slab memory.","attack_vector":"Local with network-configuration privilege (CAP_NET_ADMIN), including inside a user namespace on many setups. Triggered by installing a steering rule with a large destination list.","remediation":"Upgrade the host kernel to 5.17 or the 5.16.12 stable backport. Rolling reboot. Interim: cap the number of destinations your SDN controller or CNI is allowed to program per rule.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48932","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-48932.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-48979","cve":"CVE-2022-48979","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix array index out of bound error in DCN32 DML","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48979","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-48990","cve":"CVE-2022-48990","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix use-after-free during gpu recovery","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48990","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49025","cve":"CVE-2022-49025","aliases":["net/mlx5e use-after-free reverting termination table"],"title":"Linux kernel mlx5_core eswitch offloads (termination tables): TENANT ISOLATION: adding a multi-destination eswitch rule that partially fails leaves a stale…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch offloads (termination tables)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: adding a multi-destination eswitch rule that partially fails leaves a stale termination-table pointer, and releasing the rule triggers a use-after-free in the eswitch - the exact subsystem that enforces which VF sees which traffic. Corruption here is a plausible route to host kernel control from a workload that only has network-namespace privilege.","attack_vector":"A local user who can add and delete tc flower rules with CAP_NET_ADMIN - obtainable in an unprivileged user namespace (unshare -Urn), so reachable from inside many container runtimes, not just from host root.","remediation":"Upgrade the host kernel to 6.1 or a stable backport (5.4.226, 5.10.158, 5.15.82, 6.0.12). Rolling reboot of the fleet. Interim: disable unprivileged user namespaces (kernel.unprivileged_userns_clone=0 / user.max_user_namespaces=0) where your container runtime does not need them - a sysctl config change, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49025","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2022/CVE-2022-49025.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49203","cve":"CVE-2022-49203","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/display: Fix double free during GPU reset on DC streams","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49203","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49530","cve":"CVE-2022-49530","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A use-after-free in the amdgpu power management (SMU/powerplay). Freed kernel memory is reachable again…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu power management (SMU/powerplay). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/pm: fix double free in si_parse_power_table()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49530","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49969","cve":"CVE-2022-49969","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: clear optc underflow before turn off odm clock","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49969","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50035","cve":"CVE-2022-50035","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix use-after-free on amdgpu_bo_list mutex","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50035","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50303","cve":"CVE-2022-50303","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: A double free in the amdkfd (KFD compute driver, /dev/kfd). The same allocation is…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A double free in the amdkfd (KFD compute driver, /dev/kfd). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amdkfd: Fix double release compute pasid","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50303","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50354","cve":"CVE-2022-50354","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix kfd_process_device_init_vm error handling","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50354","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50393","cve":"CVE-2022-50393","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the DRM scheduler /…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: SDMA update use unlocked iterator","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50393","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50528","cve":"CVE-2022-50528","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":7.8,"severity":"high","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: Fix memory leakage","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50528","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0179","cve":"CVE-2023-0179","aliases":[],"title":"Linux kernel (netfilter): Integer overflow in nft_payload_copy_vlan - stack leak plus local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow in nft_payload_copy_vlan - stack leak plus local privilege escalation","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-0179"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0182","cve":"CVE-2023-0182","aliases":[],"title":"GPU Display Driver: Local privesc to host root (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (kernel buffer overflow)","attack_vector":"Any tenant with a container holding /dev/nvidia*","remediation":"Driver upgrade; drain + reboot node, evict tenant workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0182","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0266","cve":"CVE-2023-0266","aliases":[],"title":"Linux kernel (ALSA): Use-after-free in snd_ctl_elem_read - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ALSA)","year":"2023","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in snd_ctl_elem_read - local privilege escalation [KEV]","attack_vector":"Local user with access to /dev/snd","remediation":"Livepatchable; otherwise drain + reboot. Low exposure on headless GPU nodes - blacklist sound modules","references":["https://access.redhat.com/security/cve/CVE-2023-0266"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-1017","cve":"CVE-2023-1017","aliases":[],"title":"TPM 2.0 reference implementation: Out-of-bounds write in `CryptParameterDecryption` — code execution inside the TPM or a bricked TPM. If the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TPM 2.0 reference implementation","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in `CryptParameterDecryption` — code execution inside the TPM or a bricked TPM. If the TPM is the root of trust for attestation, a compromised TPM invalidates every measured-boot claim the cloud makes to its tenants","attack_vector":"Local, low privilege","remediation":"TPM firmware update from the TPM/platform vendor; TPM firmware updates frequently clear the TPM, which invalidates sealed keys and any disk encryption bound to PCRs — this is why fleets skip it","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-1017"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20548","cve":"CVE-2023-20548","aliases":[],"title":"AMD Secure Processor - TOCTOU race: MULTI-TENANT ISOLATION: A time-of-check-to-time-of-use race in the ASP lets an attacker swap a value between…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - TOCTOU race","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A time-of-check-to-time-of-use race in the ASP lets an attacker swap a value between validation and use, corrupting secure-processor memory. Winning the race costs integrity, confidentiality or availability of the ASP depending on what gets corrupted - and the ASP is the component vouching for the node's confidential-computing posture.","attack_vector":"Local. Requires the attacker to run concurrently with the ASP operation and to be able to modify the memory being checked, i.e. host-privileged code.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20548","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20555","cve":"CVE-2023-20555","aliases":[],"title":"AMD SMM - memory corruption (AMD-SB-4003): MULTI-TENANT ISOLATION: Memory corruption reachable in System Management Mode. Same class as the rest of the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SMM - memory corruption (AMD-SB-4003)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory corruption reachable in System Management Mode. Same class as the rest of the SMM cluster - a ring-0 attacker escalates into the one execution context that no hypervisor, kernel or EDR can observe, and the foothold survives OS reinstallation.","attack_vector":"Local, ring-0 privilege required.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20555","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-20598","cve":"CVE-2023-20598","aliases":[],"title":"AMD Radeon Graphics driver - IOCTL granting arbitrary I/O port and physical memory access: MULTI-TENANT ISOLATION: Improper privilege management in the AMD graphics driver lets an authenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Radeon Graphics driver - IOCTL granting arbitrary I/O port and physical memory access","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper privilege management in the AMD graphics driver lets an authenticated attacker craft an IOCTL that gives I/O control over arbitrary hardware ports or physical memory. This is about as broad as a driver bug gets: arbitrary physical memory access from a user-issued IOCTL is a complete bypass of kernel memory protection, so any workload holding the GPU device node can read every other tenant's memory and take the host. AMD's PSP driver component (AMDPSP) was the affected surface.","attack_vector":"Local, authenticated, via a GPU driver IOCTL - i.e. reachable by anything that has the GPU device handed to it, which on a GPU cloud is every tenant container.","remediation":"Update the AMD graphics driver package and reload the driver or reboot the node. Driver-speed fix - no BIOS, no VBIOS - so there is no excuse for leaving it. Treat as top priority on any node where untrusted workloads get a GPU device node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20598","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-21987","cve":"CVE-2023-21987","aliases":[],"title":"Oracle VirtualBox: Core component flaw allowing a low-privileged guest user to take over the host VirtualBox installation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Oracle VirtualBox","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Core component flaw allowing a low-privileged guest user to take over the host VirtualBox installation","attack_vector":"Tenant VM guest","remediation":"VirtualBox update + VM restart. VirtualBox is not a production hypervisor - its presence on a neocloud node is itself the finding","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-21987"],"status":"curated"},{"id":"CVE-2023-25505","cve":"CVE-2023-25505","aliases":[],"title":"NVIDIA DGX BMC (IPMI handler): Buffer overflow in the IPMI handler of the NVIDIA DGX BMC — code execution on the BMC of a GPU node","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (IPMI handler)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Buffer overflow in the IPMI handler of the NVIDIA DGX BMC — code execution on the BMC of a GPU node","attack_vector":"Local / IPMI","remediation":"DGX BMC firmware update through NVIDIA's own bundle; DGX firmware bundles are monolithic, so this pulls in unrelated component updates and a longer maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25505"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25515","cve":"CVE-2023-25515","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): The driver parses unexpected untrusted data, reaching code execution, privilege escalation and information…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver parses unexpected untrusted data, reaching code execution, privilege escalation and information disclosure from an unprivileged local account on either OS. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5468. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25515","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-822"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-25519","cve":"CVE-2023-25519","aliases":[],"title":"NVIDIA BlueField: Privilege escalation from incorrect user management on the DPU","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA BlueField","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Privilege escalation from incorrect user management on the DPU","attack_vector":"Local","remediation":"BlueField DOCA/BFB update; requires draining the node because the DPU carries the tenant's network path","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25519"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-286"]},{"id":"CVE-2023-25527","cve":"CVE-2023-25527","aliases":[],"title":"DGX H100 BMC: Kernel memory corruption","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Kernel memory corruption -> arbitrary code exec on BMC","attack_vector":"Authenticated local BMC access","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-119"]},{"id":"CVE-2023-2598","cve":"CVE-2023-2598","aliases":[],"title":"Linux kernel (io_uring): io_uring fixed-buffer registration gives out-of-bounds access to physical memory - full host compromise from…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (io_uring)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"io_uring fixed-buffer registration gives out-of-bounds access to physical memory - full host compromise from an unprivileged process","attack_vector":"Any tenant process in a container with io_uring enabled","remediation":"Livepatchable; otherwise drain + reboot. Strategic answer for a neocloud is to disable io_uring in the default container seccomp profile (`kernel.io_uring_disabled=2` on 6.6+)","references":["https://access.redhat.com/security/cve/CVE-2023-2598"],"status":"curated","fleet":{"ubiquity":"Very common - io_uring is on by default in modern kernels and is heavily used by high-throughput data loaders on GPU nodes","remediation_pain":"`node-reboot` - kernel 6.4-rc1+; the practical stopgap is disabling io_uring via sysctl, which degrades storage throughput for training jobs","pain_class":"node-reboot","why_fleet_wide":"Out-of-bounds physical-memory access from `IORING_REGISTER_BUFFERS` gives a low-privilege local user full root; from inside a container with io_uring allowed, that is a host takeover on a shared GPU node"}},{"id":"CVE-2023-31019","cve":"CVE-2023-31019","aliases":[],"title":"GPU Display Driver (Windows): Local privesc via named-pipe server access","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc via named-pipe server access","attack_vector":"Local low-priv user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-284"]},{"id":"CVE-2023-31248","cve":"CVE-2023-31248","aliases":[],"title":"Linux kernel (nf_tables): Use-after-free in nft_chain_lookup_byid() - local root (Pwn2Own Vancouver chain)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in nft_chain_lookup_byid() - local root (Pwn2Own Vancouver chain)","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-31248"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-31324","cve":"CVE-2023-31324","aliases":[],"title":"AMD Secure Processor - XGMI Trusted Agent (TOCTOU): MULTI-TENANT ISOLATION: A TOCTOU race in the ASP's XGMI Trusted Agent lets an attacker modify…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - XGMI Trusted Agent (TOCTOU)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A TOCTOU race in the ASP's XGMI Trusted Agent lets an attacker modify Infinity-Fabric/XGMI link commands after they are validated but before they execute. XGMI is the coherent interconnect binding MI-series GPUs together in a node, so tampering with its trusted-agent commands is tampering with the fabric that carries other tenants' model traffic between accelerators.","attack_vector":"Local, host-privileged, and specific to multi-GPU platforms that actually use XGMI - which is every MI200/MI250/MI300 node in an 8-GPU configuration.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. This one is squarely an Instinct-node issue, not a generic EPYC one: prioritise it on MI250/MI300 hosts where XGMI is carrying real inter-GPU traffic.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31324","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-32233","cve":"CVE-2023-32233","aliases":[],"title":"Linux kernel (nf_tables): Use-after-free in nf_tables anonymous-set batch processing - unprivileged local user to root, public exploit","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in nf_tables anonymous-set batch processing - unprivileged local user to root, public exploit","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot. Disable unprivileged user namespaces to blunt it","references":["https://access.redhat.com/security/cve/CVE-2023-32233"],"status":"curated","fleet":{"ubiquity":"Universal - `nf_tables` is enabled by default in most distributions and every kernel through 6.3.1 is affected","remediation_pain":"`node-reboot` - kernel upgrade to 6.3.2+; mitigation is blocking `CAP_NET_ADMIN`/user namespaces, which many tenant workloads legitimately need","pain_class":"node-reboot","why_fleet_wide":"Use-after-free in nf_tables batch processing gives arbitrary kernel read/write and root from an unprivileged local user - in a container with a user namespace this is a straight escape to the GPU host"}},{"id":"CVE-2023-3269","cve":"CVE-2023-3269","aliases":[],"title":"Linux kernel (mm VMA): StackRot: privilege escalation via non-RCU-protected VMA traversal","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (mm VMA)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"StackRot: privilege escalation via non-RCU-protected VMA traversal; affects 6.1-6.4","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-3269"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-3390","cve":"CVE-2023-3390","aliases":[],"title":"Linux kernel (nf_tables): UAF in nft_set_lookup_global after mixed named/anonymous set batches - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"UAF in nft_set_lookup_global after mixed named/anonymous set batches - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-3390"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-34195","cve":"CVE-2023-34195","aliases":["INSYDE-SA-2023052"],"title":"Insyde InsydeH2O (SystemFirmwareManagementRuntimeDxe, GetImage method): The firmware reads a runtime UEFI variable called GetImageProgress and then calls it as a function pointer.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SystemFirmwareManagementRuntimeDxe, GetImage method)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The firmware reads a runtime UEFI variable called GetImageProgress and then calls it as a function pointer. An attacker sets that variable from the OS to point at code they control and the firmware jumps to it during the DXE phase. This is a firmware-update-service driver, so the attacker ends up executing inside the machinery responsible for validating the next BIOS image - a direct route to a persistent, self-reinstalling firmware implant on a GPU node.","attack_vector":"Local admin/root on the host OS with UEFI variable write access, then a reboot or a firmware-management call that reaches GetImage.","remediation":"OEM BIOS update carrying the fixed Insyde kernel (5.0-5.5 affected). Firmware flash, one reboot per node. No configuration mitigates it. Interim hardening: restrict OS-side UEFI variable writes, and where the platform supports it verify that capsule updates require a signed payload so an implant cannot re-flash itself through the same service.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34195","https://www.insyde.com/security-pledge/SA-2023052"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34319","cve":"CVE-2023-34319","aliases":["XSA-432"],"title":"Linux (Xen netback): Buffer overrun in netback due to an unusual packet - guest attacks dom0","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux (Xen netback)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Buffer overrun in netback due to an unusual packet - guest attacks dom0","attack_vector":"Tenant VM guest","remediation":"dom0 kernel patch; livepatchable, otherwise dom0 reboot (evacuates every guest on the host)","references":["https://xenbits.xen.org/xsa/advisory-432.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-34322","cve":"CVE-2023-34322","aliases":["XSA-438"],"title":"Xen (64-bit PV): Top-level shadow reference dropped too early for 64-bit PV guests - privilege escalation to host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (64-bit PV)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Top-level shadow reference dropped too early for 64-bit PV guests - privilege escalation to host","attack_vector":"Tenant VM guest (PV)","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-438.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-34325","cve":"CVE-2023-34325","aliases":["XSA-443"],"title":"Xen (libfsimage/pygrub): Multiple vulnerabilities in libfsimage disk handling - a hostile guest disk image compromises the toolstack…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (libfsimage/pygrub)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Multiple vulnerabilities in libfsimage disk handling - a hostile guest disk image compromises the toolstack in dom0","attack_vector":"Tenant-supplied disk image","remediation":"Patch libfsimage and run pygrub de-privileged (see also XSA-508); toolstack restart, no full host reboot","references":["https://xenbits.xen.org/xsa/advisory-443.html"],"status":"curated"},{"id":"CVE-2023-34332","cve":"CVE-2023-34332","aliases":["AMI-SA-2023010"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC): Untrusted pointer dereference in the BMC that a low-privileged actor can turn into code execution or a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Untrusted pointer dereference in the BMC that a low-privileged actor can turn into code execution or a controller crash. As a second-stage bug it is how an attacker who already got a foothold - a read-only monitoring account, a low-privilege Redfish user, or a partial exploit of one of the network bugs - upgrades to full BMC control and firmware persistence.","attack_vector":"AMI's CVSS vector scores this as local access with low privileges required, while AMI's own prose calls it reachable from the local network; treat it as reachable by anyone who already holds a low-privilege position on or adjacent to the BMC. In a fleet, the realistic precondition is a leaked low-tier BMC credential - which is common, because BMC passwords are frequently shared across a whole rack or SKU by the provisioning system.","remediation":"Firmware flash to SPx_12.7 / SPx_13.6, out-of-band per node, ODM-gated. Alongside the flash, the cheap wins are config-only: give every BMC a unique password (kill any shared default from the deployment template), delete unused BMC accounts, and drop any monitoring account down to the minimum Redfish role.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023010.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34332"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34853","cve":"CVE-2023-34853","aliases":[],"title":"Supermicro X12DPG-QR BIOS 1.4b: Control-flow hijack inside platform firmware, driven by an NVRAM variable that host-side privileged code can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro X12DPG-QR BIOS 1.4b","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Control-flow hijack inside platform firmware, driven by an NVRAM variable that host-side privileged code can set. The attacker moves from OS root into firmware-privileged execution and can establish a UEFI-resident implant. On a X12DPG-QR - a dual-socket GPU platform board - the payoff is a persistent presence on an accelerator node that reimaging does not touch and that the operator has no host-side way to detect. A buffer overflow reachable by manipulating the SmcSecurityEraseSetupVar UEFI variable, i.e. an NVRAM variable that firmware trusts without validating its contents.","attack_vector":"Local, host-side privileged code that can write UEFI variables. On Linux that means root with access to efivarfs. This is the standard escalation available to anyone who has rented the metal or otherwise obtained host root.","remediation":"BIOS flash to a fixed image per Supermicro's August 2023 BIOS advisory, delivered as the X12DPG-QR BIOS package. Requires a host reboot. A partial hardening step that costs nothing: restrict or remove write access to efivarfs in tenant-facing bare-metal images, which raises the bar for any NVRAM-variable attack, not just this one. It does not substitute for the flash, because a tenant with root can usually reach the variable store another way.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34853","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/34xxx/CVE-2023-34853.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-35001","cve":"CVE-2023-35001","aliases":[],"title":"Linux kernel (nf_tables): Stack out-of-bounds read/write in nft_byteorder_eval() - local root (Pwn2Own)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Stack out-of-bounds read/write in nft_byteorder_eval() - local root (Pwn2Own)","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-35001"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-4004","cve":"CVE-2023-4004","aliases":[],"title":"Linux kernel (netfilter pipapo): UAF from improper element removal in nft_pipapo_remove() - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter pipapo)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"UAF from improper element removal in nft_pipapo_remove() - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4004"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-4147","cve":"CVE-2023-4147","aliases":[],"title":"Linux kernel (nf_tables): UAF adding a rule with NFTA_RULE_CHAIN_ID - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"UAF adding a rule with NFTA_RULE_CHAIN_ID - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4147"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-43068","cve":"CVE-2023-43068","aliases":[],"title":"Dell SmartFabric Storage Software (restricted shell in SSH): TENANT ISOLATION: OS command injection escaping the restricted shell of the NVMe-oF fabric controller, from…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Dell SmartFabric Storage Software (restricted shell in SSH)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: OS command injection escaping the restricted shell of the NVMe-oF fabric controller, from an authenticated remote user. Escaping the restricted shell gives full control of the appliance that enforces which host can attach to which NVMe namespace — the storage equivalent of owning the fabric's ACLs. Related CLI and path-traversal issues: CVE-2023-43069, CVE-2023-43070, CVE-2023-4401.","attack_vector":"Authenticated remote user with SSH access to SmartFabric Storage Software v1.4 or earlier.","remediation":"Upgrade SmartFabric Storage Software past v1.4 — appliance software upgrade plus restart. Restrict SSH to a management bastion, and audit the NVMe-oF zoning/namespace-masking configuration afterwards rather than assuming the patch is sufficient.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-43068","https://nvd.nist.gov/vuln/detail/CVE-2023-43069","https://nvd.nist.gov/vuln/detail/CVE-2023-4401"],"status":"curated"},{"id":"CVE-2023-4623","cve":"CVE-2023-4623","aliases":[],"title":"Linux kernel (net/sched hfsc): Use-after-free in sch_hfsc qdisc - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched hfsc)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in sch_hfsc qdisc - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4623"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-4692","cve":"CVE-2023-4692","aliases":[],"title":"GRUB2 (NTFS filesystem parser): Out-of-bounds write parsing a crafted NTFS volume. Relevant to any node that dual-boots, mounts a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (NTFS filesystem parser)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write parsing a crafted NTFS volume. Relevant to any node that dual-boots, mounts a Windows-formatted staging volume, or is handed a raw disk between tenants - the NTFS parser runs before anything verifies the disk's provenance.","attack_vector":"An attacker-supplied NTFS volume attached to the node, including via BMC virtual media.","remediation":"grub2 package update + reboot per node. Where you never need NTFS, building GRUB without the module is a permanent fix rather than a patch treadmill.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4692","https://access.redhat.com/security/cve/CVE-2023-4692"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-4911","cve":"CVE-2023-4911","aliases":[],"title":"glibc (ld.so): Looney Tunables: buffer overflow in the ld.so GLIBC_TUNABLES parser - local root on default installs","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"glibc (ld.so)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Looney Tunables: buffer overflow in the ld.so GLIBC_TUNABLES parser - local root on default installs; used by Kinsing cloud crypto-mining crews","attack_vector":"Local user, incl. any shell inside a container that shares the host glibc","remediation":"Package update; every long-running process must be restarted to pick up the new loader, which in practice means a rolling node restart or reboot for full coverage","references":["https://access.redhat.com/security/cve/CVE-2023-4911"],"status":"curated","fleet":{"ubiquity":"Universal - glibc is in essentially every base image and on every host; Qualys got root on default Fedora, Ubuntu 22.04/23.04 and Debian 12/13","remediation_pain":"`node-drain` for the host glibc (SUID binaries must be re-executed) plus a **rebuild of every container image** in the fleet - the pain is the image fan-out, not the host","pain_class":"node-drain","why_fleet_wide":"Buffer overflow in `GLIBC_TUNABLES` parsing gives local root from any SUID binary, so it converts every low-privilege foothold - in a container or on the host - into root, across every image and host simultaneously"}},{"id":"CVE-2023-5058","cve":"CVE-2023-5058","aliases":["VU#811862"],"title":"Phoenix SecureCore Technology 4 (boot splash screen image parsing): The firmware parses a user-supplied boot logo image without validating it, giving denial of service or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phoenix SecureCore Technology 4 (boot splash screen image parsing)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The firmware parses a user-supplied boot logo image without validating it, giving denial of service or arbitrary code execution in the DXE phase. This is the Phoenix instance of the wider image-parser problem in UEFI firmware: the logo is attacker-replaceable data sitting inside the firmware volume, parsed by privileged code long before Secure Boot has any say. For an operator the uncomfortable part is that a custom boot logo is a supported, documented OEM feature, so the write path exists by design.","attack_vector":"An attacker who can replace the boot logo image in the firmware volume - requiring firmware-write access from the OS (root plus a writable ESP or an unlocked SPI region), then a reboot.","remediation":"OEM BIOS update on the fixed SecureCore Technology 4 build. Firmware flash, reboot per node. Config-side hardening that helps immediately: ensure the SPI flash and the logo storage region are write-protected at the platform level, and do not deploy custom OEM boot logos on nodes where that means leaving the region writable. Track this alongside the LogoFAIL family from other IBVs - the same image-parser class was found across Insyde, AMI and Phoenix, so a fleet with mixed OEMs needs all three checked.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5058","https://www.kb.cert.org/vuls/id/811862"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-51042","cve":"CVE-2023-51042","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (In the Linux kernel before 6.4.12, amdgpu_cs_wait_all_fences in drivers/gpu/drm/amd/amdgpu/amdgpu_cs.c has a fence use-after-free.)…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: In the Linux kernel before 6.4.12, amdgpu_cs_wait_all_fences in drivers/gpu/drm/amd/amdgpu/amdgpu_cs.c has a fence use-after-free.","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-51042","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52624","cve":"CVE-2023-52624","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Wake DMCUB before executing GPINT commands","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52624","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52678","cve":"CVE-2023-52678","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Confirm list is non-empty before utilizing list_first_entry in kfd_topology.c","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52678","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52691","cve":"CVE-2023-52691","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A double free in the amdgpu power management (SMU/powerplay). The same allocation is released twice…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu power management (SMU/powerplay). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/pm: fix a double-free in si_dpm_init","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52691","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52812","cve":"CVE-2023-52812","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: check num of link levels when update pcie param","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52812","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52816","cve":"CVE-2023-52816","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix shift out-of-bounds issue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52816","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52818","cve":"CVE-2023-52818","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd): An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: Fix UBSAN array-index-out-of-bounds for SMU7","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52818","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52825","cve":"CVE-2023-52825","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: Fix a race condition of vram buffer unref in svm code","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52825","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52913","cve":"CVE-2023-52913","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: GEM context registration making a context visible to userspace before it is fully owned, so a second thread can free it mid-setup.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52913","https://git.kernel.org/stable/c/ae278887193110dfeb857ea63e243a3851fbb0bc"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52916","cve":"CVE-2023-52916","aliases":[],"title":"ASPEED video engine capture driver (drivers/media/platform/aspeed) - iKVM path: The video engine writes past its capture buffer when the host is driving a 1600x900 mode, corrupting whatever…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED video engine capture driver (drivers/media/platform/aspeed) - iKVM path","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The video engine writes past its capture buffer when the host is driving a 1600x900 mode, corrupting whatever BMC kernel memory sits after it. The reported repro is mundane: run iKVM with virtual media mounted while the BMC is under memory pressure. Since the resolution is chosen by the host, a tenant with control of the server's display output can steer the BMC into the corrupting mode on demand. Realistically this is a BMC crash - which on a GPU node means losing remote power control and console right when you need it - but out-of-bounds writes driven by attacker-chosen geometry are the raw material for something worse.","attack_vector":"The host side picks the video mode, so any tenant with root on the bare-metal node can select it. Triggering the corruption additionally needs the BMC's iKVM/video capture to be running, which is the normal state on fleets that leave remote console available.","remediation":"Kernel fix backported into stable; reaching your fleet means a BMC firmware flash per node, out-of-band, waiting on the ODM rebase. Config-only stopgap: disable the iKVM/video capture service on nodes that do not need graphical remote console. On a GPU fleet that is usually acceptable - operators run serial-over-LAN and Redfish, not KVM - and it removes both this bug and the video-engine DMA issue below without touching firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52916","https://git.kernel.org/stable/c/4c823e4027dd1d6e88c31028dec13dd19bc7b02d","https://lists.debian.org/debian-lts-announce/2025/03/msg00001.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-52921","cve":"CVE-2023-52921","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix possible UAF in amdgpu_cs_pass1()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52921","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52930","cve":"CVE-2023-52930","aliases":[],"title":"Linux i915 GPU kernel driver (GEM tiling): MULTI-TENANT ISOLATION: A double-free reachable by racing I915_GEM_SET_TILING from multiple threads.…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (GEM tiling)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A double-free reachable by racing I915_GEM_SET_TILING from multiple threads. Double-free in the kernel slab allocator is the most directly weaponisable class here - it gives an attacker with GPU access a well-understood route to arbitrary kernel write and therefore to the host and every co-tenant on the node.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52930","https://git.kernel.org/stable/c/0769f997a7b6d5cb8336db0b4ec3d2d311b8097c"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53009","cve":"CVE-2023-53009","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Add sync after creating vram bo","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53009","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53077","cve":"CVE-2023-53077","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix shift-out-of-bounds in CalculateVMAndRowBytes","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53077","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53087","cve":"CVE-2023-53087","aliases":[],"title":"Linux i915 GPU kernel driver (active barrier tracking): Non-idle barriers were misused as fence trackers, corrupting kernel lists when i915 perf is in use. Users hit…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (active barrier tracking)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Non-idle barriers were misused as fence trackers, corrupting kernel lists when i915 perf is in use. Users hit this as an oops on list corruption - node down, jobs lost.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53087","https://git.kernel.org/stable/c/5c7591b8574c52c56b3994c2fbef1a3a311b5715"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53090","cve":"CVE-2023-53090","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Fix an illegal memory access","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53090","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53378","cve":"CVE-2023-53378","aliases":[],"title":"Linux i915 GPU kernel driver (display page table objects): The buffer object backing a display page table was not treated as a framebuffer, letting it be moved or…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (display page table objects)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"The buffer object backing a display page table was not treated as a framebuffer, letting it be moved or reused while the display engine still pointed at it. Practical outcome on a headless GPU node is a driver crash rather than a tenant boundary break, but on nodes that do run display output it can surface other memory on screen.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53378","https://git.kernel.org/stable/c/3413881e1ecc3cba722a2e87ec099692eed5be28"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53471","cve":"CVE-2023-53471","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/gfx): A race condition or locking defect in the amdgpu power management (SMU/powerplay). Concurrent paths touch…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/gfx)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu power management (SMU/powerplay). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu/gfx: disable gfx9 cp_ecc_error_irq only when enabling legacy gfx ras","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53471","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53545","cve":"CVE-2023-53545","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: unmap and remove csa_va properly","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53545","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53552","cve":"CVE-2023-53552","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: requests belonging to GuC virtual engines outliving the engine, so userspace-held request references dangle.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53552","https://git.kernel.org/stable/c/5eefc5307c983b59344a4cb89009819f580c84fa"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53707","cve":"CVE-2023-53707","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix integer overflow in amdgpu_cs_pass1","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53707","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53753","cve":"CVE-2023-53753","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix mapping to non-allocated address","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53753","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53806","cve":"CVE-2023-53806","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: populate subvp cmd info only for the top pipe","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53806","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53816","cve":"CVE-2023-53816","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: fix potential kgd_mem UAFs","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53816","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53819","cve":"CVE-2023-53819","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (amdgpu)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: amdgpu: validate offset_in_bo of drm_amdgpu_gem_va","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53819","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-54168","cve":"CVE-2023-54168","aliases":["RDMA/mlx4 prevent shift wrapping in set_user_sq_size()"],"title":"Linux kernel mlx4_ib (legacy ConnectX-3 RDMA): Same class of bug on the older mlx4 stack: the user-supplied log_sq_bb_count shift wraps in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx4_ib (legacy ConnectX-3 RDMA)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Same class of bug on the older mlx4 stack: the user-supplied log_sq_bb_count shift wraps in set_user_sq_size(). Any tenant with RDMA access on a ConnectX-3-era host gets a user-controlled shift into kernel sizing logic. Matters for operators still running legacy mlx4 nodes alongside a modern fleet.","attack_vector":"Local, low-privileged process creating an RDMA queue pair on an mlx4 device.","remediation":"Upgrade the host kernel to 6.4 or a stable backport (4.19.283, 5.4.243, 5.10.180, 5.15.111, 6.1.28, 6.2.15, 6.3.2). Host reboot. Strategically, this is a prompt to retire remaining ConnectX-3/mlx4 hardware rather than keep patching it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54168","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2023/CVE-2023-54168.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-54202","cve":"CVE-2023-54202","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: a race in the perf config ioctl where a guessable object id lets two threads free the same OA config.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54202","https://git.kernel.org/stable/c/240b1502708858b5e3f10b6dc5ca3f148a322fef"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-6817","cve":"CVE-2023-6817","aliases":[],"title":"Linux kernel (netfilter pipapo): Inactive elements mishandled in nft_pipapo_walk - use-after-free, local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter pipapo)","year":"2023","cvss_score":7.8,"severity":"high","kev":false,"impact":"Inactive elements mishandled in nft_pipapo_walk - use-after-free, local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-6817"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0071","cve":"CVE-2024-0071","aliases":[],"title":"GPU Display Driver: Local privesc (kernel buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (kernel buffer over-read)","attack_vector":"Any tenant with a container; vGPU guest","remediation":"Driver upgrade (551.61 / 550.54.14 branch); drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0071","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0073","cve":"CVE-2024-0073","aliases":[],"title":"GPU Display Driver: Local privesc (improper privilege management)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (improper privilege management)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0073","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0077","cve":"CVE-2024-0077","aliases":[],"title":"vGPU Manager: Guest-to-host privesc (improper privilege management)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host privesc (improper privilege management)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager on hypervisor; evacuate guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0077","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-285"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0084","cve":"CVE-2024-0084","aliases":[],"title":"vGPU Manager: Guest-to-host privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host privesc","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0084","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0089","cve":"CVE-2024-0089","aliases":[],"title":"GPU Display Driver: Local privesc (improper resource validation / uninit memory)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (improper resource validation / uninit memory)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0089","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-665"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0090","cve":"CVE-2024-0090","aliases":[],"title":"GPU Display Driver: Local privesc to host root (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0090","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","fleet":{"ubiquity":"Universal - Windows and Linux driver, all datacenter SKUs","remediation_pain":"`node-reboot` (driver replacement)","pain_class":"node-reboot","why_fleet_wide":"Out-of-bounds write reachable from an unprivileged local user through the driver API: any tenant process on a GPU node can attempt host code execution"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"]},{"id":"CVE-2024-0091","cve":"CVE-2024-0091","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0091","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","fleet":{"ubiquity":"Universal - same driver","remediation_pain":"`node-reboot`","pain_class":"node-reboot","why_fleet_wide":"Untrusted-pointer dereference via a driver API call from a low-privileged tenant process, giving DoS / info disclosure / tampering on a shared host"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-822"]},{"id":"CVE-2024-0099","cve":"CVE-2024-0099","aliases":[],"title":"vGPU Manager: Guest-to-host escape (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape (buffer overflow)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0099","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0107","cve":"CVE-2024-0107","aliases":[],"title":"GPU Display Driver: Local privesc (buffer over-read in driver)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (buffer over-read in driver)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0107","https://github.com/NVIDIA/product-security/tree/main/2024/5557"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0117","cve":"CVE-2024-0117","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: An out-of-bounds read in the Windows user mode driver layer chains to code execution and privilege…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds read in the Windows user mode driver layer chains to code execution and privilege escalation. Note the CVSS assumes user interaction, which on a render/VDI host is trivially satisfied. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0117","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0118","cve":"CVE-2024-0118","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A second out-of-bounds read in the Windows user mode driver layer with the same code-execution outcome. Only…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second out-of-bounds read in the Windows user mode driver layer with the same code-execution outcome. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0118","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0119","cve":"CVE-2024-0119","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A third out-of-bounds read in the Windows user mode driver layer reaching code execution and privilege…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A third out-of-bounds read in the Windows user mode driver layer reaching code execution and privilege escalation. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0119","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0120","cve":"CVE-2024-0120","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A fourth out-of-bounds read in the Windows user mode driver layer with the same impact. Only matters to you…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A fourth out-of-bounds read in the Windows user mode driver layer with the same impact. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0120","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0121","cve":"CVE-2024-0121","aliases":[],"title":"NVIDIA GPU Display Driver - Windows user mode layer: A fifth out-of-bounds read in the Windows user mode driver layer, fixed in the same bulletin as the other…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows user mode layer","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A fifth out-of-bounds read in the Windows user mode driver layer, fixed in the same bulletin as the other four. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5586. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0121","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0127","cve":"CVE-2024-0127","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A tenant who has compromised their own guest kernel feeds bad input to the host GPU…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A tenant who has compromised their own guest kernel feeds bad input to the host GPU kernel driver and reaches code execution and privilege escalation on the host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Guest root is the assumed starting point, which is exactly what every vGPU tenant has. This is the cleanest guest-to-host escape candidate in the 2024 set.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5586. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0127","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0146","cve":"CVE-2024-0146","aliases":[],"title":"vGPU Manager: Guest-to-host escape via GPU firmware buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape via GPU firmware buffer overflow","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager + GPU firmware; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0146","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0582","cve":"CVE-2024-0582","aliases":[],"title":"Linux kernel (io_uring): Page use-after-free via io_uring buffer-ring mmap - unprivileged local user to root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (io_uring)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Page use-after-free via io_uring buffer-ring mmap - unprivileged local user to root","attack_vector":"Any tenant process in a container with io_uring enabled","remediation":"Livepatchable; otherwise drain + reboot. Or disable io_uring for tenant containers","references":["https://access.redhat.com/security/cve/CVE-2024-0582"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-1086","cve":"CVE-2024-1086","aliases":[],"title":"Linux kernel (nf_tables): Use-after-free in nft_verdict_init() - double-free to local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in nft_verdict_init() - double-free to local root; weaponised public exploit with a very high success rate [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable on supported kernels (Canonical/TuxCare/Ksplice all shipped it); otherwise drain + reboot. The single highest-priority container-escape CVE of the 2024 set - treat unpatched nodes as compromised-by-default in a shared-tenant fleet","references":["https://access.redhat.com/security/cve/CVE-2024-1086"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26586","cve":"CVE-2024-26586","aliases":["mlxsw spectrum_acl_tcam fix stack corruption"],"title":"Linux kernel mlxsw (Spectrum switch ASIC ACL TCAM): On Spectrum-2 and newer, firmware reports more than 16 ACLs per group but the driver's register layout was…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxsw (Spectrum switch ASIC ACL TCAM)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"On Spectrum-2 and newer, firmware reports more than 16 ACLs per group but the driver's register layout was never widened, so putting more than 16 ACLs in a group corrupts the kernel stack and panics the switch. Triggered by adding tc filters with decreasing priority in alternating order - a shape that ordinary policy automation produces. Relevant to anyone running Linux-based switch control on Spectrum silicon.","attack_vector":"Local on the switch with network-configuration privilege - an operator or automation system installing tc filters. Not remotely reachable.","remediation":"Upgrade the switch's kernel to 6.8 or a stable backport (5.10.209, 5.15.148, 6.1.79, 6.6.14, 6.7.2). On a Cumulus/NVOS switch that means an OS image upgrade and a switch reload - a fabric rolling window. Interim: constrain your ACL automation so a single group never exceeds 16 ACLs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26586","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-26586.json"],"status":"curated"},{"id":"CVE-2024-26592","cve":"CVE-2024-26592","aliases":[],"title":"Linux kernel (ksmbd): Use-after-free in ksmbd_tcp_new_connection() - in-kernel SMB server","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ksmbd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in ksmbd_tcp_new_connection() - in-kernel SMB server","attack_vector":"Unauthenticated network (if ksmbd is exposed)","remediation":"Livepatchable; otherwise drain + reboot. Correct answer for a neocloud is to not ship ksmbd at all - blacklist the module fleet-wide","references":["https://access.redhat.com/security/cve/CVE-2024-26592"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26656","cve":"CVE-2024-26656","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix use-after-free bug","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26656","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26699","cve":"CVE-2024-26699","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix array-index-out-of-bounds in dcn35_clkmgr","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26699","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26728","cve":"CVE-2024-26728","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: fix null-pointer dereference on edid reading","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26728","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26797","cve":"CVE-2024-26797","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Prevent potential buffer overflow in map_hw_resources","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26797","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26913","cve":"CVE-2024-26913","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix dcn35 8k30 Underflow/Corruption Issue","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26913","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26914","cve":"CVE-2024-26914","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix incorrect mpc_combine array size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26914","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26922","cve":"CVE-2024-26922","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate the parameters of bo mapping operations more clearly","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26922","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26939","cve":"CVE-2024-26939","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the VMA destroy path racing against retire, freeing a virtual-memory area object still in use.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26939","https://git.kernel.org/stable/c/0e45882ca829b26b915162e8e86dbb1095768e9e"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27400","cve":"CVE-2024-27400","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu): A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: once more fix the call oder in amdgpu_ttm_move() v2","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27400","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27459","cve":"CVE-2024-27459","aliases":[],"title":"OpenVPN: Stack overflow in the interactive service","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Stack overflow in the interactive service -> local privilege escalation on the client host","attack_vector":"Local","remediation":"Control-plane: jump-host and operator workstation update","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27459"],"status":"curated"},{"id":"CVE-2024-31583","cve":"CVE-2024-31583","aliases":[],"title":"PyTorch (mobile interpreter): Use-after-free in `torch/csrc/jit/mobile/interpreter.cpp`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (mobile interpreter)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free in `torch/csrc/jit/mobile/interpreter.cpp`","attack_vector":"Customer-supplied mobile/lite model file","remediation":"Ship torch >= 2.2.0 in base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31583"],"status":"curated"},{"id":"CVE-2024-31858","cve":"CVE-2024-31858","aliases":["INTEL-SA-01124","CVE-2025-33000","INTEL-SA-01373","CVE-2022-21804","CVE-2020-12333"],"title":"Intel QuickAssist Technology (QAT) software and drivers - QAT software before 2.2.0, with a 2025 batch through 2.6.0: Out-of-bounds write in the QAT software stack giving an authenticated local…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel QuickAssist Technology (QAT) software and drivers - QAT software before 2.2.0, with a 2025 batch through 2.6.0","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in the QAT software stack giving an authenticated local user privilege escalation, with a further improper-input-validation escalation (CVSS 8.8) in the 2025 batch and an earlier credential-exposure issue in the Linux QAT package. QAT is the crypto and compression offload engine on Xeon platforms - it terminates TLS and does bulk compression for storage paths, so it handles key material by design, and it is a DMA-capable PCIe device. A local escalation through the QAT driver is a container-to-root path on nodes where QAT is enabled, and QAT's position in the TLS path makes credential exposure in the same stack materially worse than a generic driver bug.","attack_vector":"Authenticated local user on the host with access to the QAT device interfaces. Where QAT is exposed into containers or VMs for offload, that is the tenant.","remediation":"Update the QAT driver and software package to 2.2.0 or later (2.6.0+ for the 2025 batch) - a software/driver update from Intel, not a firmware flash, so it can go out with a service restart or reboot rather than a full firmware maintenance window. If QAT is not actually in use on a node, unbind and blacklist the driver rather than leaving an unused DMA-capable offload path exposed to tenants. Where you do expose QAT to tenants, review whether the crypto offload path is carrying keys that a tenant-side escalation would reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31858","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01124.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01373.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-31891","cve":"CVE-2024-31891","aliases":[],"title":"IBM Storage Scale GUI (local privilege escalation): A local privilege escalation in the Storage Scale GUI available to an actor with command-line access to the…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale GUI (local privilege escalation)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A local privilege escalation in the Storage Scale GUI available to an actor with command-line access to the GUI service account. Escalating to root on a GUI node is escalating on a node that holds cluster-wide storage credentials, which is why this scores higher operationally than 'it's just the web UI' suggests.","attack_vector":"Local, requires command-line access as the GUI service user on Storage Scale GUI 5.1.9.0-5.1.9.6 or 5.2.0.0-5.2.1.1.","remediation":"Upgrade the Storage Scale GUI. Service-level upgrade with a restart; filesystem I/O is unaffected. Companion CSV-handling issue CVE-2024-31892 is fixed in the same range.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31891","https://nvd.nist.gov/vuln/detail/CVE-2024-31892"],"status":"curated"},{"id":"CVE-2024-33656","cve":"CVE-2024-33656","aliases":["AMI-SA-2024003"],"title":"AMI AptioV UEFI BIOS (SmmComputrace DXE module): The SmmComputrace DXE module leaks stack and global memory to a local attacker, giving up the addresses and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SmmComputrace DXE module)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The SmmComputrace DXE module leaks stack and global memory to a local attacker, giving up the addresses and secrets needed to defeat firmware memory protections and escalate to arbitrary code execution and OS security bypass. Computrace is the anti-theft persistence module - it exists specifically to survive OS reinstall, so a bug in it lands in code designed for durability. Most datacenter operators do not use Computrace at all and do not realise the module is compiled into their BIOS anyway.","attack_vector":"Local, low privileges, no interaction. Any code on the host OS can start reading. On a bare-metal GPU rental the tenant qualifies without doing anything unusual.","remediation":"BIOS update from your server vendor with the fixed AptioV build - firmware flash plus host reboot per node, vendor-gated, and AMI names only 'AptioV' as the fix version so you have to confirm the specific BIOS release with your OEM. The useful config-only step here is subtractive: check whether Computrace is enabled in BIOS setup on your fleet and disable it, since datacenter operators almost never need it and it removes the attack surface with a setup change plus one reboot rather than a firmware flash campaign.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024003.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-33656"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-33657","cve":"CVE-2024-33657","aliases":["AMI-SA-2024003"],"title":"AMI AptioV UEFI BIOS (SMM modules): An SMM vulnerability letting a privileged local attacker execute arbitrary code in System Management Mode…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SMM modules)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An SMM vulnerability letting a privileged local attacker execute arbitrary code in System Management Mode, manipulate SMM stack memory, and leak SMRAM contents into kernel space. SMRAM is supposed to be opaque to the OS; leaking it hands the attacker firmware secrets and the layout information needed to build a reliable bootkit. Once code runs in SMM the attacker is above the hypervisor and can survive OS reinstall, so a node that was compromised once should be considered compromised until its firmware is reflashed and verified, not just reimaged.","attack_vector":"Local, low privileges required per AMI's CVSS vector, no user interaction. Needs code on the host - which on any node running untrusted tenant workloads or on any node after an initial OS compromise is a given.","remediation":"BIOS update from your server vendor carrying the fixed AptioV build - firmware flash plus a full host reboot, per node. AMI's advisory names only 'AptioV' as the fix version rather than a specific BKC, so you must confirm with your OEM which BIOS release for your SKU actually contains it; do not assume 'latest' covers it. No config-only mitigation for an SMM bug.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024003.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-33657"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-35817","cve":"CVE-2024-35817","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: amdgpu_ttm_gart_bind set gtt bound flag","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-35931","cve":"CVE-2024-35931","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Skip do PCI error slot reset during RAS recovery","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35931","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36914","cve":"CVE-2024-36914","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Skip on writeback when it's not applicable","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36914","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36971","cve":"CVE-2024-36971","aliases":[],"title":"Linux kernel (net routing): Use-after-free in network route management (__dst_negative_advice) - actively exploited","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net routing)","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in network route management (__dst_negative_advice) - actively exploited [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-36971"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-38080","cve":"CVE-2024-38080","aliases":[],"title":"Microsoft Hyper-V: Hyper-V elevation of privilege, exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Hyper-V elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user / guest on the host","remediation":"Windows update + host reboot with live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38080"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-38459","cve":"CVE-2024-38459","aliases":[],"title":"langchain-experimental (Python REPL): Python REPL exposed without an opt-in","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"langchain-experimental (Python REPL)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Python REPL exposed without an opt-in","attack_vector":"Untrusted agent input","remediation":"Upgrade past 0.0.61","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38459"],"status":"curated"},{"id":"CVE-2024-38581","cve":"CVE-2024-38581","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/mes): A use-after-free in the amdgpu firmware, ACPI and IP-block initialisation. Freed kernel memory is reachable…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/mes)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu firmware, ACPI and IP-block initialisation. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu/mes: fix use-after-free issue","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38581","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-39291","cve":"CVE-2024-39291","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix buffer size in gfx_v9_4_3_init_ cp_compute_microcode() and rlc_microcode()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39291","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-39471","cve":"CVE-2024-39471","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: add error handle to avoid out-of-bounds","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39471","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-40990","cve":"CVE-2024-40990","aliases":["RDMA/mlx5 add check for srq max_sge attribute"],"title":"Linux kernel mlx5_ib (shared receive queue): The max_sge attribute for a shared receive queue is taken from the user and used unchecked. A tenant process…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (shared receive queue)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The max_sge attribute for a shared receive queue is taken from the user and used unchecked. A tenant process creating an SRQ supplies a value the kernel trusts - memory corruption from an ordinary verbs call available to any RDMA workload on the node.","attack_vector":"Local, low-privileged process with RDMA verbs access on an mlx5 device.","remediation":"Upgrade the host kernel to 6.10 or a stable backport (5.10.221, 5.15.162, 6.1.96, 6.6.36, 6.9.7). Rolling reboot of the RDMA fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40990","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-40990.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-41008","cve":"CVE-2024-41008","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: change vm->task_info handling","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41008","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-41011","cve":"CVE-2024-41011","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: don't allow mapping the MMIO HDP page with large pages","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41011","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-41061","cve":"CVE-2024-41061","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix array-index-out-of-bounds in dml2/FCLKChangeSupport","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41061","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-41092","cve":"CVE-2024-41092","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: revocation of fence registers racing with their use, leaving a dangling fence.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41092","https://git.kernel.org/stable/c/06dec31a0a5112a91f49085e8a8fa1a82296d5c7"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42064","cve":"CVE-2024-42064","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Skip pipe if the pipe idx not set properly","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42064","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42117","cve":"CVE-2024-42117","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: ASSERT when failing to find index by plane/stream id","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42117","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42118","cve":"CVE-2024-42118","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Do not return negative stream id for array","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42118","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42119","cve":"CVE-2024-42119","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Skip finding free audio for unknown engine_id","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42119","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42120","cve":"CVE-2024-42120","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check pipe offset before setting vblank","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42120","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42121","cve":"CVE-2024-42121","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check index msg_id before read or write","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42121","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-4467","cve":"CVE-2024-4467","aliases":[],"title":"QEMU (qemu-img): `qemu-img info` on an untrusted qcow2 image reaches arbitrary host file read/write","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (qemu-img)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"`qemu-img info` on an untrusted qcow2 image reaches arbitrary host file read/write","attack_vector":"Tenant-supplied disk image processed by the control plane","remediation":"Update qemu-img and never run it untrusted-unsandboxed; no reboot. Directly relevant to any \"bring your own VM image\" feature","references":["https://access.redhat.com/security/cve/CVE-2024-4467"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-44977","cve":"CVE-2024-44977","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Validate TA binary size","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44977","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45019","cve":"CVE-2024-45019","aliases":["net/mlx5e take state lock during tx timeout reporter"],"title":"Linux kernel mlx5_core TX timeout devlink health reporter: The TX timeout recovery path runs without the state lock, so it races channel teardown. The racing side is…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core TX timeout devlink health reporter","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The TX timeout recovery path runs without the state lock, so it races channel teardown. The racing side is driven by unprivileged statistics reads - cat /proc/net/dev, ip -s link, or anything scraping /sys/class/net/*/statistics/*. That requeues the stats work while the channels are being freed, giving a reliable use-after-free read of groomable slab memory from an ordinary unprivileged account. Every Prometheus node-exporter on the fleet is polling exactly these files.","attack_vector":"Local unprivileged user reading standard network statistics, combined with a TX timeout or channel reconfiguration. No special capability needed.","remediation":"Upgrade the host kernel to a build carrying the fix (6.11 and the corresponding stable backports). Rolling reboot of the fleet. No practical config mitigation - you are not going to stop metrics collection.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45019","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-45019.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-4610","cve":"CVE-2024-4610","aliases":[],"title":"Arm Mali GPU kernel driver: Use-after-free in the Bifrost/Valhall GPU kernel driver - a local non-privileged user reaches freed kernel…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Arm Mali GPU kernel driver","year":"2024","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in the Bifrost/Valhall GPU kernel driver - a local non-privileged user reaches freed kernel memory; exploited in the wild [KEV]","attack_vector":"Local user with GPU device access","remediation":"Driver update + reboot. Not applicable to NVIDIA/AMD datacentre GPUs, but directly relevant to any Arm-based (Grace, Ampere Altra) node that ships a Mali display GPU, and it is the closest existing precedent for tenant-reachable GPU-driver privesc","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4610"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46725","cve":"CVE-2024-46725","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix out-of-bounds write warning","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46725","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46729","cve":"CVE-2024-46729","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Fix incorrect size calculation for loop","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46729","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46730","cve":"CVE-2024-46730","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Ensure array index tg_inst won't be -1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46730","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46746","cve":"CVE-2024-46746","aliases":[],"title":"Linux HID/amd_sfh - driver_data freed after HID device destruction: A use-after-free in the AMD Sensor Fusion Hub HID driver: driver_data is freed in the wrong order relative to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux HID/amd_sfh - driver_data freed after HID device destruction","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the AMD Sensor Fusion Hub HID driver: driver_data is freed in the wrong order relative to hid_destroy_device(), so callbacks touch memory that is already gone. Kernel UAF is a privilege-escalation primitive. AMD SFH is a client-platform driver and unlikely to be loaded on an Instinct server - but it is compiled into stock distro kernels, and a driver that is present but unneeded is attack surface you are carrying for nothing.","attack_vector":"Local, on hosts where the amd_sfh driver is loaded.","remediation":"Fixed in the Linux kernel; take the distro update and reboot. Better: blacklist amd_sfh on server images. Auditing your GPU nodes for client-platform drivers that autoload and are never used is a cheap one-off that shrinks the kernel attack surface permanently.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46746"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46803","cve":"CVE-2024-46803","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Check debug trap enable before write dbg_ev_file","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46803","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46804","cve":"CVE-2024-46804","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add array index check for hdcp ddc access","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46804","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46812","cve":"CVE-2024-46812","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Skip inactive planes within ModeSupportAndSystemConfiguration","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46812","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46813","cve":"CVE-2024-46813","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check link_index before accessing dc->links[]","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46813","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46814","cve":"CVE-2024-46814","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check msg_id before processing transcation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46814","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46816","cve":"CVE-2024-46816","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Stop amdgpu_dm initialize when link nums greater than max_links","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46816","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46818","cve":"CVE-2024-46818","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check gpio_id before used as array index","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46818","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46820","cve":"CVE-2024-46820","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: remove irq disabling in vcn 5 suspend","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46820","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46821","cve":"CVE-2024-46821","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/pm: Fix negative array index read","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46821","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46836","cve":"CVE-2024-46836","aliases":[],"title":"ASPEED USB device controller driver (drivers/usb/gadget/udc/aspeed_udc.c): The BMC presents itself to the host over USB - that is how OpenBMC does virtual media, USB-network (the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED USB device controller driver (drivers/usb/gadget/udc/aspeed_udc.c)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"The BMC presents itself to the host over USB - that is how OpenBMC does virtual media, USB-network (the host-to-BMC management channel) and HID for iKVM. The endpoint index coming from the host is used without a bounds check, so the host walks past the endpoint array into adjacent BMC kernel memory. The direction of travel is what matters here: the host is the untrusted side on a rented bare-metal node, and this is a host-controlled index into BMC kernel structures. It is a genuine tenant-to-BMC crossing, not a BMC-local bug.","attack_vector":"Requires the ability to drive USB control traffic from the host to the BMC's USB gadget - root on the bare-metal server, which any tenant of that node has. The gadget is enabled by default on OpenBMC platforms that offer virtual media or a USB management NIC.","remediation":"Kernel patch backported across stable trees; on real fleets it arrives only in a new BMC firmware image, so per-node out-of-band flash with the usual ODM lag and brick risk. Partial config mitigation: unbind or do not compose the virtual-media and USB-gadget functions on nodes that do not need them. That is often viable for GPU nodes provisioned by network boot, less so where the ODM's own management path rides the USB NIC.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46836","https://git.kernel.org/stable/c/b2a50ffdd1a079869a62198a8d1441355c513c7c","https://lists.debian.org/debian-lts-announce/2025/01/msg00001.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-46850","cve":"CVE-2024-46850","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid race between dcn35_set_drr() and dc_state_destruct()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46850","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46871","cve":"CVE-2024-46871","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Correct the defined value for AMDGPU_DMUB_NOTIFICATION_MAX","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46871","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49895","cve":"CVE-2024-49895","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index out of bounds in DCN30 degamma hardware format translation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49895","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49969","cve":"CVE-2024-49969","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index out of bounds in DCN30 color transformation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49969","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49989","cve":"CVE-2024-49989","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/display: fix double free issue during amdgpu module unload","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49989","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49991","cve":"CVE-2024-49991","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: amdkfd_free_gtt_mem clear the correct pointer","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49991","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50221","cve":"CVE-2024-50221","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/pm: Vangogh: Fix kernel memory out of bounds write","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50221","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50282","cve":"CVE-2024-50282","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu GEM/VM/command-submission ioctl surface - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: add missing size check in amdgpu_debugfs_gprwave_read()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50282","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53133","cve":"CVE-2024-53133","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Handle dml allocation failure to avoid crash","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53133","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56551","cve":"CVE-2024-56551","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu kernel driver core. Freed kernel memory is reachable…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu kernel driver core. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix usage slab after free","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56551","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56608","cve":"CVE-2024-56608","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix out-of-bounds access in 'dcn21_link_encoder_create'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56608","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56695","cve":"CVE-2024-56695","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Use dynamic allocation for CU occupancy array in 'kfd_get_cu_occupancy()'","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56695","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56775","cve":"CVE-2024-56775","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A double free in the amdgpu display core (DC/DM). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amd/display: Fix handling of plane refcount","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56775","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56784","cve":"CVE-2024-56784","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Adding array index check to prevent memory corruption","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56784","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57801","cve":"CVE-2024-57801","aliases":[],"title":"Linux kernel mlx5_core eswitch vport representors / IPsec FS: TENANT ISOLATION: during driver unload the vport representor private struct is freed before unregister_netdev…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch vport representors / IPsec FS","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: during driver unload the vport representor private struct is freed before unregister_netdev runs, so the kernel walks freed representor state across every VF on the box. Vport representors are the per-tenant hooks in switchdev mode; a use-after-free walking all of them on a shared host is a host-kernel corruption reachable through ordinary driver lifecycle events.","attack_vector":"Local - triggered on mlx5 driver unload/reload on a switchdev SR-IOV host. Reachable by anyone who can induce a driver reload (operator action, firmware reset flow, or a fault path an attacker provokes).","remediation":"Upgrade the host kernel to 6.13 or a stable backport (6.6.70, 6.12.9). Rolling reboot. Until then, avoid mlx5 driver unload/reload on live switchdev hosts - use full node reboots instead of in-place driver restarts.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57801","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-57801.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57918","cve":"CVE-2024-57918","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: fix page fault due to max surface definition mismatch","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57918","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57921","cve":"CVE-2024-57921","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Add a lock when accessing the buddy trim function","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57921","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-5998","cve":"CVE-2024-5998","aliases":[],"title":"LangChain (`FAISS.deserialize_from_bytes`): Pickle deserialization of an untrusted vector index","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (`FAISS.deserialize_from_bytes`)","year":"2024","cvss_score":7.8,"severity":"high","kev":false,"impact":"Pickle deserialization of an untrusted vector index","attack_vector":"Customer-supplied FAISS index file from shared storage","remediation":"Upgrade; a vector index is a pickle and must be treated as a model artifact","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-5998"],"status":"curated"},{"id":"CVE-2025-0678","cve":"CVE-2025-0678","aliases":[],"title":"GRUB2 (squashfs): Integer overflow in the squash4 filesystem module leading to out-of-bounds write and possible Secure Boot…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (squashfs)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow in the squash4 filesystem module leading to out-of-bounds write and possible Secure Boot bypass","attack_vector":"Local, crafted filesystem image","remediation":"GRUB2 update plus dbx revocation; squashfs is used by live/netboot images, so this lands directly on the bare-metal provisioning path","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0678"],"status":"curated"},{"id":"CVE-2025-10155","cve":"CVE-2025-10155","aliases":[],"title":"picklescan: Improper input validation lets a crafted pickle evade scanning","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper input validation lets a crafted pickle evade scanning","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 0.0.30; layer scanning with format restriction","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10155"],"status":"curated"},{"id":"CVE-2025-1753","cve":"CVE-2025-1753","aliases":[],"title":"LlamaIndex CLI: OS command injection via the `--files` argument","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LlamaIndex CLI","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"OS command injection via the `--files` argument","attack_vector":"Attacker-influenced filename in an automated pipeline","remediation":"Upgrade past 0.12.20","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1753"],"status":"curated"},{"id":"CVE-2025-21333","cve":"CVE-2025-21333","aliases":[],"title":"Microsoft Hyper-V: Heap-based buffer overflow in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Heap-based buffer overflow in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user on the Hyper-V host (or a container using the VSP path)","remediation":"Windows update + host reboot with VM live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21333"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21334","cve":"CVE-2025-21334","aliases":[],"title":"Microsoft Hyper-V: Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user on the Hyper-V host","remediation":"Windows update + host reboot with VM live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21334"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21335","cve":"CVE-2025-21335","aliases":[],"title":"Microsoft Hyper-V: Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Microsoft Hyper-V","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Use-after-free in the NT Kernel Integration VSP - elevation of privilege, exploited in the wild [KEV]","attack_vector":"Local user on the Hyper-V host","remediation":"Windows update + host reboot with live-migration","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21335"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21714","cve":"CVE-2025-21714","aliases":["RDMA/mlx5 fix implicit ODP use after free"],"title":"Linux kernel mlx5_ib on-demand paging (ODP): Implicit on-demand-paging memory region destroy work can be queued twice, so the second run touches an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib on-demand paging (ODP)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Implicit on-demand-paging memory region destroy work can be queued twice, so the second run touches an already-freed MR - a refcount underflow and use-after-free. ODP is what lets RDMA register memory lazily, which is heavily used by large-model training frameworks, so this is reachable from ordinary tenant RDMA registration patterns on a GPU node.","attack_vector":"Local, low-privileged - a tenant process registering and tearing down ODP memory regions through libibverbs.","remediation":"Upgrade the host kernel to 6.14 or a stable backport (6.12.13, 6.13.2). Rolling reboot of nodes where ODP is enabled.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21714","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-21714.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21756","cve":"CVE-2025-21756","aliases":[],"title":"Linux kernel (vsock): vsock binding not kept until socket destruction - use-after-free, local root with a public exploit","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (vsock)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"vsock binding not kept until socket destruction - use-after-free, local root with a public exploit","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Livepatchable; otherwise drain + reboot. Blacklist vsock where unused","references":["https://access.redhat.com/security/cve/CVE-2025-21756"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21780","cve":"CVE-2025-21780","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: avoid buffer overflow attach in smu_sys_set_pp_table()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21780","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21842","cve":"CVE-2025-21842","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: amdkfd: properly free gang_ctx_bo when failed to init user queue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21842","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21882","cve":"CVE-2025-21882","aliases":["net/mlx5 vport QoS cleanup on error"],"title":"Linux kernel mlx5_core eswitch vport QoS scheduling: TENANT ISOLATION: when enabling per-vport QoS fails, the scheduling node is leaked and the pointer left…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch vport QoS scheduling","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: when enabling per-vport QoS fails, the scheduling node is leaked and the pointer left dangling. Per-VF QoS is the mechanism that stops one tenant's VF from starving the others' bandwidth, so failures here both leak host kernel memory and undermine the rate-limiting you sold as an isolation guarantee.","attack_vector":"Local, low-privileged - reached through the VF QoS configuration path on an SR-IOV host. An attacker who can make QoS enablement fail (resource exhaustion, invalid rates) reaches the bad path.","remediation":"Upgrade the host kernel to 6.14 or the 6.13.6 stable backport. Rolling reboot of SR-IOV hosts. No firmware flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21882","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-21882.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21968","cve":"CVE-2025-21968","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: Fix slab-use-after-free on hdcp_work","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21968","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21985","cve":"CVE-2025-21985","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix out-of-bound accesses","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21985","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21991","cve":"CVE-2025-21991","aliases":[],"title":"Linux x86/microcode/AMD - out-of-bounds on CPU-less NUMA nodes: The AMD microcode loader iterated every NUMA node and unconditionally touched per-CPU data for each node's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/microcode/AMD - out-of-bounds on CPU-less NUMA nodes","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The AMD microcode loader iterated every NUMA node and unconditionally touched per-CPU data for each node's first CPU - which does not exist on CPU-less NUMA nodes. The result is an out-of-bounds kernel access during microcode load. CPU-less NUMA nodes are not exotic on AI hardware: they are what you get with CXL memory expanders and certain GPU/HBM topologies, so this fires on exactly the machines an AI operator runs.","attack_vector":"Local; triggered during microcode loading on systems with CPU-less NUMA nodes. Not attacker-controlled so much as a crash you hit on affected topologies.","remediation":"Fixed in the Linux kernel microcode loader. Take the distro kernel update and reboot. No firmware or BIOS step. If you run CXL-attached memory or other configurations that produce memory-only NUMA nodes, treat this as a stability fix worth taking promptly rather than a security backlog item.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21991"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2025-22836","cve":"CVE-2025-22836","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (an integer overflow)…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (an integer overflow) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22836","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-22893","cve":"CVE-2025-22893","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (insufficient control-flow…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (insufficient control-flow management) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22893","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23244","cve":"CVE-2025-23244","aliases":[],"title":"GPU Display Driver: Local privesc (insufficient access control)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (insufficient access control)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23244","https://github.com/NVIDIA/product-security/tree/main/2025/5630"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-863"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23264","cve":"CVE-2025-23264","aliases":[],"title":"Megatron-LM: Arbitrary code execution in the training job","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution in the training job","attack_vector":"Malicious model/checkpoint or dataset","remediation":"Bump Megatron-LM in training images; rebuild; treat checkpoints as untrusted input","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23264","https://github.com/NVIDIA/product-security/tree/main/2025/5663"],"status":"curated","fleet":{"ubiquity":"Common - Megatron is the reference large-model training stack for customers doing pretraining on rented clusters","remediation_pain":"`hot-patch` (upgrade to v0.12.1, rebuild training images)","pain_class":"hot-patch","why_fleet_wide":"A malicious file supplied to the Python component triggers code injection; on a shared training cluster the attacker executes inside a job that already holds cluster-wide storage and fabric credentials"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23265","cve":"CVE-2025-23265","aliases":[],"title":"Megatron-LM: Arbitrary code execution","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution","attack_vector":"Malicious model/checkpoint","remediation":"Bump Megatron-LM; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23265","https://github.com/NVIDIA/product-security/tree/main/2025/5663"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23276","cve":"CVE-2025-23276","aliases":[],"title":"GPU Display Driver: Local privesc via file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc via file permissions","attack_vector":"Local user on the node","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23276","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-552"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23283","cve":"CVE-2025-23283","aliases":[],"title":"vGPU Manager: Guest-to-host escape (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape (stack buffer overflow)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23283","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23284","cve":"CVE-2025-23284","aliases":[],"title":"vGPU Manager: Guest-to-host escape (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Guest-to-host escape (stack buffer overflow)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate guest VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23284","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-121"]},{"id":"CVE-2025-23294","cve":"CVE-2025-23294","aliases":[],"title":"NVIDIA WebDataset: Arbitrary code execution with elevated permissions from the data-loading library. WebDataset exists to stream…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA WebDataset","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution with elevated permissions from the data-loading library. WebDataset exists to stream shards from object storage, so the trust boundary is whoever can write to your bucket. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5658 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23294","https://github.com/NVIDIA/product-security/tree/main/2025/5658"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2025-23295","cve":"CVE-2025-23295","aliases":[],"title":"NVIDIA Apex: A Python component injects code from a malicious file. Apex is a near-universal dependency in mixed-precision…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Apex","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python component injects code from a malicious file. Apex is a near-universal dependency in mixed-precision training images. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5680 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23295","https://github.com/NVIDIA/product-security/tree/main/2025/5680"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23296","cve":"CVE-2025-23296","aliases":[],"title":"NVIDIA Isaac-GR00T N1: A Python component injects code from crafted input into the robotics-model pipeline, which in practice runs…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac-GR00T N1","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python component injects code from crafted input into the robotics-model pipeline, which in practice runs on GPU cluster nodes. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5681 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23296","https://github.com/NVIDIA/product-security/tree/main/2025/5681"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23298","cve":"CVE-2025-23298","aliases":[],"title":"NVIDIA Merlin Transformers4Rec: A Python dependency permits code injection into the recommender training job. In an AI datacenter this is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Merlin Transformers4Rec","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python dependency permits code injection into the recommender training job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5683 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23298","https://github.com/NVIDIA/product-security/tree/main/2025/5683"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23303","cve":"CVE-2025-23303","aliases":[],"title":"NVIDIA NeMo Framework: Deserialization of untrusted data reaches remote code execution when a crafted artifact is loaded. In an AI…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Deserialization of untrusted data reaches remote code execution when a crafted artifact is loaded. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5686 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23303","https://github.com/NVIDIA/product-security/tree/main/2025/5686"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-23304","cve":"CVE-2025-23304","aliases":[],"title":"NVIDIA NeMo Framework: Loading a .nemo file with crafted metadata injects code at model-load time - the model file itself is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Loading a .nemo file with crafted metadata injects code at model-load time - the model file itself is the payload. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5686 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23304","https://github.com/NVIDIA/product-security/tree/main/2025/5686"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-22"]},{"id":"CVE-2025-23305","cve":"CVE-2025-23305","aliases":[],"title":"NVIDIA Megatron-LM: A code-injection flaw in the tools component executes attacker-controlled code inside the training job. In an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A code-injection flaw in the tools component executes attacker-controlled code inside the training job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5685 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23305","https://github.com/NVIDIA/product-security/tree/main/2025/5685"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23306","cve":"CVE-2025-23306","aliases":[],"title":"NVIDIA Megatron-LM: megatron/training/arguments.py injects code from malicious input - the argument-parsing path, which every…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"megatron/training/arguments.py injects code from malicious input - the argument-parsing path, which every training launch goes through. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5685 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23306","https://github.com/NVIDIA/product-security/tree/main/2025/5685"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23307","cve":"CVE-2025-23307","aliases":[],"title":"NVIDIA NeMo Curator: A malicious file processed by the data-curation pipeline injects code. Curator exists to chew through large…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Curator","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A malicious file processed by the data-curation pipeline injects code. Curator exists to chew through large untrusted corpora, so the untrusted-input assumption is the product. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5690 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23307","https://github.com/NVIDIA/product-security/tree/main/2025/5690"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23312","cve":"CVE-2025-23312","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the retrieval-services component injects code into the running job. In an AI datacenter this…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the retrieval-services component injects code into the running job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23312","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23313","cve":"CVE-2025-23313","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the NLP component injects code into the running job. In an AI datacenter this is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the NLP component injects code into the running job. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23313","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23314","cve":"CVE-2025-23314","aliases":[],"title":"NVIDIA NeMo Framework: A second NLP-component code-injection path with the same reach. In an AI datacenter this is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second NLP-component code-injection path with the same reach. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23314","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23315","cve":"CVE-2025-23315","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the export-and-deploy component injects code - notable because export/deploy typically runs…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the export-and-deploy component injects code - notable because export/deploy typically runs with registry push credentials. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5689 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23315","https://github.com/NVIDIA/product-security/tree/main/2025/5689"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23348","cve":"CVE-2025-23348","aliases":[],"title":"NVIDIA Megatron-LM: The pretrain_gpt script injects code from crafted data. In an AI datacenter this is the model-and-data supply…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The pretrain_gpt script injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23348","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23349","cve":"CVE-2025-23349","aliases":[],"title":"NVIDIA Megatron-LM: tasks/orqa/unsupervised/nq.py injects code from crafted data. In an AI datacenter this is the model-and-data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"tasks/orqa/unsupervised/nq.py injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23349","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23352","cve":"CVE-2025-23352","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A malicious guest drives the Virtual GPU Manager into using an uninitialised pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A malicious guest drives the Virtual GPU Manager into using an uninitialised pointer, reaching code execution, privilege escalation and information disclosure on the host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Uninitialised-pointer use under guest control is a strong escape primitive - prioritise this above the null-deref bugs in the same bulletin.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5703. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23352","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-824"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-23353","cve":"CVE-2025-23353","aliases":[],"title":"NVIDIA Megatron-LM: The msdp preprocessing script injects code from crafted data - preprocessing is where untrusted corpora…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The msdp preprocessing script injects code from crafted data - preprocessing is where untrusted corpora arrive. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23353","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23354","cve":"CVE-2025-23354","aliases":[],"title":"NVIDIA Megatron-LM: The ensemble_classifier script injects code from crafted data. In an AI datacenter this is the model-and-data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The ensemble_classifier script injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5698 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23354","https://github.com/NVIDIA/product-security/tree/main/2025/5698"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23357","cve":"CVE-2025-23357","aliases":[],"title":"NVIDIA Megatron-LM: A script in the repository injects code from crafted data. In an AI datacenter this is the model-and-data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A script in the repository injects code from crafted data. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5712 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23357","https://github.com/NVIDIA/product-security/tree/main/2025/5712"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-23361","cve":"CVE-2025-23361","aliases":[],"title":"NVIDIA NeMo Framework: Malicious input causes improper control of code generation, reaching code execution. In an AI datacenter this…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Malicious input causes improper control of code generation, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5718 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23361","https://github.com/NVIDIA/product-security/tree/main/2025/5718"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-24303","cve":"CVE-2025-24303","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (a missing check for an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (a missing check for an exceptional condition) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24303","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-24484","cve":"CVE-2025-24484","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (improper input validation)…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (improper input validation) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24484","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-32463","cve":"CVE-2025-32463","aliases":[],"title":"sudo: Local privilege escalation via the `--chroot` option - any local user reaches root on default sudoers, no…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"sudo","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Local privilege escalation via the `--chroot` option - any local user reaches root on default sudoers, no sudo rights needed [KEV]","attack_vector":"Local user, incl. inside a container that ships sudo","remediation":"Package update (sudo >= 1.9.17p1); no reboot. Also rebuild every container base image that includes sudo - the host fix does not cover tenant images","references":["https://access.redhat.com/security/cve/CVE-2025-32463"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-33178","cve":"CVE-2025-33178","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the BERT services component injects code. In an AI datacenter this is the model-and-data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the BERT services component injects code. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5718 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33178","https://github.com/NVIDIA/product-security/tree/main/2025/5718"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33183","cve":"CVE-2025-33183","aliases":[],"title":"NVIDIA Isaac-GR00T N1.5: A Python component injects code from crafted input. In an AI datacenter this is the model-and-data supply…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac-GR00T N1.5","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A Python component injects code from crafted input. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5725 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33183","https://github.com/NVIDIA/product-security/tree/main/2025/5725"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33184","cve":"CVE-2025-33184","aliases":[],"title":"NVIDIA Isaac-GR00T N1.5: A second Python code-injection path in the same release. In an AI datacenter this is the model-and-data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Isaac-GR00T N1.5","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second Python code-injection path in the same release. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5725 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33184","https://github.com/NVIDIA/product-security/tree/main/2025/5725"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33189","cve":"CVE-2025-33189","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An out-of-bounds write in SROOT firmware reaches code execution and privilege escalation in the root-of-trust…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds write in SROOT firmware reaches code execution and privilege escalation in the root-of-trust context. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33189","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33204","cve":"CVE-2025-33204","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data in the NLP and LLM components injects code. In an AI datacenter this is the model-and-data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data in the NLP and LLM components injects code. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5729 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33204","https://github.com/NVIDIA/product-security/tree/main/2025/5729"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33206","cve":"CVE-2025-33206","aliases":[],"title":"Nsight Graphics: Local code exec via command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Nsight Graphics","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local code exec via command injection","attack_vector":"Unprivileged local user on a dev node","remediation":"Upgrade Nsight packages in dev images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33206","https://github.com/NVIDIA/product-security/tree/main/2026/5738"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2025-33217","cve":"CVE-2025-33217","aliases":[],"title":"GPU Display Driver: Local privesc to host root (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc to host root (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33217","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-33218","cve":"CVE-2025-33218","aliases":[],"title":"GPU Display Driver: Local privesc (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (integer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33218","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-33219","cve":"CVE-2025-33219","aliases":[],"title":"GPU Display Driver: Local privesc (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (integer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33219","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","fleet":{"ubiquity":"Universal - the Linux kernel module is on every NVIDIA compute node","remediation_pain":"`node-drain` then `node-reboot` (kernel module replacement)","pain_class":"node-reboot","why_fleet_wide":"Integer overflow in the Linux kernel module reachable from a tenant process: code execution at elevated privilege, DoS, or access to another tenant's data on the same host"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"]},{"id":"CVE-2025-33220","cve":"CVE-2025-33220","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A malicious guest causes the Virtual GPU Manager to access heap memory after it has…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A malicious guest causes the Virtual GPU Manager to access heap memory after it has been freed, reaching host code execution and privilege escalation. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. A guest-triggerable host-side use-after-free is the highest-value vGPU bug in the 2026 bulletin set; treat it as a presumed escape until proven otherwise.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5747. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33220","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-33226","cve":"CVE-2025-33226","aliases":[],"title":"NVIDIA NeMo Framework: Crafted data reaches code injection and privilege escalation in the job context. In an AI datacenter this is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Crafted data reaches code injection and privilege escalation in the job context. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5736 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33226","https://github.com/NVIDIA/product-security/tree/main/2025/5736"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33233","cve":"CVE-2025-33233","aliases":[],"title":"Merlin Transformers4Rec: RCE via unsanitized input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Merlin Transformers4Rec","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsanitized input","attack_vector":"Malicious model/dataset","remediation":"Bump the package; rebuild recsys images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33233","https://github.com/NVIDIA/product-security/tree/main/2026/5761"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33234","cve":"CVE-2025-33234","aliases":[],"title":"NVIDIA runx: Arbitrary command exec via shell metacharacter injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA runx","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary command exec via shell metacharacter injection","attack_vector":"Local user invoking the tool","remediation":"Upgrade runx on nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33234","https://github.com/NVIDIA/product-security/tree/main/2026/5764"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2025-33235","cve":"CVE-2025-33235","aliases":[],"title":"NVIDIA Resiliency Extension: A race condition in the checkpointing core reaches information disclosure, data tampering and privilege…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Resiliency Extension","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition in the checkpointing core reaches information disclosure, data tampering and privilege escalation. Checkpoints are the training run's crown jewels - a tampering primitive here means silently poisoned model state that survives every restart.","attack_vector":"Local, low privileges. An account on a node participating in the checkpoint write path.","remediation":"Update the Resiliency Extension per bulletin 5746 and rebuild training images. Cost: image rebuild and job restart. Verify checkpoint integrity out-of-band if you suspect exposure - a corrupted checkpoint does not announce itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33235","https://github.com/NVIDIA/product-security/tree/main/2025/5746"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-362"]},{"id":"CVE-2025-33236","cve":"CVE-2025-33236","aliases":[],"title":"NeMo Framework: Arbitrary Python code exec via code injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary Python code exec via code injection","attack_vector":"Malicious model/config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33236","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33239","cve":"CVE-2025-33239","aliases":[],"title":"Megatron-Bridge: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33239","https://github.com/NVIDIA/product-security/tree/main/2026/5781"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33240","cve":"CVE-2025-33240","aliases":[],"title":"Megatron-Bridge: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33240","https://github.com/NVIDIA/product-security/tree/main/2026/5781"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33241","cve":"CVE-2025-33241","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33241","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33243","cve":"CVE-2025-33243","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33243","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33246","cve":"CVE-2025-33246","aliases":[],"title":"NeMo Framework: Local command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local command injection","attack_vector":"Malicious config / local user","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33246","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"]},{"id":"CVE-2025-33247","cve":"CVE-2025-33247","aliases":[],"title":"Megatron-LM: Local privesc / RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc / RCE via unsafe pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-LM; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33247","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33248","cve":"CVE-2025-33248","aliases":[],"title":"Megatron-LM: RCE via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-LM","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-LM; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33248","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33249","cve":"CVE-2025-33249","aliases":[],"title":"NeMo Framework: Local command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local command injection","attack_vector":"Malicious config","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33249","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"]},{"id":"CVE-2025-33250","cve":"CVE-2025-33250","aliases":[],"title":"NeMo Framework: RCE via unsafe object deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe object deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33250","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33251","cve":"CVE-2025-33251","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33251","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2025-33252","cve":"CVE-2025-33252","aliases":[],"title":"NeMo Framework: RCE via pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via pickle deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33252","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33253","cve":"CVE-2025-33253","aliases":[],"title":"NeMo Framework: RCE via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33253","https://github.com/NVIDIA/product-security/tree/main/2026/5762"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-37854","cve":"CVE-2025-37854","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: Fix mode1 reset crash issue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37854","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37903","cve":"CVE-2025-37903","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: Fix slab-use-after-free in hdcp","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37903","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37911","cve":"CVE-2025-37911","aliases":[],"title":"Linux bnxt_en driver (ethtool coredump / bnxt_get_coredump): Out-of-bounds memcpy when retrieving a firmware coredump via ethtool — the returned DMA length can exceed the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (ethtool coredump / bnxt_get_coredump)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds memcpy when retrieving a firmware coredump via ethtool — the returned DMA length can exceed the buffer the driver allocated, corrupting kernel memory. What makes this operationally awkward is that `ethtool -w` is exactly what your support workflow runs when a NIC misbehaves, so the diagnostic step is the trigger.","attack_vector":"Local privileged user running an ethtool coredump against the NIC, with firmware returning an over-long length. Relevant if tenants have root on bare metal, or if a compromised NIC firmware can influence the returned length.","remediation":"Kernel/driver upgrade plus host reboot. Until patched, avoid `ethtool -w` on Broadcom NICs in your automated diagnostics.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37911"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38080","cve":"CVE-2025-38080","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Increase block_sequence array size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38080","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38091","cve":"CVE-2025-38091","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: check stream id dml21 wrapper to get plane_id","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38091","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38098","cve":"CVE-2025-38098","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Don't treat wb connector as physical in create_validate_stream_for_sink","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38098","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38109","cve":"CVE-2025-38109","aliases":[],"title":"NVIDIA BlueField (mlx5 ECVF): Use-after-free during ECVF vport unload in the mlx5 driver — kernel-level memory corruption on the host…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA BlueField (mlx5 ECVF)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free during ECVF vport unload in the mlx5 driver — kernel-level memory corruption on the host attached to the DPU","attack_vector":"Local, host kernel","remediation":"Host kernel update plus driver refresh; requires a node drain because the mlx5 driver carries the tenant's data path","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38109"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-38352","cve":"CVE-2025-38352","aliases":[],"title":"Linux kernel (posix-cpu-timers): TOCTOU race between handle_posix_cpu_timers() and posix_cpu_timer_del() - local privilege escalation…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (posix-cpu-timers)","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"TOCTOU race between handle_posix_cpu_timers() and posix_cpu_timer_del() - local privilege escalation, exploited in the wild [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot. No capabilities or namespaces needed, so container hardening does not mitigate it","references":["https://access.redhat.com/security/cve/CVE-2025-38352"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38361","cve":"CVE-2025-38361","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check dce_hwseq before dereferencing it","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38361","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38389","cve":"CVE-2025-38389","aliases":[],"title":"Linux i915 GPU kernel driver (GT timeline / VMA allocation): A timeline is left held when VMA allocation fails, so the error path leaks a reference and the driver wedges.…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (GT timeline / VMA allocation)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A timeline is left held when VMA allocation fails, so the error path leaks a reference and the driver wedges. Shows up as hung GPU submissions and an unresponsive card after memory pressure - which on a busy training node is a common condition, not a rare one.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38389","https://git.kernel.org/stable/c/40e09506aea1fde1f3e0e04eca531bbb23404baf"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38455","cve":"CVE-2025-38455","aliases":[],"title":"Linux KVM/SVM - SEV/SEV-ES intra-host migration during vCPU creation: MULTI-TENANT ISOLATION: KVM permitted SEV/SEV-ES intra-host migration while vCPU creation was still in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM/SVM - SEV/SEV-ES intra-host migration during vCPU creation","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: KVM permitted SEV/SEV-ES intra-host migration while vCPU creation was still in flight, producing a race on confidential-VM state. Migration racing against vCPU setup means encrypted vCPU state can be moved or referenced while half-built - a route to host memory corruption driven from the VM lifecycle path.","attack_vector":"Through the KVM ioctl interface used for migration, reachable by the VMM process - so a compromised orchestrator or VMM.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. Interim control: disable intra-host migration for SEV guests in your VMM configuration.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38455"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38598","cve":"CVE-2025-38598","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix use-after-free in amdgpu_userq_suspend+0x51a/0x5a0","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38598","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38722","cve":"CVE-2025-38722","aliases":[],"title":"habanalabs kernel driver (dma-buf export path): MULTI-TENANT ISOLATION: A use-after-free in the habanalabs dma-buf export path: the driver installs a file…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (dma-buf export path)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the habanalabs dma-buf export path: the driver installs a file descriptor into the process table and then keeps using the object, so a second thread that closes the fd first frees memory the driver is still writing. dma-buf is exactly the mechanism used to hand accelerator memory to another process or device, so a successful exploit is a kernel-memory write reachable from an unprivileged accelerator user.","attack_vector":"Any local user with access to the habanalabs device node - on Kubernetes that is any pod granted a Gaudi device. Requires a deliberate race, not a lucky one.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38722","https://git.kernel.org/stable/c/33927f3d0ecdcff06326d6e4edb6166aed42811c"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39810","cve":"CVE-2025-39810","aliases":[],"title":"Linux bnxt_en driver (ring defaults vs traffic classes on ifdown): Memory corruption when firmware resources change while the interface is down, because the default-ring…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (ring defaults vs traffic classes on ifdown)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Memory corruption when firmware resources change while the interface is down, because the default-ring calculation assumes no traffic classes have been created. Relevant to RoCE clusters specifically: traffic classes are how you carve out lossless queues for RDMA with PFC, so any cluster running RoCEv2 with DCB has traffic classes configured and is in the affected configuration by construction.","attack_vector":"Local — an interface down/up cycle combined with a firmware resource change, on a host with traffic classes configured.","remediation":"Kernel/driver upgrade plus host reboot. Until then, avoid ifdown/ifup cycles on RoCE-configured Broadcom interfaces as a routine operational step — use link-level draining at the switch instead.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39810"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39906","cve":"CVE-2025-39906","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: remove oem i2c adapter on finish","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39906","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39979","cve":"CVE-2025-39979","aliases":["net/mlx5 fs, fix UAF in flow counter release"],"title":"Linux kernel mlx5_core flow steering / flow counters (hardware steering): Use-after-free releasing the hardware-steering action of a local flow counter - the refcount and mutex were…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core flow steering / flow counters (hardware steering)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Use-after-free releasing the hardware-steering action of a local flow counter - the refcount and mutex were never initialized and the counter struct can already be freed when the rule is deleted. Notably it is reached through an ib_uverbs ioctl into mlx5_ib_destroy_flow, so a tenant holding an RDMA verbs handle can drive host kernel memory corruption in the NIC's flow-steering tables.","attack_vector":"Local, low-privileged - reachable from a userspace RDMA verbs handle destroying a flow, not only from privileged tc/devlink paths.","remediation":"Upgrade the host kernel to 6.17 or the 6.16.10 stable backport. Rolling reboot of the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39979","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-39979.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40202","cve":"CVE-2025-40202","aliases":[],"title":"The Linux kernel's IPMI driver message-handling layer: A use-after-free in a kernel driver reachable from the host's IPMI device nodes gives a local attacker a path…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"The Linux kernel's IPMI driver message-handling layer","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in a kernel driver reachable from the host's IPMI device nodes gives a local attacker a path to kernel memory corruption and, from there, to privilege escalation on the host. On a bare-metal GPU node the significance is the direction of travel: this is a route from an unprivileged tenant process, through the kernel, toward the interface that talks to the BMC. It is the host-side half of the out-of-band security story, and it is the half that operators tend not to inventory because it lives in the kernel rather than in firmware. The per-user message limit was miscounted in several paths, producing a use-after-free. This is the in-kernel code every host uses to talk to its own BMC over the KCS or SSIF interface.","attack_vector":"A local process on the host with access to the IPMI character devices (/dev/ipmi*). On many stock server images those permissions are looser than they should be, and any container or tenant workload given access to them is in position.","remediation":"Kernel update and reboot - which on a GPU node means draining long-running training jobs, so it lands in the same expensive maintenance window as everything else kernel-level. There is a genuinely effective config-only mitigation that costs nothing: unless a workload needs in-band IPMI, do not expose /dev/ipmi* to it, and consider blacklisting the ipmi_devintf module entirely on tenant-facing bare metal. Most operators poll their BMCs over the network anyway and do not need the in-band path at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40202","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/40xxx/CVE-2025-40202.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40354","cve":"CVE-2025-40354","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: increase max link count and fix link->enc NULL pointer access","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40354","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-41244","cve":"CVE-2025-41244","aliases":[],"title":"VMware Aria Operations / VMware Tools: Local privilege escalation to root inside a managed VM via SDMP service discovery","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware Aria Operations / VMware Tools","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"Local privilege escalation to root inside a managed VM via SDMP service discovery; exploited in the wild since Oct 2024 [KEV]","attack_vector":"Local non-admin user inside a managed guest VM","remediation":"Update VMware Tools / Aria Operations in guest images; no host reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41244"],"status":"curated"},{"id":"CVE-2025-53000","cve":"CVE-2025-53000","aliases":[],"title":"nbconvert: Template-driven conversion executes attacker content","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"nbconvert","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Template-driven conversion executes attacker content","attack_vector":"Customer-supplied notebook converted by a pipeline","remediation":"Upgrade; report-generation pipelines ingest tenant notebooks","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-53000"],"status":"curated"},{"id":"CVE-2025-6018","cve":"CVE-2025-6018","aliases":[],"title":"PAM (pam-config): Local user is treated as `allow_active` in PAM - first half of a public unprivileged-to-root chain on…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"PAM (pam-config)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local user is treated as `allow_active` in PAM - first half of a public unprivileged-to-root chain on SUSE/openSUSE","attack_vector":"Local user (SSH session counts)","remediation":"Package update; no reboot. Chain with CVE-2025-6019 - patch both or neither is meaningful","references":["https://access.redhat.com/security/cve/CVE-2025-6018"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-67450","cve":"CVE-2025-67450","aliases":["ETN-VA-2025-1027"],"title":"Eaton UPS Companion (EUC) executable - library loading: Insecure library loading in the shipped executable gives arbitrary code execution to an attacker with access…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS Companion (EUC) executable - library loading","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insecure library loading in the shipped executable gives arbitrary code execution to an attacker with access to the software package. Persistent code on hosts that talk to the UPS, running with the privileges power-management agents typically hold.","attack_vector":"Local, requires access to the software package or its directory on the host.","remediation":"Update to the fixed EUC version. Lock down the install directory permissions - power-management agents are installed to writable locations more often than they should be.","references":["https://www.eaton.com/content/dam/eaton/company/news-insights/cybersecurity/security-bulletins/etn-va-2025-1027.pdf"],"status":"curated"},{"id":"CVE-2025-68174","cve":"CVE-2025-68174","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd): A division by zero in the amdkfd (KFD compute driver, /dev/kfd), reachable with attacker-influenced…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"A division by zero in the amdkfd (KFD compute driver, /dev/kfd), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: amd/amdkfd: enhance kfd process check in switch partition","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68174","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68286","cve":"CVE-2025-68286","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check NULL before accessing","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68286","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68793","cve":"CVE-2025-68793","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix a job->pasid access race in gpu recovery","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68793","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68801","cve":"CVE-2025-68801","aliases":["mlxsw spectrum_router fix neighbour use-after-free"],"title":"Linux kernel mlxsw (Spectrum switch router, neighbour table): The driver stored neighbour pointers without holding a reference, taking one only when the neighbour was used…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxsw (Spectrum switch router, neighbour table)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"The driver stored neighbour pointers without holding a reference, taking one only when the neighbour was used by a nexthop - a slab use-after-free when updating neighbour entries. Reproduced on an NVIDIA SN5600. Neighbour churn on a busy leaf is normal operation, so this is a switch crash that arrives on its own schedule and takes a rack's uplinks with it.","attack_vector":"Local on the switch, driven by neighbour table churn - which an attacker on an attached network can amplify by cycling ARP/ND entries.","remediation":"Upgrade the switch OS to a build carrying kernel 6.19 or a stable backport (5.10.248, 5.15.198, 6.1.160, 6.6.120, 6.12.64, 6.18.3). Switch OS upgrade and reload - fabric rolling window, one switch at a time with ECMP draining.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68801","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-68801.json"],"status":"curated"},{"id":"CVE-2025-71092","cve":"CVE-2025-71092","aliases":[],"title":"Linux bnxt_re RoCE driver (bnxt_re_copy_err_stats out-of-bounds write): Out-of-bounds write in the Broadcom RoCE driver's error-statistics copy, introduced when three RoCE hardware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (bnxt_re_copy_err_stats out-of-bounds write)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Out-of-bounds write in the Broadcom RoCE driver's error-statistics copy, introduced when three RoCE hardware counters were added past the end of the existing array. Reading RDMA counters is something monitoring agents do constantly on an AI cluster, so the vulnerable path runs on a schedule whether or not anyone attacks it.","attack_vector":"Triggered by reading RoCE hardware counters — reachable from any local process permitted to query RDMA statistics, including monitoring agents.","remediation":"Kernel/driver upgrade plus host reboot. Interim: stop polling RoCE hardware counters on Broadcom adapters, which costs you fabric observability — usually a worse trade than patching.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71092"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-71130","cve":"CVE-2025-71130","aliases":[],"title":"Linux i915 GPU kernel driver (execbuffer VMA array): MULTI-TENANT ISOLATION: The execbuffer VMA array was not zero-initialised, so uninitialised kernel stack/heap…","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Linux i915 GPU kernel driver (execbuffer VMA array)","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The execbuffer VMA array was not zero-initialised, so uninitialised kernel stack/heap contents could be acted on or leaked back through the GPU submission path. Info-leak-grade on its own, and useful as the KASLR-defeating first stage for one of the neighbouring i915 use-after-frees.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71130","https://git.kernel.org/stable/c/0336188cc85d0eab8463bd1bbd4ded4e9602de8b"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-8747","cve":"CVE-2025-8747","aliases":[],"title":"Keras: Safe-mode bypass in Keras 3.0.0–3.10.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras","year":"2025","cvss_score":7.8,"severity":"high","kev":false,"impact":"Safe-mode bypass in Keras 3.0.0–3.10.0 → arbitrary code execution","attack_vector":"Customer-supplied `.keras` file","remediation":"Upgrade past 3.10.0; bypass of the CVE-2025-1550 fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8747"],"status":"curated","fleet":{"ubiquity":"Very common - the bypass of the 3.9.0 fix, so operators who patched once are still exposed","remediation_pain":"**Image rebuild** again - the second round, which is exactly the fleet-wide-patch fatigue pattern","pain_class":"other","why_fleet_wide":"Reuse of internal functionality re-enables arbitrary code execution on `load_model`, proving the model-file-as-code class is not closable by a single patch"}},{"id":"CVE-2025-8875","cve":"CVE-2025-8875","aliases":[],"title":"N-able N-central: Deserialization of untrusted data allowing local code execution on the RMM server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2025","cvss_score":7.8,"severity":"high","kev":true,"impact":"[KEV] Deserialization of untrusted data allowing local code execution on the RMM server","attack_vector":"Local","remediation":"Control-plane: patch to 2025.3.1+; the RMM reaches every managed host","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8875"],"status":"curated"},{"id":"CVE-2026-12484","cve":"CVE-2026-12484","aliases":[],"title":"Keras (`TorchModuleWrapper`): Unsafe deserialization of attacker-controlled PyTorch pickle inside a Keras model","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (`TorchModuleWrapper`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Unsafe deserialization of attacker-controlled PyTorch pickle inside a Keras model","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 3.15.0; cross-framework pickle re-entry defeats Keras' own safe mode","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-12484"],"status":"curated"},{"id":"CVE-2026-1839","cve":"CVE-2026-1839","aliases":[],"title":"HuggingFace transformers (`Trainer._load_rng_state`): Arbitrary code execution when a training run resumes from a checkpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"HuggingFace transformers (`Trainer._load_rng_state`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Arbitrary code execution when a training run resumes from a checkpoint","attack_vector":"Customer-supplied or poisoned checkpoint directory in shared storage","remediation":"Tenant-owned code; provider must prevent cross-tenant writes to checkpoint directories","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1839"],"status":"curated"},{"id":"CVE-2026-23198","cve":"CVE-2026-23198","aliases":[],"title":"Linux KVM - irqfd routing type clobbered on deassign: MULTI-TENANT ISOLATION: Deassigning a KVM_IRQFD clobbers the irqfd's copy of the interrupt routing entry…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - irqfd routing type clobbered on deassign","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Deassigning a KVM_IRQFD clobbers the irqfd's copy of the interrupt routing entry, leaving stale or wrong routing behind. Interrupt routing decides which guest receives which interrupt, so corrupting it on teardown is a cross-VM correctness failure on the interrupt path - and on a GPU host, irqfd is exactly how passed-through accelerator interrupts reach their guest.","attack_vector":"Through the KVM ioctl interface, from the VMM process managing guests - reachable when devices are hot-unplugged or VMs torn down.","remediation":"Fixed in the Linux kernel. Distro kernel update plus host reboot; no firmware step. Relevant to any fleet doing GPU passthrough with dynamic device attach/detach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23198"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23856","cve":"CVE-2026-23856","aliases":["DSA-2026-077"],"title":"Dell iDRAC Service Module (iSM) for Windows and Linux: Improper access control in the host-side iDRAC Service Module lets a low-privilege local user escalate on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC Service Module (iSM) for Windows and Linux","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper access control in the host-side iDRAC Service Module lets a low-privilege local user escalate on the host. iSM is the agent that bridges the operating system to the iDRAC over the internal USB-NIC / passthrough channel, so it is the component that deliberately crosses the host-to-BMC boundary. Escalating through it gets an attacker host-level privilege and puts them next to a channel that talks to the service processor - the pivot direction that turns a tenant workload compromise into an out-of-band one. Affects iSM for Windows before 6.0.3.1 and iSM for Linux before 5.4.1.1.","attack_vector":"Host-side, local, low privilege - a tenant workload or any account on the operating system where iSM is installed. Not reachable from the management VLAN; the exposure is entirely inside the host.","remediation":"Upgrade the iSM package on the host to 6.0.3.1 (Windows) or 5.4.1.1 (Linux) or later. This is a host-side package update and a service restart, not a firmware flash - no host reboot in the normal case, so no job drain. Config-only alternative worth considering on nodes that do not need it: uninstall iSM entirely, or disable the iDRAC host USB-NIC passthrough (iDRAC Settings > OS to iDRAC Pass-through), which removes the host-to-BMC channel at the cost of in-band iDRAC access and some OS-level telemetry.","references":["https://www.dell.com/support/kbdoc/en-us/000426282/dsa-2026-077-security-update-for-dell-idrac-service-module-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2026-23856"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-24141","cve":"CVE-2026-24141","aliases":[],"title":"NVIDIA Model Optimizer: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Model Optimizer","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious model artifact","remediation":"Bump ModelOpt; rebuild optimization images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24141","https://github.com/NVIDIA/product-security/tree/main/2026/5798"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24149","cve":"CVE-2026-24149","aliases":[],"title":"NVIDIA Megatron-LM: A further script-level code-injection path, filed under the same class as the 2025 set. In an AI datacenter…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A further script-level code-injection path, filed under the same class as the 2025 set. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5712 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24149","https://github.com/NVIDIA/product-security/tree/main/2025/5712"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2026-24150","cve":"CVE-2026-24150","aliases":[],"title":"NVIDIA Megatron-LM: Checkpoint loading reaches remote code execution when a user loads a crafted checkpoint - the highest-value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Checkpoint loading reaches remote code execution when a user loads a crafted checkpoint - the highest-value path in this family, since checkpoints move between organisations routinely. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5769 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24150","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24151","cve":"CVE-2026-24151","aliases":[],"title":"NVIDIA Megatron-LM: The inferencing path reaches remote code execution on crafted input. In an AI datacenter this is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"The inferencing path reaches remote code execution on crafted input. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5769 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24151","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24152","cve":"CVE-2026-24152","aliases":[],"title":"NVIDIA Megatron-LM: A second checkpoint-loading remote code execution path. In an AI datacenter this is the model-and-data supply…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Megatron-LM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A second checkpoint-loading remote code execution path. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5769 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24152","https://github.com/NVIDIA/product-security/tree/main/2026/5769"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24155","cve":"CVE-2026-24155","aliases":[],"title":"NeMo Framework: RCE via malicious YAML deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via malicious YAML deserialization","attack_vector":"Malicious config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24155","https://github.com/NVIDIA/product-security/tree/main/2026/5839"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2026-24157","cve":"CVE-2026-24157","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24157","https://github.com/NVIDIA/product-security/tree/main/2026/5800"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24159","cve":"CVE-2026-24159","aliases":[],"title":"NeMo Framework: RCE via unsafe deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization on model load","attack_vector":"Malicious model","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24159","https://github.com/NVIDIA/product-security/tree/main/2026/5800"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24162","cve":"CVE-2026-24162","aliases":[],"title":"NVIDIA Merlin Transformers4Rec: Improper deserialization of untrusted data reaches code execution and information disclosure. In an AI…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Merlin Transformers4Rec","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper deserialization of untrusted data reaches code execution and information disclosure. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5838 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24162","https://github.com/NVIDIA/product-security/tree/main/2026/5838"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24165","cve":"CVE-2026-24165","aliases":[],"title":"BioNeMo Framework: RCE via malicious pickled data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via malicious pickled data","attack_vector":"Malicious model artifact","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24165","https://github.com/NVIDIA/product-security/tree/main/2026/5808"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24183","cve":"CVE-2026-24183","aliases":[],"title":"NVIDIA Cumulus Linux: Improper privilege management in the user-management component lets an unprivileged switch user escalate to…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Cumulus Linux","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Improper privilege management in the user-management component lets an unprivileged switch user escalate to full switch administration.","attack_vector":"Local on the switch, unprivileged. Anyone with any shell account on the Cumulus switch - including read-only monitoring accounts.","remediation":"Upgrade Cumulus Linux per bulletin 5817. Cost: switch reboot and link flap; sequence leaf-by-leaf so the fabric keeps ECMP paths. Audit who actually has shell on your switches while you are at it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24183","https://github.com/NVIDIA/product-security/tree/main/2026/5817"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24190","cve":"CVE-2026-24190","aliases":[],"title":"GPU Display Driver: Local privesc (insufficient permission checks)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (insufficient permission checks)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24190","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-862"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24191","cve":"CVE-2026-24191","aliases":[],"title":"GPU Display Driver: Local privesc (synchronization issue)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (synchronization issue)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24191","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-367"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24192","cve":"CVE-2026-24192","aliases":[],"title":"GPU Display Driver: Local privesc (integer overflow in address calc)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (integer overflow in address calc)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24192","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-681"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24193","cve":"CVE-2026-24193","aliases":[],"title":"GPU Display Driver: Local privesc (buffer overflow in GPU command processing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc (buffer overflow in GPU command processing)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24193","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24194","cve":"CVE-2026-24194","aliases":[],"title":"GPU Display Driver / GPU firmware: Local privesc via improper GPU firmware parameter validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver / GPU firmware","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Local privesc via improper GPU firmware parameter validation","attack_vector":"Any tenant with a container","remediation":"Driver + GPU firmware upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24194","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-281"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24216","cve":"CVE-2026-24216","aliases":[],"title":"BioNeMo Framework: RCE via insecure deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BioNeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization on model load","attack_vector":"Malicious model","remediation":"Bump BioNeMo; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24216","https://github.com/NVIDIA/product-security/tree/main/2026/5831"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24221","cve":"CVE-2026-24221","aliases":[],"title":"NVTabular: RCE via unsafe pickle deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVTabular","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe pickle deserialization","attack_vector":"Malicious dataset artifact","remediation":"Bump NVTabular; rebuild data images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24221","https://github.com/NVIDIA/product-security/tree/main/2026/5851"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24228","cve":"CVE-2026-24228","aliases":[],"title":"NeMo Framework: RCE via unsafe object deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe object deserialization","attack_vector":"Malicious config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24228","https://github.com/NVIDIA/product-security/tree/main/2026/5839"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24237","cve":"CVE-2026-24237","aliases":[],"title":"NVTabular: RCE via insecure deserialization in data processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVTabular","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization in data processing","attack_vector":"Malicious dataset artifact","remediation":"Bump NVTabular; rebuild data images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24237","https://github.com/NVIDIA/product-security/tree/main/2026/5851"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24238","cve":"CVE-2026-24238","aliases":[],"title":"TensorRT: Info disclosure / code exec (OOB read in tensor parsing)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Info disclosure / code exec (OOB read in tensor parsing)","attack_vector":"Malicious engine/model file","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24238","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-129"]},{"id":"CVE-2026-24240","cve":"CVE-2026-24240","aliases":[],"title":"Megatron-Bridge: RCE via insecure deserialization of model configs","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization of model configs","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24240","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24242","cve":"CVE-2026-24242","aliases":[],"title":"Megatron-Bridge: SSRF in file operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"SSRF in file operations","attack_vector":"Tenant-supplied checkpoint URI","remediation":"Bump Megatron-Bridge; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24242","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-918"]},{"id":"CVE-2026-24243","cve":"CVE-2026-24243","aliases":[],"title":"Megatron-Bridge: RCE via unsafe deserialization in weight loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization in weight loading","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24243","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24244","cve":"CVE-2026-24244","aliases":[],"title":"Megatron-Bridge: RCE via deserialization of untrusted model data","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via deserialization of untrusted model data","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24244","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24245","cve":"CVE-2026-24245","aliases":[],"title":"Megatron-Bridge: RCE via insecure config deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure config deserialization","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24245","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24246","cve":"CVE-2026-24246","aliases":[],"title":"Megatron-Bridge: Validation bypass via incorrect type comparison","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Validation bypass via incorrect type comparison","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24246","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-470"]},{"id":"CVE-2026-24247","cve":"CVE-2026-24247","aliases":[],"title":"Megatron-Bridge: RCE via unsafe deserialization in module loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization in module loading","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24247","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24248","cve":"CVE-2026-24248","aliases":[],"title":"Megatron-Bridge: RCE via malicious deserialized objects","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via malicious deserialized objects","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24248","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2026-24249","cve":"CVE-2026-24249","aliases":[],"title":"Megatron-Bridge: RCE via unsafe evaluation of loaded config","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via unsafe evaluation of loaded config","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24249","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"]},{"id":"CVE-2026-24250","cve":"CVE-2026-24250","aliases":[],"title":"NeMo Framework: Command injection in script processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Command injection in script processing","attack_vector":"Malicious config artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24250","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24251","cve":"CVE-2026-24251","aliases":[],"title":"Megatron-Bridge: RCE via deserialization of untrusted objects","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Megatron-Bridge","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via deserialization of untrusted objects","attack_vector":"Malicious checkpoint","remediation":"Bump Megatron-Bridge; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24251","https://github.com/NVIDIA/product-security/tree/main/2026/5841"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24252","cve":"CVE-2026-24252","aliases":[],"title":"NVIDIA NeMo Framework: OS command injection reaches code execution with the job's privileges. In an AI datacenter this is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"OS command injection reaches code execution with the job's privileges. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5839 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24252","https://github.com/NVIDIA/product-security/tree/main/2026/5839"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2026-24268","cve":"CVE-2026-24268","aliases":[],"title":"TensorRT: Code exec (OOB write in tensor manipulation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Code exec (OOB write in tensor manipulation)","attack_vector":"Malicious engine/model file","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24268","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-122"]},{"id":"CVE-2026-24272","cve":"CVE-2026-24272","aliases":[],"title":"TensorRT: Code exec (buffer overflow in tensor dimension handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Code exec (buffer overflow in tensor dimension handling)","attack_vector":"Malicious engine/model file","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24272","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-122"]},{"id":"CVE-2026-27905","cve":"CVE-2026-27905","aliases":[],"title":"BentoML (`safe_extract_tarfile`): Tar extraction escape despite the \"safe\" helper","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (`safe_extract_tarfile`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Tar extraction escape despite the \"safe\" helper","attack_vector":"Customer-supplied bento archive","remediation":"Upgrade to 1.4.36+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27905"],"status":"curated"},{"id":"CVE-2026-31431","cve":"CVE-2026-31431","aliases":[],"title":"Linux kernel (crypto algif_aead): Incorrect resource transfer between spheres in algif_aead (reverted to out-of-place operation) - local…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (crypto algif_aead)","year":"2026","cvss_score":7.8,"severity":"high","kev":true,"impact":"Incorrect resource transfer between spheres in algif_aead (reverted to out-of-place operation) - local privilege escalation, actively exploited [KEV]","attack_vector":"Any tenant process in a container (AF_ALG socket access)","remediation":"Livepatchable; otherwise drain + reboot. Compensating control: block AF_ALG in the default container seccomp profile","references":["https://access.redhat.com/security/cve/CVE-2026-31431"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31488","cve":"CVE-2026-31488","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: Do not skip unrelated mode changes in DSC validation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31488","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31566","cve":"CVE-2026-31566","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix fence put before wait in amdgpu_amdkfd_submit_ib","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31566","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31656","cve":"CVE-2026-31656","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: a refcount underflow in the engine heartbeat park path, which frees an engine still referenced.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31656","https://git.kernel.org/stable/c/2af8b200cae3fdd0e917ecc2753b28bb40c876c1"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-32305","cve":"CVE-2026-32305","aliases":[],"title":"Traefik: mTLS bypass via SNI pre-sniffing on fragmented ClientHello packets","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"mTLS bypass via SNI pre-sniffing on fragmented ClientHello packets","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-32305"],"status":"curated"},{"id":"CVE-2026-33298","cve":"CVE-2026-33298","aliases":[],"title":"llama.cpp (`ggml_nbytes`): Integer overflow in the core ggml size calculation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`ggml_nbytes`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow in the core ggml size calculation","attack_vector":"Customer-supplied model file","remediation":"Rebuild past b7824; affects every ggml-based downstream (whisper.cpp, stable-diffusion.cpp)","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33298"],"status":"curated"},{"id":"CVE-2026-33744","cve":"CVE-2026-33744","aliases":[],"title":"BentoML (`docker.system_packages`): Command injection through the package list field","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"BentoML (`docker.system_packages`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Command injection through the package list field","attack_vector":"Customer-supplied build config","remediation":"Upgrade to 1.4.37+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33744"],"status":"curated"},{"id":"CVE-2026-35051","cve":"CVE-2026-35051","aliases":[],"title":"Traefik: Authentication bypass in ForwardAuth when trustForwardHeader=false","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass in ForwardAuth when trustForwardHeader=false","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade to 2.11.43/3.6.14+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-35051"],"status":"curated"},{"id":"CVE-2026-39858","cve":"CVE-2026-39858","aliases":[],"title":"Traefik: Authentication bypass in ForwardAuth and snippet-based auth middleware","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass in ForwardAuth and snippet-based auth middleware","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39858"],"status":"curated"},{"id":"CVE-2026-3989","cve":"CVE-2026-3989","aliases":[],"title":"SGLang (`replay_request_dump.py`): Insecure `pickle.load()` on a `.pkl` dump","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`replay_request_dump.py`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Insecure `pickle.load()` on a `.pkl` dump","attack_vector":"Customer-supplied dump file replayed by an operator during debugging","remediation":"Upgrade; operator tooling that ingests tenant artifacts is a privilege-escalation path into the provider plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3989"],"status":"curated"},{"id":"CVE-2026-40912","cve":"CVE-2026-40912","aliases":[],"title":"Traefik: Authentication bypass via StripPrefixRegex middleware","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass via StripPrefixRegex middleware","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-40912"],"status":"curated"},{"id":"CVE-2026-43128","cve":"CVE-2026-43128","aliases":["RDMA/umem fix double dma_buf_unpin in failure path"],"title":"Linux kernel InfiniBand core dmabuf umem (GPUDirect RDMA path): When mapping a dmabuf-backed RDMA memory region fails, the dmabuf is unpinned immediately but the pinned flag…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand core dmabuf umem (GPUDirect RDMA path)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"When mapping a dmabuf-backed RDMA memory region fails, the dmabuf is unpinned immediately but the pinned flag is left set, so release unpins it a second time. This is the dmabuf path that GPUDirect RDMA uses to register GPU memory for the NIC - a double-unpin here corrupts the pinning state of buffers that sit between GPU memory and the fabric, exactly the code an operator most wants to be boring.","attack_vector":"Local, low-privileged - a tenant process registering GPU memory for RDMA via dmabuf, able to induce a mapping failure (resource pressure, invalid parameters).","remediation":"Upgrade the host kernel to 7.0 or a stable backport (6.1.165, 6.6.128, 6.12.75, 6.18.16, 6.19.6). Rolling reboot of GPU nodes using GPUDirect RDMA over dmabuf.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43128","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-43128.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43206","cve":"CVE-2026-43206","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix out-of-bounds write in kfd_event_page_set()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43206","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43237","cve":"CVE-2026-43237","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Refactor amdgpu_gem_va_ioctl for Handling Last Fence Update and Timeline Management v4","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43237","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43260","cve":"CVE-2026-43260","aliases":[],"title":"Linux bnxt_en driver (RSS context delete logic): RSS contexts are not always freed in firmware when the driver deletes them, leaving stale VNIC state on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (RSS context delete logic)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RSS contexts are not always freed in firmware when the driver deletes them, leaving stale VNIC state on the adapter. Stale receive-steering state on a NIC is worth flagging in a multi-tenant context: RSS contexts and their VNICs determine which queues — and therefore which owner — receives which packets, and leaked contexts are exactly the kind of residue that should not survive a tenant teardown.","attack_vector":"Local, via repeated RSS context create/delete cycles from the host.","remediation":"Kernel/driver upgrade plus host reboot. On bare-metal handoff, a cold power cycle of the NIC clears residual adapter state regardless of driver version — worth doing between tenants anyway.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43260"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43370","cve":"CVE-2026-43370","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A use-after-free in the amdgpu firmware, ACPI and IP-block initialisation. Freed kernel memory is reachable…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu firmware, ACPI and IP-block initialisation. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix use-after-free race in VM acquire","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43370","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43627","cve":"CVE-2026-43627","aliases":[],"title":"llama.cpp (`llama_batch_init`): Integer overflow from unchecked multiplication","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`llama_batch_init`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Integer overflow from unchecked multiplication","attack_vector":"Tenant-controlled batch parameters","remediation":"Rebuild past b9058","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43627"],"status":"curated"},{"id":"CVE-2026-4372","cve":"CVE-2026-4372","aliases":[],"title":"HuggingFace transformers: Critical RCE in all versions before 5.3.0","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"HuggingFace transformers","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Critical RCE in all versions before 5.3.0","attack_vector":"Customer-supplied model repo loaded by `from_pretrained`","remediation":"Rebuild every image with transformers >= 5.3.0. Tenant-pinned versions are outside provider control","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-4372"],"status":"curated"},{"id":"CVE-2026-45853","cve":"CVE-2026-45853","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Use kvfree instead of kfree in amdgpu_gmc_get_nps_memranges()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45853","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-45878","cve":"CVE-2026-45878","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix watch_id bounds checking in debug address watch v2","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45878","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46197","cve":"CVE-2026-46197","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: validate SVM ioctl nattr against buffer size","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46197","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46263","cve":"CVE-2026-46263","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix out-of-bounds stream encoder index v3","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46263","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46311","cve":"CVE-2026-46311","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A correctness defect in the amdgpu user-mode queues (doorbell submission path) reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu user-mode queues (doorbell submission path) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/userq: fix access to stale wptr mapping","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46311","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-47472","cve":"CVE-2026-47472","aliases":[],"title":"TensorRT-LLM: RCE via insecure deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE via insecure deserialization on model load","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47472","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-47749","cve":"CVE-2026-47749","aliases":[],"title":"stable-diffusion.cpp: Memory-safety flaw in model loading","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"stable-diffusion.cpp","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Memory-safety flaw in model loading","attack_vector":"Customer-supplied diffusion model file","remediation":"Rebuild; inherits the ggml parser class","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47749"],"status":"curated"},{"id":"CVE-2026-48020","cve":"CVE-2026-48020","aliases":[],"title":"Traefik: StripPrefix middleware allows route-level authentication bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"StripPrefix middleware allows route-level authentication bypass","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade to 2.11.48/3.6.19/3.7.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-48020"],"status":"curated"},{"id":"CVE-2026-52987","cve":"CVE-2026-52987","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): MULTI-TENANT ISOLATION: A double free in the amdgpu user-mode queues (doorbell submission path). The same…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A double free in the amdgpu user-mode queues (doorbell submission path). The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amdgpu: avoid double drm_exec_fini() in userq validate","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-52987","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53136","cve":"CVE-2026-53136","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Clamp VBIOS HDMI retimer register count to array size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53136","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53137","cve":"CVE-2026-53137","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Clamp HDMI HDCP2 rx_id_list read to buffer size","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53137","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53143","cve":"CVE-2026-53143","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix buffer overflow in SDMA queue checkpoint/restore on GFX11","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53143","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53622","cve":"CVE-2026-53622","aliases":[],"title":"Traefik: HTTP/3 QUIC TLS configuration selection lets clients bypass router-specific mTLS enforcement","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"HTTP/3 QUIC TLS configuration selection lets clients bypass router-specific mTLS enforcement","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade to 3.7.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53622"],"status":"curated"},{"id":"CVE-2026-58659","cve":"CVE-2026-58659","aliases":[],"title":"PyTorch Lightning (`_load_state`): RCE by importing and executing classes named in the checkpoint","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch Lightning (`_load_state`)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"RCE by importing and executing classes named in the checkpoint","attack_vector":"Customer-supplied checkpoint","remediation":"No format-level fix — the checkpoint format is the vulnerability. Enforce safetensors","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-58659"],"status":"curated"},{"id":"CVE-2026-63840","cve":"CVE-2026-63840","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v5.3.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63840","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63841","cve":"CVE-2026-63841","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/jpeg): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v5.0.1 ring","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63841","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63842","cve":"CVE-2026-63842","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v5.0.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63842","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63843","cve":"CVE-2026-63843","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v4.0.5 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63843","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63844","cve":"CVE-2026-63844","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v4.0.3 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63844","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63845","cve":"CVE-2026-63845","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v4.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63845","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63846","cve":"CVE-2026-63846","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v3.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63846","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63847","cve":"CVE-2026-63847","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v2.5 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63847","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63848","cve":"CVE-2026-63848","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/jpeg)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/jpeg: set no_user_fence for JPEG v2.0 ring","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63848","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63849","cve":"CVE-2026-63849","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v5.0.1 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63849","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63850","cve":"CVE-2026-63850","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v5.0.0 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63850","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63851","cve":"CVE-2026-63851","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v4.0.5 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63851","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63852","cve":"CVE-2026-63852","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v4.0.3 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63852","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63853","cve":"CVE-2026-63853","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v4.0 enc ring","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63853","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63854","cve":"CVE-2026-63854","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v3.0 enc/dec rings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63854","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63855","cve":"CVE-2026-63855","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v2.5 enc/dec rings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63855","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63856","cve":"CVE-2026-63856","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/vcn: set no_user_fence for VCN v2.0 enc/dec rings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63856","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63879","cve":"CVE-2026-63879","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: fix amdgpu_hmm_range_get_pages","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63879","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63881","cve":"CVE-2026-63881","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: fix a vulnerability of integer overflow in kfd debugger","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63881","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63884","cve":"CVE-2026-63884","aliases":[],"title":"Linux i915 GPU kernel driver: MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the i915 GPU kernel driver. The general shape is that a GPU object is freed on one path while another path still holds a reference to it, so a local user with GPU access can get the kernel to read or write freed memory. Exploitability varies by heap layout, but on a GPU node every such bug is reachable from inside a container that was granted /dev/dri - the same boundary that is supposed to separate tenants. Specific trigger: the TTM buffer object being swapped out from under a purge, so the purge frees the wrong object.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63884","https://git.kernel.org/stable/c/073bcbc95e9648c976da1654c7590a8d6ee12c2d"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64097","cve":"CVE-2026-64097","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Validate GPIO pin LUT table size before iterating","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64097","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-65600","cve":"CVE-2026-65600","aliases":[],"title":"Traefik: Authentication bypass via path traversal in ReplacePathRegex","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"Authentication bypass via path traversal in ReplacePathRegex","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65600"],"status":"curated"},{"id":"CVE-2026-68104","cve":"CVE-2026-68104","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: invoke pm_genpd_remove() before freeing genpd","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68104","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68106","cve":"CVE-2026-68106","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A division by zero in the amdgpu firmware, ACPI and IP-block initialisation, reachable with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A division by zero in the amdgpu firmware, ACPI and IP-block initialisation, reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amdgpu: fix division by zero with invalid uvd dimensions","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68106","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68236","cve":"CVE-2026-68236","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu display core (DC/DM). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/display: set new_stream to NULL after release","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68236","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68245","cve":"CVE-2026-68245","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu GEM/VM/command-submission ioctl surface. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: fix lifetime issue of amdgpu_vm_get_task_info_pasid()","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68245","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68257","cve":"CVE-2026-68257","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: fix 32-bit overflow in CWSR total size calculation","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68257","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68273","cve":"CVE-2026-68273","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): MULTI-TENANT ISOLATION: A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdgpu RAS / GPU reset and recovery path. Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdgpu: Fix context pstate override handling","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68273","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72072","cve":"CVE-2026-72072","aliases":[],"title":"Linux kernel mlx5_core MACsec offload: TENANT ISOLATION: deleting an offloaded MACsec RX secure channel frees the per-SC metadata_dst with a call…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core MACsec offload","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: deleting an offloaded MACsec RX secure channel frees the per-SC metadata_dst with a call that ignores the reference count, while the RX datapath is concurrently taking a reference on it under RCU - a use-after-free reachable from the packet path. The kernel CNA notes it is reachable by cloud tenants with SR-IOV VFs, containers, or user/network namespaces without init-namespace root, so this is a container-to-host kernel corruption on nodes doing link-layer encryption.","attack_vector":"A tenant with an SR-IOV VF, a container with network-namespace capability, or a local user in a user namespace - combined with MACsec RX secure channel churn on an mlx5 interface.","remediation":"Upgrade the host kernel to 7.2 or a stable backport (6.1.178, 6.6.145, 6.12.97, 6.18.40, 7.1.5). Rolling reboot. Interim: restrict unprivileged user namespaces and do not delegate MACsec configuration into tenant namespaces.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72072","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-72072.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72449","cve":"CVE-2026-72449","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the amdkfd (KFD compute driver, /dev/kfd). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amdkfd: fix list_del corruption in kfd_criu_resume_svm","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72449","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74297","cve":"CVE-2026-74297","aliases":["RDMA/mlx5 fix undefined shift of user RQ WQE size"],"title":"Linux kernel mlx5_ib (queue pair sizing): set_rq_size() computes the receive-queue work-entry size as 1 << rq_wqe_shift from a user-supplied shift that…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (queue pair sizing)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"set_rq_size() computes the receive-queue work-entry size as 1 << rq_wqe_shift from a user-supplied shift that is only checked against being greater than 32, so shifts of 31 and 32 pass and overflow. A tenant process controls a shift that the kernel then uses to size an allocation - the classic setup for heap corruption from an RDMA verbs call.","attack_vector":"Local, low-privileged - any process that can create an RDMA queue pair on an mlx5 device.","remediation":"Upgrade the host kernel to 7.2 or a stable backport (5.10.261, 5.15.212, 6.1.178, 6.6.145, 6.12.97, 6.18.40, 7.1.5). Rolling reboot across the RDMA fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74297","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-74297.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74357","cve":"CVE-2026-74357","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu RAS / GPU reset and recovery path - a length…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu RAS / GPU reset and recovery path - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix KASAN slab-out-of-bounds in amdgpu_coredump ring dump","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74357","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74378","cve":"CVE-2026-74378","aliases":["RDMA/rxe get_srq_wqe TOCTOU","Soft-RoCE shared receive queue heap overflow"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c: TENANT ISOLATION: The shared receive queue buffer is mapped into userspace, and get_srq_wqe() validates the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"TENANT ISOLATION: The shared receive queue buffer is mapped into userspace, and get_srq_wqe() validates the num_sge field from that buffer against max_sge but then re-reads the same field to size the memcpy. A local process racing a second thread flips num_sge between the check and the copy and overflows the kernel heap. On a multi-tenant node this is a container-to-kernel escalation reachable by any tenant given an RDMA device - which, for RDMA to be useful at all, is every tenant. Full loss of confidentiality, integrity and availability on the host, meaning escape from the container onto the GPU node and everything else scheduled on it.","attack_vector":"Local, low privilege. The attacker maps the SRQ, posts a work queue entry, and has a concurrent thread rewrite num_sge in the shared mapping in the window between validation and use. Classic time-of-check-to-time-of-use on a userspace-writable structure the kernel reads twice. Requires the rdma_rxe module to be loaded and an rxe device present.","remediation":"Host reboot / kernel upgrade. Strong interim mitigation: rxe (Soft-RoCE) is a software RoCE emulation used for development and for nodes without a real RNIC - on a GPU cluster with ConnectX or BlueField hardware it is almost never needed. Unload and blacklist rdma_rxe (config change, no downtime) and this and the other rxe findings in this database all disappear at once. Audit which nodes have it loaded; it is frequently pulled in by test tooling and left behind.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-74378.json","https://nvd.nist.gov/vuln/detail/CVE-2026-74378"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74446","cve":"CVE-2026-74446","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: hold event_mutex while checkpointing CRIU events","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74446","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74447","cve":"CVE-2026-74447","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: fix uint32_t overflow in EOP ring buffer size alignment","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74447","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74449","cve":"CVE-2026-74449","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix divide-by-zero in calculate_mcache_setting on zero viewport","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74449","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74450","cve":"CVE-2026-74450","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A use-after-free in the amdgpu power management (SMU/powerplay). Freed kernel memory is reachable again…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2026","cvss_score":7.8,"severity":"high","kev":false,"impact":"A use-after-free in the amdgpu power management (SMU/powerplay). Freed kernel memory is reachable again through a later operation, so an attacker who can win the timing and reoccupy the freed slab object gets a write (or a controlled read) into live kernel memory. In practice this is a local-privilege-escalation primitive: from inside a GPU container it is a route to host kernel code execution, and from there to every other tenant's GPU memory and data on the node. Unexploited, it is a kernel panic that takes the whole node down mid-job. Upstream fix: drm/amd/pm: fix pptable use-after-free","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74450","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-15114","cve":"CVE-2020-15114","aliases":[],"title":"etcd: Gateway can be pointed at itself, causing an infinite loop and control-plane DoS","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd","year":"2020","cvss_score":7.7,"severity":"high","kev":false,"impact":"Gateway can be pointed at itself, causing an infinite loop and control-plane DoS","attack_vector":"Anyone who can influence etcd gateway config or DNS","remediation":"Rolling etcd upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15114"],"status":"curated"},{"id":"CVE-2022-24348","cve":"CVE-2022-24348","aliases":[],"title":"Argo CD: Directory traversal via Helm charts discloses credentials from other Applications' value files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Directory traversal via Helm charts discloses credentials from other Applications' value files; cross-tenant secret leak","attack_vector":"Any user with repo write access","remediation":"Rolling Argo CD upgrade; rotate any credentials stored in chart values","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24348"],"status":"curated"},{"id":"CVE-2022-24730","cve":"CVE-2022-24730","aliases":[],"title":"Argo CD: Path traversal plus improper access control in the repo-server","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Path traversal plus improper access control in the repo-server","attack_vector":"Any user with repo access","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24730"],"status":"curated"},{"id":"CVE-2022-28183","cve":"CVE-2022-28183","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An out-of-bounds read in the kernel mode layer leaks kernel memory to an unprivileged local user. Both the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"An out-of-bounds read in the kernel mode layer leaks kernel memory to an unprivileged local user. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28183","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-31666","cve":"CVE-2022-31666","aliases":[],"title":"Harbor: Missing permission validation on Webhook policies","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Missing permission validation on Webhook policies; view, update and delete another tenant's webhooks","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31666"],"status":"curated"},{"id":"CVE-2022-31670","cve":"CVE-2022-31670","aliases":[],"title":"Harbor: Missing permission validation on tag retention policies across projects","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Missing permission validation on tag retention policies across projects","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31670"],"status":"curated"},{"id":"CVE-2022-42275","cve":"CVE-2022-42275","aliases":[],"title":"DGX servers BMC: Improper access control on BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":7.7,"severity":"high","kev":false,"impact":"Improper access control on BMC","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42275","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-288","CWE-120"]},{"id":"CVE-2023-34335","cve":"CVE-2023-34335","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx 13 (IPMI handler / host SPI flash path): The multi-tenant bare-metal nightmare. An unauthenticated host - meaning code running on the server's own OS…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 13 (IPMI handler / host SPI flash path)","year":"2023","cvss_score":7.7,"severity":"high","kev":false,"impact":"The multi-tenant bare-metal nightmare. An unauthenticated host - meaning code running on the server's own OS with no BMC credentials at all - can write the host's SPI flash through the BMC's IPMI handler, bypassing Secure Boot. A tenant who rents a bare-metal GPU node for an hour can leave a BIOS-resident implant behind that the next tenant inherits, and that no OS reimage, disk wipe or node rebuild will find. For anyone selling bare-metal GPU capacity this is a tenant-isolation break, not just a firmware bug.","attack_vector":"Requires code execution on the tenant/host OS (root or equivalent), then pivots inboard over the host-BMC interface - KCS/LPC or the in-band IPMI channel - to reach the BMC's flash write path. No BMC password, no management-network access, and no physical presence. The BMC VLAN being airtight does not help you here, because the attack comes from the host side.","remediation":"BMC firmware flash to MegaRAC SPx_13.5 or later; out-of-band, per node, ODM-gated. Because the attack path is in-band, network segmentation buys you nothing - the only other lever is disabling the in-band host-to-BMC IPMI interface (KCS) in BIOS setup, which is a BIOS config change plus a reboot and will break in-band ipmitool, node health agents and most vendor management agents running on the host. If you sell bare-metal, treat this as a wipe-and-reflash-BIOS-between-tenants question rather than a patch question.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34335"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-3056","cve":"CVE-2024-3056","aliases":[],"title":"Podman: Crafted container sharing IPC creates unbounded IPC resources in /dev/shm","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"Crafted container sharing IPC creates unbounded IPC resources in /dev/shm; node resource exhaustion","attack_vector":"Any tenant workload sharing an IPC namespace","remediation":"Upgrade Podman; disable shared IPC across tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3056"],"status":"curated"},{"id":"CVE-2024-3095","cve":"CVE-2024-3095","aliases":[],"title":"LangChain (Web Research Retriever): SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LangChain (Web Research Retriever)","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"SSRF","attack_vector":"Attacker-supplied URL or retrieved content","remediation":"Upgrade; block internal egress","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-3095"],"status":"curated"},{"id":"CVE-2024-31410","cve":"CVE-2024-31410","aliases":[],"title":"CyberPower PowerPanel managed devices - shared device certificates: Every managed device uses an identical certificate derived from a hardcoded key, so any device can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel managed devices - shared device certificates","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"Every managed device uses an identical certificate derived from a hardcoded key, so any device can impersonate any other. An attacker who compromises one PDU in one rack can pose as every other device in the estate and feed the DCIM whatever telemetry they like - including telling it everything is fine while a hall overheats, or triggering automated responses that are themselves PHYSICAL actions.","attack_vector":"Anyone who obtains the key - which means anyone who obtains any single managed device, including a unit bought secondhand or pulled from an RMA pile.","remediation":"Vendor firmware and platform upgrade that issues per-device certificates. Until then, the device identity layer provides no assurance and you should not build automated power actions on top of it. Note this also breaks the trust assumption in your decommissioning process: a device leaving your estate carries the fleet key with it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31410"],"status":"curated"},{"id":"CVE-2024-8698","cve":"CVE-2024-8698","aliases":[],"title":"Keycloak: SAML signature scope determined by position, not Reference","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2024","cvss_score":7.7,"severity":"high","kev":false,"impact":"SAML signature scope determined by position, not Reference -> signature validation bypass and impersonation","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the IdP; revalidate every SAML client config","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-8698"],"status":"curated"},{"id":"CVE-2025-3937","cve":"CVE-2025-3937","aliases":["CVE-2025-3936","CVE-2025-3944","CVE-2025-3945","CVE-2025-3938","CVE-2025-3943"],"title":"Tridium Niagara Framework and Niagara Enterprise Security (before 4.10.11 / 4.14.2 / 4.15.1): A chain, not a single bug, and the chain is what matters. Password hashes are stored with insufficient…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tridium Niagara Framework and Niagara Enterprise Security (before 4.10.11 / 4.14.2 / 4.15.1)","year":"2025","cvss_score":7.7,"severity":"high","kev":false,"impact":"A chain, not a single bug, and the chain is what matters. Password hashes are stored with insufficient computational effort so they crack offline; a missing cryptographic step and an observable response discrepancy help an attacker recover or confirm credentials; incorrect permission assignment on the QNX-based JACE controllers allows file manipulation; and an argument-injection path on QNX turns that into command execution. Put together, an attacker with a foothold anywhere near the station escalates to control of the Niagara supervisor and the JACE field controllers underneath it - which is direct write access to cooling commands and setpoints for the whole site. Niagara Enterprise Security is affected too, so on sites that use it the same chain reaches door control. The operator-facing consequence is loss of thermal control over the hall plus, potentially, loss of the physical access boundary around the cages, from one credential-recovery weakness.","attack_vector":"Requires reaching the Niagara station or capturing its authentication material - so a compromised facilities workstation, a foothold on the building network, an integrator's remote path, or an internet-exposed station. The QNX permission and argument-injection pieces then apply on the JACE hardware controllers themselves, which sit on the facility VLAN and are rarely monitored by anyone.","remediation":"Upgrade the framework to 4.10.11, 4.14.2 or 4.15.1 (or later) per the Tridium/Honeywell tech bulletins. This is a supervisor software upgrade plus a firmware push to every JACE, done by the integrator - a real project, typically a scheduled outage of supervisory control while field controllers keep running standalone. After upgrading, rotate every Niagara credential, because the old hashes are assumed compromised. Compensating controls while you wait: restrict the station to a jump host, enforce MFA on the path to it, and put the JACEs behind an allow-list. If your Niagara estate is integrator-managed under a service contract, the upgrade is a billable engagement - price it now rather than discovering it during an incident.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3937","https://nvd.nist.gov/vuln/detail/CVE-2025-3945","https://docs.niagara-community.com/category/tech_bull","https://www.honeywell.com/us/en/product-security"],"status":"curated"},{"id":"CVE-2026-24177","cve":"CVE-2026-24177","aliases":[],"title":"KAI Scheduler: Missing authentication on API endpoints","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"KAI Scheduler","year":"2026","cvss_score":7.7,"severity":"high","kev":false,"impact":"Missing authentication on API endpoints -> scheduler manipulation","attack_vector":"Network-adjacent attacker inside the cluster","remediation":"Upgrade KAI Scheduler chart; add network policy in front of the API","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24177","https://github.com/NVIDIA/product-security/tree/main/2026/5818"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-306"]},{"id":"CVE-2026-9804","cve":"CVE-2026-9804","aliases":[],"title":"KubeVirt: Symlink path traversal in the virt-exportserver VMExport directory endpoint","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":7.7,"severity":"high","kev":false,"impact":"Symlink path traversal in the virt-exportserver VMExport directory endpoint","attack_vector":"Cluster user with namespace access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-9804"],"status":"curated"},{"id":"CVE-2018-3652","cve":"CVE-2018-3652","aliases":["INTEL-SA-00127"],"title":"Intel DCI (Direct Connect Interface) UEFI setting restrictions - Xeon E3 v5/v6, Xeon Scalable, Xeon D: The UEFI setting that is supposed to lock out DCI can be bypassed, re-enabling Intel's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel DCI (Direct Connect Interface) UEFI setting restrictions - Xeon E3 v5/v6, Xeon Scalable, Xeon D","year":"2018","cvss_score":7.6,"severity":"high","kev":false,"impact":"The UEFI setting that is supposed to lock out DCI can be bypassed, re-enabling Intel's closed-chassis debug path. DCI exposes JTAG-class control of the CPU over a USB port: halt cores, read and write all of physical memory including anything a tenant left resident, single-step SMM, and modify the boot chain. This is the strongest possible below-the-OS position on a node and everything it plants survives a reimage. It is squarely a tenant-handoff and colocation problem - a departing tenant with physical or smart-hands access to the chassis can leave the debug door open for the next occupant's data.","attack_vector":"Physical or near-physical access to a USB3 port on the node (or to a KVM/USB-over-IP appliance wired to one). Relevant wherever your halls have shared cages, third-party remote-hands, or hardware that transits an untrusted logistics chain.","remediation":"BIOS update from the OEM (Dell, HPE, Supermicro, Lenovo, Quanta, Wiwynn) that correctly enforces the DCI lock - reboot and drain required, and OEM releases for Xeon Scalable server boards lagged Intel's July 2018 advisory considerably. Alongside the flash: set and lock the BIOS admin password, disable DCI and USB debug explicitly in the BIOS profile, physically disable or block front-panel USB on production nodes, and treat any node that has left your physical custody as needing a full firmware re-flash plus measured-boot re-attestation before it goes back into a tenant pool.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3652","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00127.html","https://security.netapp.com/advisory/ntap-20180802-0001/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-25742","cve":"CVE-2021-25742","aliases":[],"title":"ingress-nginx: Custom nginx snippets in an Ingress annotation retrieve the ingress-nginx service-account token and therefore…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Custom nginx snippets in an Ingress annotation retrieve the ingress-nginx service-account token and therefore every Secret in the cluster","attack_vector":"Cluster user with namespace access who can create Ingress objects","remediation":"Rolling controller upgrade and set `allow-snippet-annotations: false`; rotate all cluster Secrets if exploited","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25742"],"status":"curated"},{"id":"CVE-2021-25745","cve":"CVE-2021-25745","aliases":[],"title":"ingress-nginx: Ingress `path` can be pointed at the service-account token file","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Ingress `path` can be pointed at the service-account token file","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; enable path validation","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25745"],"status":"curated"},{"id":"CVE-2021-25746","cve":"CVE-2021-25746","aliases":[],"title":"ingress-nginx: Directive injection through Ingress annotations obtains controller credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Directive injection through Ingress annotations obtains controller credentials","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; enable annotation validation","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25746"],"status":"curated"},{"id":"CVE-2021-25748","cve":"CVE-2021-25748","aliases":[],"title":"ingress-nginx: Newline character bypasses `path` sanitization","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2021","cvss_score":7.6,"severity":"high","kev":false,"impact":"Newline character bypasses `path` sanitization","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25748"],"status":"curated"},{"id":"CVE-2022-39388","cve":"CVE-2022-39388","aliases":[],"title":"Istio: Localhost access to the istiod pod lets a user impersonate any workload identity in the mesh","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":7.6,"severity":"high","kev":false,"impact":"Localhost access to the istiod pod lets a user impersonate any workload identity in the mesh","attack_vector":"An attacker with a foothold in the istiod pod","remediation":"Rolling istiod upgrade; restrict exec into istio-system","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-39388"],"status":"curated"},{"id":"CVE-2022-43758","cve":"CVE-2022-43758","aliases":[],"title":"Rancher: OS command injection through an untrusted Helm catalog URL","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":7.6,"severity":"high","kev":false,"impact":"OS command injection through an untrusted Helm catalog URL","attack_vector":"A user who can add a Helm catalog","remediation":"Upgrade Rancher; restrict catalog sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43758"],"status":"curated"},{"id":"CVE-2023-25531","cve":"CVE-2023-25531","aliases":[],"title":"DGX H100 BMC (IPMI): Credential exposure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"Credential exposure","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18; rotate IPMI credentials; disable IPMI where possible","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-522"]},{"id":"CVE-2023-34337","cve":"CVE-2023-34337","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC cryptography / HMAC): The BMC uses inadequate HMAC strength, so an attacker positioned on the management network can forge or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC cryptography / HMAC)","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"The BMC uses inadequate HMAC strength, so an attacker positioned on the management network can forge or replay authenticated material rather than having to break in. The practical outcome is impersonating a legitimate management session and issuing privileged BMC operations - power, virtual media, firmware update - without ever holding a valid password. AMI scores the scope as changed, i.e. the consequences land outside the BMC.","attack_vector":"Adjacent network with a low-privilege foothold and some user interaction, at high attack complexity. Realistically this is an attacker already sitting on the management VLAN who can observe or interpose on BMC traffic, e.g. after compromising a management jump host or a switch on that segment.","remediation":"Firmware flash to SPx_12.2 / SPx_13.0 or later - this one has been fixed for a long time, so the operator question is whether your ODM image is actually from a fixed branch, not whether AMI shipped a fix. Audit the running BMC build across the fleet before assuming it. Until then, force TLS everywhere on the BMC, disable the plain-HTTP and legacy management listeners, and keep the management plane on its own switched segment so there is nowhere to interpose.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34337"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-39347","cve":"CVE-2023-39347","aliases":[],"title":"Cilium: An attacker able to update pod labels causes Cilium to apply the wrong network policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"An attacker able to update pod labels causes Cilium to apply the wrong network policy","attack_vector":"Cluster user with namespace access","remediation":"Rolling Cilium upgrade; restrict pod label patch RBAC","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39347"],"status":"curated"},{"id":"CVE-2023-5043","cve":"CVE-2023-5043","aliases":[],"title":"ingress-nginx: Annotation injection causes arbitrary command execution in the controller pod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"Annotation injection causes arbitrary command execution in the controller pod","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade; disable snippet annotations","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5043"],"status":"curated"},{"id":"CVE-2023-5044","cve":"CVE-2023-5044","aliases":[],"title":"ingress-nginx: Code injection via the permanent-redirect annotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"Code injection via the permanent-redirect annotation","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5044"],"status":"curated"},{"id":"CVE-2023-5077","cve":"CVE-2023-5077","aliases":[],"title":"HashiCorp Vault: GCP secrets engine drops existing IAM Conditions when creating/updating rolesets","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2023","cvss_score":7.6,"severity":"high","kev":false,"impact":"GCP secrets engine drops existing IAM Conditions when creating/updating rolesets -> over-broad cloud grants","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + re-apply IAM conditions on all GCP rolesets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5077"],"status":"curated"},{"id":"CVE-2024-0122","cve":"CVE-2024-0122","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): An unauthorised action against the DLS reaches partial denial of service and disclosure of confidential…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"An unauthorised action against the DLS reaches partial denial of service and disclosure of confidential information. A DLS outage eventually strands vGPU guests when cached licences expire, so this is an availability problem for the whole vGPU estate, not just a management-plane nuisance.","attack_vector":"Adjacent network, no privileges required. Anyone with a route to the licensing appliance - which is usually the same management VLAN as everything else.","remediation":"Update the DLS appliance per bulletin 5570. Cost: appliance restart only. Confirm your licence lease duration so you know how long guests survive a DLS outage before this becomes a tenant-visible incident.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0122","https://github.com/NVIDIA/product-security/tree/main/2024/5570"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-862"]},{"id":"CVE-2024-0135","cve":"CVE-2024-0135","aliases":[],"title":"Container Toolkit / GPU Operator: Container escape to host root (insufficient input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"Container escape to host root (insufficient input validation)","attack_vector":"Any tenant that can run a container image on a GPU node","remediation":"Bump nvidia-container-toolkit package + restart containerd/docker; upgrade GPU Operator Helm chart; evict running tenant containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0135","https://github.com/NVIDIA/product-security/tree/main/2025/5599"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-653"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0136","cve":"CVE-2024-0136","aliases":[],"title":"Container Toolkit / GPU Operator: Container escape to host (insufficient input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"Container escape to host (insufficient input validation)","attack_vector":"Any tenant with a container","remediation":"Bump toolkit + restart runtime; upgrade GPU Operator chart; evict tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0136","https://github.com/NVIDIA/product-security/tree/main/2025/5599"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-653"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0148","cve":"CVE-2024-0148","aliases":[],"title":"IGX Orin bootloader: Improper access control in bootloader","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"IGX Orin bootloader","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"Improper access control in bootloader -> persistent compromise","attack_vector":"Local attacker with device access","remediation":"Flash bootloader out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0148","https://github.com/NVIDIA/product-security/tree/main/2025/5617"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-447"]},{"id":"CVE-2024-25943","cve":"CVE-2024-25943","aliases":["DSA-2024-099"],"title":"Dell iDRAC9 (IPMI 2.0 over LAN): iDRAC9 generates predictable IPMI 2.0 session IDs, so an attacker can hijack somebody else's live IPMI…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (IPMI 2.0 over LAN)","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"iDRAC9 generates predictable IPMI 2.0 session IDs, so an attacker can hijack somebody else's live IPMI session rather than authenticate as themselves. The stolen session carries whatever privilege the legitimate user had - typically chassis power control, boot device selection, and sensor/SEL access. For a GPU fleet the concrete outcomes are unauthorised power cycling of running training nodes and boot-order manipulation that sets up an attacker-controlled boot on the next restart. Spans 14G, 15G and 16G PowerEdge, which is most of the current GPU-server installed base.","attack_vector":"Anything that can reach UDP 623 on the iDRAC address, plus the ability to observe or race a legitimate IPMI session. That means the OOB management VLAN, and in practice any automation host, monitoring collector, or DCIM system that already speaks IPMI to the fleet.","remediation":"Flash iDRAC9 to 7.00.00.172 (14G) or 7.10.50.00 (15G/16G) or later - out-of-band, per-node, no host reboot and no drain of running jobs. The strong config-only mitigation here is to disable IPMI over LAN entirely: Redfish and racadm cover everything modern tooling needs, and turning IPMI off removes an entire legacy attack surface rather than patching one bug in it. Budget for the tooling migration if any of your automation still speaks ipmitool.","references":["https://www.dell.com/support/kbdoc/en-us/000226503/dsa-2024-099-security-update-for-dell-idrac9-ipmi-session-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2024-25943"],"status":"curated"},{"id":"CVE-2024-43805","cve":"CVE-2024-43805","aliases":[],"title":"JupyterLab: XSS via untrusted notebook content","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterLab","year":"2024","cvss_score":7.6,"severity":"high","kev":false,"impact":"XSS via untrusted notebook content → session compromise","attack_vector":"Customer-supplied notebook opened by another user","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43805"],"status":"curated"},{"id":"CVE-2025-23249","cve":"CVE-2025-23249","aliases":[],"title":"NeMo Framework: RCE via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"RCE via insecure deserialization","attack_vector":"Malicious model/checkpoint","remediation":"Bump NeMo in training images; rebuild; restrict checkpoint provenance","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23249","https://github.com/NVIDIA/product-security/tree/main/2025/5641"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:L/I:H/A:L","cwe":["CWE-502"]},{"id":"CVE-2025-23250","cve":"CVE-2025-23250","aliases":[],"title":"NeMo Framework: Arbitrary file write/read via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Arbitrary file write/read via path traversal","attack_vector":"Malicious model artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23250","https://github.com/NVIDIA/product-security/tree/main/2025/5641"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:L/I:H/A:L","cwe":["CWE-22"]},{"id":"CVE-2025-23251","cve":"CVE-2025-23251","aliases":[],"title":"NeMo Framework: RCE","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NeMo Framework","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"RCE","attack_vector":"Malicious model artifact","remediation":"Bump NeMo; rebuild training images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23251","https://github.com/NVIDIA/product-security/tree/main/2025/5641"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:L/I:H/A:L","cwe":["CWE-94"]},{"id":"CVE-2025-23263","cve":"CVE-2025-23263","aliases":[],"title":"Mellanox OFED: Authentication bypass in the host networking stack","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox OFED","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Authentication bypass in the host networking stack","attack_vector":"Network-adjacent attacker","remediation":"Upgrade MLNX_OFED / DOCA-Host on all nodes; driver reload requires node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23263","https://github.com/NVIDIA/product-security/tree/main/2025/5654"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:L","cwe":["CWE-279"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-23343","cve":"CVE-2025-23343","aliases":[],"title":"NVIDIA NVDebug tool: NVDebug can be induced to write files into restricted components, reaching data tampering and information…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVDebug tool","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"NVDebug can be induced to write files into restricted components, reaching data tampering and information disclosure with a changed scope on a DGX/HGX platform host.","attack_vector":"Adjacent network, low privileges, user interaction, high complexity. Narrow, but the target is a platform-management host.","remediation":"Update NVDebug per bulletin 5696. Cost: trivial - replace the tool bundle. No node drain.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23343","https://github.com/NVIDIA/product-security/tree/main/2025/5696"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:R/S:C/C:H/I:H/A:H","cwe":["CWE-22"]},{"id":"CVE-2025-33203","cve":"CVE-2025-33203","aliases":[],"title":"NVIDIA NeMo Agent Toolkit (Web UI): The chat API endpoint is vulnerable to server-side request forgery, so an attacker makes the toolkit issue…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Agent Toolkit (Web UI)","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"The chat API endpoint is vulnerable to server-side request forgery, so an attacker makes the toolkit issue requests to internal endpoints - cloud metadata services and in-cluster APIs being the obvious targets. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5726 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33203","https://github.com/NVIDIA/product-security/tree/main/2025/5726"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-918"]},{"id":"CVE-2025-3744","cve":"CVE-2025-3744","aliases":[],"title":"HashiCorp Nomad Enterprise: Jobs using the policy-override option bypass mandatory Sentinel policies","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Nomad Enterprise","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Jobs using the policy-override option bypass mandatory Sentinel policies","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade Nomad 1.10.1/1.9.9/1.8.13; audit recently submitted jobs","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3744"],"status":"curated"},{"id":"CVE-2025-4123","cve":"CVE-2025-4123","aliases":[],"title":"Grafana: Client path traversal + open redirect","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"Client path traversal + open redirect -> load an attacker-hosted frontend plugin and run arbitrary JS","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; disable anonymous access","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-4123"],"status":"curated"},{"id":"CVE-2025-47290","cve":"CVE-2025-47290","aliases":[],"title":"containerd: TOCTOU during image unpack: a crafted image can arbitrarily modify the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2025","cvss_score":7.6,"severity":"high","kev":false,"impact":"TOCTOU during image unpack: a crafted image can arbitrarily modify the host filesystem","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain; restrict tenant registries","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-47290"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2011-3997","cve":"CVE-2011-3997","aliases":[],"title":"Opengear console server: Authentication bypass in the console server allowing remote attackers to modify settings and reach every…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Opengear console server","year":"2011","cvss_score":7.5,"severity":"high","kev":false,"impact":"Authentication bypass in the console server allowing remote attackers to modify settings and reach every attached device's serial console — switches, BMCs, storage controllers","attack_vector":"Network / OOB LAN, unauthenticated","remediation":"Console-server firmware upgrade; older units are EOL and the practical fix is replacing the appliance, which means a scheduled loss of out-of-band access to the rack","references":["https://nvd.nist.gov/vuln/detail/CVE-2011-3997"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2013-4786","cve":"CVE-2013-4786","aliases":[],"title":"IPMI 2.0 RAKP (all vendors): Protocol design flaw — RAKP message 2 returns an HMAC over the password hash to any unauthenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IPMI 2.0 RAKP (all vendors)","year":"2013","cvss_score":7.5,"severity":"high","kev":false,"impact":"Protocol design flaw — RAKP message 2 returns an HMAC over the password hash to any unauthenticated requester, enabling offline cracking of every BMC account on the fleet","attack_vector":"Network / IPMI over LAN, unauthenticated","remediation":"Cannot be patched — it is the IPMI 2.0 spec. Only real remediation is disabling IPMI-over-LAN entirely and moving to Redfish with strong per-node unique credentials, which breaks legacy provisioning tooling","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4786"],"status":"curated","fleet":{"ubiquity":"universal - IPMI-over-LAN is enabled on essentially every server BMC unless deliberately disabled","remediation_pain":"unpatchable-mitigate-only - **this is a flaw in the IPMI 2.0 specification itself**, so no firmware fixes it; the only remediation is disabling IPMI-over-LAN fleet-wide or hard-isolating UDP/623, which breaks tooling that depends on it","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"A vulnerable BMC hands out a password-derived HMAC-SHA1 before authentication, so any host that can reach the management network harvests offline-crackable BMC credentials for every node at once - and BMC passwords are typically identical across a fleet built from one golden config."}},{"id":"CVE-2017-5926","cve":"CVE-2017-5926","aliases":["AnC-style MMU side channel"],"title":"AMD processors - page table walk traces in the last-level cache: MULTI-TENANT ISOLATION: The MMU's page table walks during address translation leave traces in the last-level…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - page table walk traces in the last-level cache","year":"2017","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The MMU's page table walks during address translation leave traces in the last-level cache, which is shared across cores on an AMD socket. A side-channel attack on the MMU recovers information about a victim's virtual address layout - and since the LLC is shared, this reaches across cores, not just across SMT threads, so core pinning does not help.","attack_vector":"Local, co-resident on the same socket as the victim. Works across cores because the last-level cache is the shared resource.","remediation":"**Effectively unpatchable in hardware** - shared last-level cache is a design property, not a bug. Mitigation is architectural: do not co-schedule mutually untrusted tenants on the same socket, and where the threat model demands it, allocate whole nodes rather than slices. On a GPU fleet that maps naturally onto whole-node allocation for sensitive customers, which most operators already offer as a premium tier.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5926"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2018-0090","cve":"CVE-2018-0090","aliases":[],"title":"Cisco NX-OS (management interface ACL): The ACL you put on the management interface is not enforced, so traffic you believe is being dropped reaches…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (management interface ACL)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"The ACL you put on the management interface is not enforced, so traffic you believe is being dropped reaches the switch's control plane anyway. Every 'we restricted the mgmt interface to the jump host' assumption becomes false. It does not by itself grant access, but it silently removes the compensating control that most operators rely on for every other switch CVE in this list.","attack_vector":"Unauthenticated, remote — any host with IP reachability to mgmt0, even one the ACL was supposed to block.","remediation":"NX-OS upgrade plus reload. Do not treat management-interface ACLs as a substitute for real network segmentation — put the management interface on a physically or VRF-separated OOB network, which is a config/topology change and the durable fix.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0090"],"status":"curated"},{"id":"CVE-2018-1211","cve":"CVE-2018-1211","aliases":[],"title":"Dell iDRAC7 / iDRAC8 (web server URI parser): Directory traversal in the BMC's own HTTP front end lets an attacker with no credentials read files off the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC7 / iDRAC8 (web server URI parser)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal in the BMC's own HTTP front end lets an attacker with no credentials read files off the iDRAC filesystem. In practice that is reconnaissance that turns into access: configuration, session material, and stored secrets that let the attacker come back as an authenticated user. Relevant to fleets that still run 12G/13G PowerEdge as head nodes, storage, or staging boxes alongside the GPU racks - the older tier tends to be the one nobody re-flashed.","attack_vector":"Anything routable to the iDRAC web port on the out-of-band management VLAN, unauthenticated. No host access and no valid account required.","remediation":"Flash iDRAC7/iDRAC8 to 2.52.52.52 or later - out-of-band, per-node, no host reboot. On a fleet this old the real cost is inventory: finding which nodes are still on pre-2.52 firmware. Config-only interim: ACL the iDRAC web port to the jump hosts only. Note the original Dell TechCenter advisory URL is dead; NVD carries the authoritative version data.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-1211","https://www.dell.com/support/kbdoc/en-us/000131409/dsa-2018-004-idrac-vulnerabilities"],"status":"curated"},{"id":"CVE-2018-12922","cve":"CVE-2018-12922","aliases":[],"title":"Emerson/Vertiv Liebert IntelliSlot Web Card (config/configUser.htm, config/configTelnet.htm): The IntelliSlot card is the network brain bolted into Liebert CRAC/CRAH units, condensers and thermal…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Emerson/Vertiv Liebert IntelliSlot Web Card (config/configUser.htm, config/configTelnet.htm)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"The IntelliSlot card is the network brain bolted into Liebert CRAC/CRAH units, condensers and thermal management gear. A remote attacker can reconfigure its access control and telnet settings by hitting the config pages directly, which means they can create their own administrative access and keep it. From there the card exposes the unit's operating parameters - fan command, compressor enable, setpoints, alarm relays. Turning off or de-rating the CRAHs serving a GPU hall is a fleet-wide availability kill: the racks do not have hours of ride-through, they have minutes. Equally damaging and much quieter is suppressing the alarm path, so the operations team's first indication of a thermal event is GPUs dropping off the fabric rather than a temperature alarm.","attack_vector":"Remote HTTP to the card, no authentication. These cards live on the facility monitoring VLAN alongside the PDU network cards and environmental sensors. Many operators inherited that VLAN from the building and treat it as trusted. Note that IntelliSlot cards are frequently also reachable from the DCIM/monitoring server, so compromise of a monitoring host is a direct path in.","remediation":"Replace the card. The IntelliSlot Web Card generation covered here is end-of-life; Vertiv's answer is migration to a current Unity/IS-UNITY card, which is a hardware swap per cooling unit and needs the unit taken off network (not necessarily off cooling). Until then: isolate the card VLAN, block telnet outright at the switch, and put an ACL so only the DCIM collector can reach the card's HTTP port. If you lease, this is landlord equipment - ask for the card model and firmware inventory in writing and treat 'IntelliSlot Web Card' in the answer as a finding.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12922","https://www.seebug.org/vuldb/ssvid-97372"],"status":"curated"},{"id":"CVE-2018-15664","cve":"CVE-2018-15664","aliases":[],"title":"Docker / moby: `docker cp` symlink-exchange TOCTOU gives arbitrary host read/write as root","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"`docker cp` symlink-exchange TOCTOU gives arbitrary host read/write as root","attack_vector":"Any tenant workload on a node whose operator uses docker cp","remediation":"Upgrade Docker Engine; restart daemon (drains containers unless live-restore is on)","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15664"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2018-5254","cve":"CVE-2018-5254","aliases":[],"title":"Arista EOS (BGP UPDATE): Malformed path attribute in a BGP UPDATE from a peer causes denial of service","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (BGP UPDATE)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"Malformed path attribute in a BGP UPDATE from a peer causes denial of service","attack_vector":"Network, from a BGP peer","remediation":"EOS upgrade; relevant where the fabric peers with a transit provider or a customer-controlled router","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-5254"],"status":"curated"},{"id":"CVE-2018-7185","cve":"CVE-2018-7185","aliases":["Zero Origin timestamp / 'Zero-o'"],"title":"ntpd (protocol engine, zero-origin timestamp): Continually sending packets with a zero-origin timestamp lets a remote attacker disrupt an ntpd peer…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ntpd (protocol engine, zero-origin timestamp)","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"Continually sending packets with a zero-origin timestamp lets a remote attacker disrupt an ntpd peer association. Cheap, stateless, and it needs nothing but the ability to send UDP to port 123 — which is open on far more cluster nodes than operators realise, because NTP is usually configured once at image-build time and never reviewed.","attack_vector":"Remote, unauthenticated — UDP packets to the NTP port.","remediation":"Upgrade ntp to 4.2.8p11 or later and restart. Also firewall UDP/123 so only your internal time servers can reach cluster nodes — a host or fabric ACL change, applied live, that removes most of the NTP attack surface at once.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7185"],"status":"curated"},{"id":"CVE-2019-11253","cve":"CVE-2019-11253","aliases":[],"title":"Kubernetes (kube-apiserver): \"Billion laughs\": malicious YAML/JSON payload consumes all apiserver memory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"\"Billion laughs\": malicious YAML/JSON payload consumes all apiserver memory; control-plane DoS","attack_vector":"Any authorized cluster user, and unauthenticated if anonymous auth is on","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11253"],"status":"curated"},{"id":"CVE-2019-13509","cve":"CVE-2019-13509","aliases":[],"title":"Docker / moby: Docker Engine in debug mode writes secrets into the debug log","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Docker Engine in debug mode writes secrets into the debug log","attack_vector":"Anyone with node or log-pipeline read access","remediation":"Turn off daemon debug mode; rotate leaked secrets; scrub log store","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-13509"],"status":"curated"},{"id":"CVE-2019-16884","cve":"CVE-2019-16884","aliases":[],"title":"runc: AppArmor restriction bypass via mount-target check flaw","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"AppArmor restriction bypass via mount-target check flaw; container can mount over /proc","attack_vector":"Malicious image","remediation":"Replace runc binary; drain node to restart existing containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16884"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2019-16919","cve":"CVE-2019-16919","aliases":[],"title":"Harbor: Broken access control allows creating robot accounts with push/pull rights to projects the user does not own","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Broken access control allows creating robot accounts with push/pull rights to projects the user does not own; cross-tenant image poisoning","attack_vector":"Project administrator of any project","remediation":"Upgrade Harbor; audit robot accounts","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16919"],"status":"curated"},{"id":"CVE-2019-18948","cve":"CVE-2019-18948","aliases":[],"title":"Arista EOS (VxLAN agent): Malformed ARP packets crash the VxLAN software forwarding agent — a tenant VM can take down overlay forwarding","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (VxLAN agent)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Malformed ARP packets crash the VxLAN software forwarding agent — a tenant VM can take down overlay forwarding","attack_vector":"Network, from inside a tenant VLAN","remediation":"EOS upgrade with failover; notable because the trigger comes from tenant traffic, not the management plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-18948"],"status":"curated"},{"id":"CVE-2019-20423","cve":"CVE-2019-20423","aliases":[],"title":"Lustre ptlrpc / mdt modules (client-driven server panic family): FABRIC DOS: the head of a family of ten Lustre defects (CVE-2019-20423 through CVE-2019-20432) that all share…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Lustre ptlrpc / mdt modules (client-driven server panic family)","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: the head of a family of ten Lustre defects (CVE-2019-20423 through CVE-2019-20432) that all share one shape — the server does not validate fields in packets sent by a client, so any client can panic the metadata or object storage server with a malformed RPC. Ten separate ways for one tenant's node to take down the filesystem that every other tenant's training job is reading from. In a shared-storage AI cluster this is the cheapest available cross-tenant denial of service: one machine, one packet, everyone's jobs stall.","attack_vector":"Any mounted Lustre client. No credentials, no escalation — the client is inherently trusted by the protocol.","remediation":"Upgrade Lustre servers to 2.12.3 or later and restart the MDS/OSS nodes; failover pairs limit but do not eliminate the I/O pause. There is no config workaround for the parsing itself. What you can do immediately is network-level: restrict which hosts can reach LNet, and stop treating 'the tenant's compute node' as a trusted peer of your storage servers.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-20423","https://nvd.nist.gov/vuln/detail/CVE-2019-20425","https://nvd.nist.gov/vuln/detail/CVE-2019-20432"],"status":"curated"},{"id":"CVE-2019-6193","cve":"CVE-2019-6193","aliases":["LEN-29477"],"title":"Lenovo XClarity Administrator (LXCA) - unauthenticated config file access: Unauthenticated access to LXCA configuration files, which contain usernames, license keys and IP addresses.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Administrator (LXCA) - unauthenticated config file access","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated access to LXCA configuration files, which contain usernames, license keys and IP addresses. LXCA is the fleet management console that talks to every XCC it manages, so an unauthenticated read of its configuration hands an attacker a map of the whole management estate - which addresses to hit, which account names to try - without touching a single node. Combined with any of the XCC authorisation bypasses of the same era, that map is the difference between a blind scan and a targeted walk through the fleet. Affects LXCA before 2.6.6.","attack_vector":"Anything routable to the LXCA appliance, unauthenticated. No credentials, no host access, no user interaction.","remediation":"Upgrade LXCA to 2.6.6 or later. Note the upgrade path: you must be on 2.6.0 before you can install the 2.6.6 fix bundle, so this is a two-step appliance upgrade, not a single jump. It is still a single appliance rather than a per-node campaign - no node reboots, no job drain, only LXCA's own downtime. Rotate anything the exposed configuration named, and put LXCA on a restricted segment rather than the general management VLAN.","references":["https://support.lenovo.com/us/en/product_security/LEN-29477","https://nvd.nist.gov/vuln/detail/CVE-2019-6193"],"status":"curated"},{"id":"CVE-2019-9946","cve":"CVE-2019-9946","aliases":[],"title":"CNI portmap plugin: portmap inserts rules ahead of the KUBE-SERVICES chain, so hostPort traffic bypasses NetworkPolicy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CNI portmap plugin","year":"2019","cvss_score":7.5,"severity":"high","kev":false,"impact":"portmap inserts rules ahead of the KUBE-SERVICES chain, so hostPort traffic bypasses NetworkPolicy","attack_vector":"Any pod on the cluster network","remediation":"Upgrade the CNI plugins package on every node; block hostPort for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9946"],"status":"curated"},{"id":"CVE-2020-11487","cve":"CVE-2020-11487","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): A hard-coded RSA-1024 key with weak ciphers in the BMC firmware means the encryption protecting BMC sessions…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"A hard-coded RSA-1024 key with weak ciphers in the BMC firmware means the encryption protecting BMC sessions and data is decryptable by anyone holding the firmware image - so passively captured management traffic can be read. Notably this one lists all DGX A100 BMC firmware versions as affected, not just older DGX-1/DGX-2 builds.","attack_vector":"Anyone who can capture traffic on the management network, plus anyone who can download the firmware - which is everyone.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11487"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11489","cve":"CVE-2020-11489","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): Default SNMP community strings on the DGX BMC. Anyone who can reach the BMC over SNMP enumerates hardware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Default SNMP community strings on the DGX BMC. Anyone who can reach the BMC over SNMP enumerates hardware inventory, sensor data and configuration without authenticating - cheap reconnaissance that maps your fleet before a real attack. DGX-1 before 3.38.30, DGX-2 before 1.06.06.","attack_vector":"Anyone with UDP reach to the BMC's SNMP port on the management network.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks. Also change or disable SNMP community strings explicitly after flashing - a firmware update does not necessarily reset a configured default.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11489"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11615","cve":"CVE-2020-11615","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): Hard-coded RC4 key in the DGX BMC firmware. RC4 is broken to begin with and the key is public, so anything…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Hard-coded RC4 key in the DGX BMC firmware. RC4 is broken to begin with and the key is public, so anything the BMC protects with it should be treated as cleartext. DGX-1 before BMC 3.38.30.","attack_vector":"Anyone who can capture BMC traffic or read BMC-stored data, plus anyone who downloads the firmware image.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11615"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11616","cve":"CVE-2020-11616","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): The PRNG used by the IPMI implementation in the BMC's JSOL package is not cryptographically strong, so…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"The PRNG used by the IPMI implementation in the BMC's JSOL package is not cryptographically strong, so session identifiers and other 'random' values are predictable. Combined with the rest of this bulletin it removes the guesswork from hijacking BMC/IPMI sessions. DGX-1 before BMC 3.38.30.","attack_vector":"Anyone with network reach to the BMC's IPMI service.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11616"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-11868","cve":"CVE-2020-11868","aliases":[],"title":"ntpd (NTP.org reference implementation): An off-path attacker can block a node's unauthenticated time synchronization by spoofing the source IP in a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ntpd (NTP.org reference implementation)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"An off-path attacker can block a node's unauthenticated time synchronization by spoofing the source IP in a server-mode packet. Clock drift is an underrated cluster failure mode: it breaks TLS validity windows, Kerberos, distributed tracing correlation, checkpoint ordering, and any scheduler that reasons about deadlines. An attacker who can freeze your clocks without touching your data plane has a quiet, hard-to-attribute lever.","attack_vector":"Off-path attacker able to spoof source addresses toward the NTP client. Does not require being on the path between client and server.","remediation":"Upgrade ntp to 4.2.8p14 or later (or migrate to chrony, which most modern distributions default to) and restart the service. Package upgrade, no reboot. The structural fix is authenticated time — NTS or symmetric-key NTP — plus internal stratum-1 sources rather than public pools; that is a config and topology change and it is what actually removes this class.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11868","https://nvd.nist.gov/vuln/detail/CVE-2020-13817"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2020-12965","cve":"CVE-2020-12965","aliases":["Transient Execution of Non-canonical Accesses"],"title":"AMD processors - transient non-canonical loads and stores using lower 48 address bits: MULTI-TENANT ISOLATION: Combined with specific software sequences, AMD CPUs transiently execute non-canonical…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - transient non-canonical loads and stores using lower 48 address bits","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Combined with specific software sequences, AMD CPUs transiently execute non-canonical loads and stores using only the lower 48 address bits - so an access that should fault instead speculatively reads a truncated address. That truncation can land inside another security domain's memory, and the result is observable through the usual cache channels. Data leakage across the boundaries the address canonicality check was supposed to enforce.","attack_vector":"Local, needs the victim to contain a specific software sequence, so exploitability depends on what is running - but on a shared node you do not control what your tenants run.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Kernel-side mitigations exist for the known sequences; take the distro kernel update as well as the firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12965","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-15383","cve":"CVE-2020-15383","aliases":[],"title":"Brocade Fabric OS (config and secnotify processes): Running a routine security scan against the SAN switch crashes the config and secnotify processes. Worth…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS (config and secnotify processes)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Running a routine security scan against the SAN switch crashes the config and secnotify processes. Worth carrying because it inverts the usual advice: the compliance activity you are required to perform is itself the outage. Operators who scan their storage fabric on a schedule have been taking unexplained SAN switch faults from their own tooling.","attack_vector":"Any security scanner reaching the switch's management services — no attacker required.","remediation":"Fabric OS upgrade to v9.0.0 / v8.2.2d / v8.2.1e or later, firmware install plus reboot. Until then, exclude FOS management addresses from automated vulnerability scans or scan them only in a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15383"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-25632","cve":"CVE-2020-25632","aliases":[],"title":"GRUB2 (rmmod command): Use-after-free in the rmmod command. Unloading a module whose dependencies are still live leaves dangling…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (rmmod command)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Use-after-free in the rmmod command. Unloading a module whose dependencies are still live leaves dangling pointers GRUB will later call through, which is a clean primitive for arbitrary pre-boot execution and another Secure Boot bypass.","attack_vector":"Local, via GRUB command line or a controlled grub.cfg. On a bare-metal fleet, any tenant who had console or root on the node.","remediation":"grub2 package update + reboot per node. If you leave the GRUB command line unlocked on your image, set a GRUB password as a stopgap - it does not fix the bug but it removes the easiest path to it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-25632","https://access.redhat.com/security/cve/CVE-2020-25632"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-27174","cve":"CVE-2020-27174","aliases":[],"title":"Firecracker: Unbounded serial console buffer growth leaks host memory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unbounded serial console buffer growth leaks host memory; host exhaustion from inside a guest","attack_vector":"Any tenant guest VM","remediation":"Upgrade Firecracker","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27174"],"status":"curated"},{"id":"CVE-2020-27749","cve":"CVE-2020-27749","aliases":[],"title":"GRUB2 (grub_parser_split_cmdline): Stack buffer overflow from variable expansion in the GRUB command line. Attacker gets code execution before…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (grub_parser_split_cmdline)","year":"2020","cvss_score":7.5,"severity":"high","kev":false,"impact":"Stack buffer overflow from variable expansion in the GRUB command line. Attacker gets code execution before the kernel and can therefore load an unsigned kernel or plant a bootkit that no in-OS EDR will see.","attack_vector":"Anyone who can type at the GRUB prompt or supply grub.cfg - local console, serial console server, or BMC KVM.","remediation":"grub2 package update + reboot per node. Set a GRUB password and lock the serial/BMC console as a partial mitigation in the meantime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27749","https://access.redhat.com/security/cve/CVE-2020-27749"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-23201","cve":"CVE-2021-23201","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: a privileged user can craft microcode that the GPU's internal microcontroller accepts…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: a privileged user can craft microcode that the GPU's internal microcontroller accepts as valid. That means loading attacker-controlled firmware onto the GPU, below the driver and below the OS - persistent, invisible to host tooling, and NVIDIA explicitly notes the scope may extend to other components. On a bare-metal GPU rental this is the tenant-persistence scenario: tenant A leaves code on the card that outlives the reprovision and is there when tenant B arrives. Verify with the vendor whether a full VBIOS and firmware reflash between tenants is sufficient.","attack_vector":"A user with elevated privileges on the host. On rented bare metal that is the tenant; on a managed cluster it is anyone who got root on a node.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained. Because this is a firmware-persistence risk, also add a firmware/VBIOS attestation or reflash step to your bare-metal tenant handoff - patching alone does not tell you whether a previous tenant already loaded something.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23201"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-23217","cve":"CVE-2021-23217","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: a privileged user can time a DMA write from the GPU's internal microcontroller to…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: a privileged user can time a DMA write from the GPU's internal microcontroller to land inside a specific window and corrupt code execution, with impact NVIDIA says may extend to other components. A DMA engine writing outside its lane is the mechanism by which a GPU compromises its host.","attack_vector":"A user with elevated privileges on the GPU host, able to time operations precisely.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23217"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26406","cve":"CVE-2021-26406","aliases":[],"title":"AMD SEV / SEV-ES - Owner's Certificate Authority (OCA) certificate parsing: Insufficient validation when parsing OCA certificates in the SEV and SEV-ES user application crashes the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES - Owner's Certificate Authority (OCA) certificate parsing","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Insufficient validation when parsing OCA certificates in the SEV and SEV-ES user application crashes the host. OCA certificates come from whoever owns the platform's SEV identity, so this is a malformed-input crash on the certificate path that underpins SEV ownership and attestation - a guest or tooling that supplies a bad certificate takes the machine down.","attack_vector":"Reachable by whatever supplies OCA certificates to the SEV stack, which in a managed confidential-computing service is the control plane or the tenant-facing provisioning path.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26406","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26923","cve":"CVE-2021-26923","aliases":[],"title":"Argo CD: /api/version leaks internal system information without authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"/api/version leaks internal system information without authentication","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26923"],"status":"curated"},{"id":"CVE-2021-28361","cve":"CVE-2021-28361","aliases":["CVE-2019-9547"],"title":"SPDK iSCSI target (before 20.01.01) and SPDK vhost target (before 19.01): A zero-length PDU sent where data is expected crashes the SPDK iSCSI target on a NULL pointer dereference.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SPDK iSCSI target (before 20.01.01) and SPDK vhost target (before 19.01)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"A zero-length PDU sent where data is expected crashes the SPDK iSCSI target on a NULL pointer dereference. The companion vhost defect lets a guest VM build a circular descriptor chain that partially wedges the SPDK vhost target. Both are single-packet or single-guest denial of service against a userspace process that is serving block storage to many tenants at once - the iSCSI one needs no authentication and no valid target, and the vhost one is reachable from inside any guest whose virtio-blk/virtio-scsi device SPDK is backing. On a node running SPDK as the storage datapath for a rack of VMs, one guest kills storage for all of them.","attack_vector":"iSCSI variant: any host that can connect to the SPDK iSCSI target port (3260) on the storage network, unauthenticated. vhost variant: a malicious or compromised guest VM whose virtio block device is served by SPDK vhost - i.e. a paying tenant.","remediation":"Upgrade SPDK past 20.01.01 (iSCSI) and 19.01 (vhost) and restart the target process, which disconnects all sessions on that node. Anyone still on an SPDK that old is likely running a vendored fork inside a storage appliance image, so the practical action is to identify the SPDK version compiled into your storage service, not to check a package manifest. For the vhost issue there is no configuration mitigation - the attacker is inside the guest by definition, which is the whole point of the exposure.","references":["https://github.com/spdk/spdk/commit/eca42c66092b9031711afe215fbc1891ee55f143","https://nvd.nist.gov/vuln/detail/CVE-2021-28361","https://nvd.nist.gov/vuln/detail/CVE-2019-9547"],"status":"curated"},{"id":"CVE-2021-28505","cve":"CVE-2021-28505","aliases":[],"title":"Arista EOS (VXLAN match rule in IPv4 ACL): TENANT ISOLATION: if an IPv4 access list contains a VXLAN match rule, that rule and every rule after it in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (VXLAN match rule in IPv4 ACL)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"TENANT ISOLATION: if an IPv4 access list contains a VXLAN match rule, that rule and every rule after it in the list ignore the IP protocol you specified. Your ACL silently permits or denies far more than you wrote. Anyone using ACLs to keep tenant overlays apart, or to fence off a storage VLAN, is enforcing something other than what is in the config — and `show access-list` will not tell you.","attack_vector":"Any traffic subject to the affected ACL. No attacker capability needed beyond being on a path the ACL was supposed to control.","remediation":"EOS upgrade plus reload. Immediate mitigation: reorder access lists so VXLAN match rules come last, or split them into a separate list — a live config change that restores correct enforcement of the remaining rules. Audit every ACL in the fabric for VXLAN match rules before assuming you are unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28505"],"status":"curated"},{"id":"CVE-2021-28508","cve":"CVE-2021-28508","aliases":[],"title":"Arista EOS (TerminAttr / IPsec): TerminAttr leaks IPsec sensitive material in plaintext to authorized users — fabric encryption keys exposed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (TerminAttr / IPsec)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"TerminAttr leaks IPsec sensitive material in plaintext to authorized users — fabric encryption keys exposed through the telemetry agent","attack_vector":"Local/network","remediation":"EOS + TerminAttr upgrade plus rotation of any IPsec keys that were live on the affected switches","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28508"],"status":"curated"},{"id":"CVE-2021-3695","cve":"CVE-2021-3695","aliases":[],"title":"GRUB2 (PNG reader): A crafted PNG in the boot splash path causes an out-of-bounds write in GRUB. Boot logos and themes are…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (PNG reader)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"A crafted PNG in the boot splash path causes an out-of-bounds write in GRUB. Boot logos and themes are attacker-writable data on almost every image, so this is a low-effort way to turn cosmetic files into pre-boot code execution.","attack_vector":"Anyone who can replace a theme/splash image on the boot partition - local root, previous tenant, or a tampered golden image.","remediation":"grub2 package update + reboot. Strip custom boot themes from your golden image if you do not need them; it removes the attack surface entirely at zero operational cost.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3695","https://access.redhat.com/security/cve/CVE-2021-3695"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-3697","cve":"CVE-2021-3697","aliases":[],"title":"GRUB2 (JPEG reader): Crafted JPEG in the boot path drives a heap out-of-bounds write in GRUB. This is the GRUB-side sibling of the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (JPEG reader)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted JPEG in the boot path drives a heap out-of-bounds write in GRUB. This is the GRUB-side sibling of the firmware image-parser problem that LogoFAIL exploited a year later - same idea, different layer.","attack_vector":"Attacker-writable splash/theme file on the boot partition.","remediation":"grub2 package update + reboot. Removing boot theme images from the image is a real mitigation here.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3697","https://access.redhat.com/security/cve/CVE-2021-3697"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-3748","cve":"CVE-2021-3748","aliases":[],"title":"QEMU (virtio-net): Heap use-after-free in virtio_net_receive_rcu - guest-to-host code execution in the QEMU process","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (virtio-net)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Heap use-after-free in virtio_net_receive_rcu - guest-to-host code execution in the QEMU process","attack_vector":"Tenant VM guest","remediation":"QEMU package update + restart each VM's qemu process. Live-migrate to patched hosts to avoid tenant downtime; GPU-passthrough VMs cannot live-migrate, so this becomes a scheduled drain","references":["https://access.redhat.com/security/cve/CVE-2021-3748"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-38960","cve":"CVE-2021-38960","aliases":["IBM X-Force 212047"],"title":"IBM OpenBMC OP920 / OP930 / OP940: An unauthenticated caller retrieves sensitive information from the BMC. Same operational shape as the 2024…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC OP920 / OP930 / OP940","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated caller retrieves sensitive information from the BMC. Same operational shape as the 2024 bmcweb URI disclosure three years earlier and on the same product line, which is the useful signal: pre-auth information exposure on this BMC stack is recurring, not a one-off. For a fleet operator the leaked material is inventory and configuration detail that lets an attacker pick which nodes to attack and with what.","attack_vector":"Unauthenticated network access to the BMC's management interface.","remediation":"Fixed in later OP920/OP930/OP940 firmware - a per-node system firmware update requiring a maintenance window. Because this class keeps recurring on the same stack, the durable control is the network boundary: BMCs on an isolated management VLAN reachable only from bastion hosts, so pre-auth disclosure bugs have no audience. That is config-only and it covers the next one too.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-38960","https://www.ibm.com/support/pages/node/6529322","https://exchange.xforce.ibmcloud.com/vulnerabilities/212047"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-39295","cve":"CVE-2021-39295","aliases":["GHSA-gg9x-v835-m48q"],"title":"OpenBMC phosphor-net-ipmid (IPMI LAN+): Sibling finding to the authentication bypass, from the same Google report. Crafted IPMI messages take the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC phosphor-net-ipmid (IPMI LAN+)","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Sibling finding to the authentication bypass, from the same Google report. Crafted IPMI messages take the BMC's IPMI daemon down without any credentials. Losing IPMI on its own is survivable if you run Redfish, but the operational shape is bad: an unauthenticated packet source on the management VLAN can knock out out-of-band management across every ASPEED node simultaneously, which is precisely when you would want it - during an incident, or to blind an operator while something else happens on the hosts.","attack_vector":"Unauthenticated, network, UDP 623 on the BMC. Same reachability precondition as the authentication bypass.","remediation":"Same fix and same delivery cost as the authentication bypass: post-2.9 OpenBMC via a per-node out-of-band BMC firmware flash. Config-only mitigation is the same and is the right first move: turn off IPMI over LAN and run Redfish, or ACL UDP 623 to your management jump hosts. If you are already flashing for CVE-2021-39296 you get this one in the same image.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-39295","https://github.com/google/security-research/security/advisories/GHSA-gg9x-v835-m48q","https://github.com/openbmc/openbmc/issues/3811"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-43798","cve":"CVE-2021-43798","aliases":[],"title":"Grafana: Unauthenticated directory traversal via /public/plugins/<id>/","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2021","cvss_score":7.5,"severity":"high","kev":true,"impact":"[KEV] Unauthenticated directory traversal via /public/plugins/<id>/ -> read local files incl. grafana.db","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rotate every datasource credential stored in grafana.db","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43798"],"status":"curated"},{"id":"CVE-2021-45454","cve":"CVE-2021-45454","aliases":["AMP-SB-0003","PLATYPUS on Ampere","power telemetry side channel"],"title":"Ampere Altra before SRP 1.08b and Altra Max before SRP 2.05 - power telemetry exposed through the Linux HWmon interface: Unprivileged readers get fine-grained CPU power telemetry, which is a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra before SRP 1.08b and Altra Max before SRP 2.05 - power telemetry exposed through the Linux HWmon interface","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unprivileged readers get fine-grained CPU power telemetry, which is a data-dependent side channel: power correlates with the operands being processed, so it leaks key material and other secrets from workloads sharing the socket. On a multi-tenant Arm node this is a cross-tenant leak that needs no memory access at all - just a file read in sysfs. It is also a nuisance for anyone selling confidential inference, because the power trace of a model serving run is itself informative about the workload.","attack_vector":"Any unprivileged local user or container on an Altra / Altra Max host with the HWmon power sensors exposed. Containers that inherit the host sysfs make this trivially available to tenants.","remediation":"Update to Altra SRP 1.08b / Altra Max SRP 2.05 or later, which restricts the telemetry. Flash + reboot + drain. Cheaper interim control that works today: restrict access to the HWmon power sensors (root-only permissions, do not bind-mount host /sys into tenant containers, drop the sensor nodes from the container's device allowlist). Losing per-core power telemetry costs you some capacity-planning visibility - decide whether your scheduler actually consumes it before turning it off fleet-wide.","references":["https://amperecomputing.com/products/security-bulletins/platypus.html","https://nvd.nist.gov/vuln/detail/CVE-2021-45454"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-1708","cve":"CVE-2022-1708","aliases":[],"title":"CRI-O: Unbounded ExecSync output exhausts node memory or disk","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unbounded ExecSync output exhausts node memory or disk; node DoS","attack_vector":"Anyone with kube API exec access","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-1708"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-23635","cve":"CVE-2022-23635","aliases":[],"title":"Istio: Crafted message crashes istiod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted message crashes istiod; mesh-wide config distribution stops","attack_vector":"Any pod that can reach istiod, i.e. any tenant pod","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23635"],"status":"curated"},{"id":"CVE-2022-23648","cve":"CVE-2022-23648","aliases":[],"title":"containerd: Crafted image config allows arbitrary host file read by containers launched via the CRI plugin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted image config allows arbitrary host file read by containers launched via the CRI plugin","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23648"],"status":"curated","fleet":{"ubiquity":"Universal - containerd is the runtime under nearly every Kubernetes-based GPU cloud; affects <1.6.1 / 1.5.10 / 1.4.12","remediation_pain":"`daemon-restart` - containerd upgrade; with `--restart` semantics containers may survive, but the fleet-wide rollout still means touching every node","pain_class":"daemon-restart","why_fleet_wide":"A specially crafted *image config* - i.e. something a customer supplies - mounts read-only copies of arbitrary host files into the container, bypassing Pod Security Policy; any tenant who can push an image reads host secrets on every node they land on"}},{"id":"CVE-2022-23818","cve":"CVE-2022-23818","aliases":[],"title":"AMD SEV-SNP - VM_HSAVE_PA MSR validation: MULTI-TENANT ISOLATION: Insufficient validation of the VM_HSAVE_PA model-specific register lets a malicious…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - VM_HSAVE_PA MSR validation","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient validation of the VM_HSAVE_PA model-specific register lets a malicious hypervisor point the host-state save area somewhere it should not be, breaking SEV-SNP guest memory integrity. The RMP is supposed to make it impossible for the host to write guest pages; this is a way around that using an MSR the host legitimately controls.","attack_vector":"Requires host/hypervisor privilege - the SEV-SNP adversary model exactly. No guest cooperation needed.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23818","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-25882","cve":"CVE-2022-25882","aliases":[],"title":"ONNX: Directory traversal via `external_data` field in the tensor proto","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal via `external_data` field in the tensor proto","attack_vector":"Customer-supplied ONNX model file","remediation":"Upgrade onnx >= 1.13.0. Every ONNX ingest path must resolve external-data paths against a jail","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-25882"],"status":"curated"},{"id":"CVE-2022-26353","cve":"CVE-2022-26353","aliases":[],"title":"QEMU (virtio-net): Map leaking on error during receive - guest-triggered host memory exhaustion / DoS","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (virtio-net)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Map leaking on error during receive - guest-triggered host memory exhaustion / DoS","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart or live-migration","references":["https://access.redhat.com/security/cve/CVE-2022-26353"],"status":"curated"},{"id":"CVE-2022-27649","cve":"CVE-2022-27649","aliases":[],"title":"Podman: Containers started with non-empty default inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Containers started with non-empty default inheritable capabilities","attack_vector":"Any tenant workload","remediation":"Upgrade Podman; restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27649"],"status":"curated"},{"id":"CVE-2022-2809","cve":"CVE-2022-2809","aliases":["GHSA-g3qc-375m-h66j"],"title":"OpenBMC bmcweb multipart_parser (Redfish / web UI HTTP front end): bmcweb is the single process behind Redfish, the web UI, serial-over-LAN and the KVM websocket - kill it and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb multipart_parser (Redfish / web UI HTTP front end)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"bmcweb is the single process behind Redfish, the web UI, serial-over-LAN and the KVM websocket - kill it and you lose every out-of-band control path at once. A multipart form body containing a long header line with no colon writes one byte off the end of a heap buffer, and the write can be repeated in a loop. The reported outcome is denial of service, but the advisory classifies it as both heap and stack out-of-bounds writes reachable without authentication, which is a materially worse posture than the DoS framing suggests. Fleet impact: an unauthenticated source on the management VLAN can flatten Redfish on every ASPEED node.","attack_vector":"Unauthenticated HTTP(S) to bmcweb on the BMC's management interface. Multipart upload endpoints are reachable pre-auth in affected versions, so no credentials and no host access are needed.","remediation":"Fixed in bmcweb 2.13 / OpenBMC 2.13 (Gerrit 56796 and 56868). Delivery is a BMC firmware flash - per node, out-of-band, and gated on when your ODM last rebased bmcweb, which for many server vendors is a long time. Check the bmcweb version your image reports before assuming you are covered. Config-only mitigation: restrict which hosts can reach the BMC's HTTPS port to your management jump boxes, which is the same control that limits half of this cluster.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2809","https://github.com/openbmc/bmcweb/security/advisories/GHSA-g3qc-375m-h66j"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-2827","cve":"CVE-2022-2827","aliases":[],"title":"AMI MegaRAC: User enumeration — lets an attacker map valid BMC accounts before credential attack","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"User enumeration — lets an attacker map valid BMC accounts before credential attack","attack_vector":"Network","remediation":"BMC firmware update; low individual severity but it is the reconnaissance step for the rest of the MegaRAC chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2827"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-29179","cve":"CVE-2022-29179","aliases":[],"title":"Cilium: After a container escape, an attacker can install eBPF programs and take over the node dataplane","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"After a container escape, an attacker can install eBPF programs and take over the node dataplane","attack_vector":"An attacker who already escaped a container","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29179"],"status":"curated"},{"id":"CVE-2022-29225","cve":"CVE-2022-29225","aliases":[],"title":"Envoy: Decompressor accumulates unbounded data","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Decompressor accumulates unbounded data; memory exhaustion of the proxy","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29225"],"status":"curated"},{"id":"CVE-2022-30283","cve":"CVE-2022-30283","aliases":["INSYDE-SA-2022063"],"title":"Insyde InsydeH2O (UsbCoreDxe USB working buffer, DMA TOCTOU): UsbCoreDxe builds its USB transaction working buffer outside SMRAM while the code consuming it runs inside…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbCoreDxe USB working buffer, DMA TOCTOU)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"UsbCoreDxe builds its USB transaction working buffer outside SMRAM while the code consuming it runs inside SMM. The driver tries to sanitise pointers against a list of known-good buffer locations, but a pointer that misses the list is used anyway - so DMA tampering mid-transaction corrupts SMRAM and escalates to ring -2. A more subtle failure than the rest of the batch: the validation exists, it just does not fail closed.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in the kernel releases named in the advisory (Insyde does not enumerate per-kernel versions for this one). Disabling USB legacy support on headless nodes reduces how often the vulnerable transaction path runs. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30283","https://www.insyde.com/security-pledge/SA-2022063"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31600","cve":"CVE-2022-31600","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An integer overflow in SmmCore, chainable from another bug, reaches SMM code execution and secure-boot-level…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"An integer overflow in SmmCore, chainable from another bug, reaches SMM code execution and secure-boot-level compromise. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31600","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-3409","cve":"CVE-2022-3409","aliases":["GHSA-g3qc-375m-h66j"],"title":"OpenBMC bmcweb multipart_parser (second variant found during the CVE-2022-2809 fix): The second bug the fuzzer found while the first one was being patched - same parser, same unauthenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb multipart_parser (second variant found during the CVE-2022-2809 fix)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"The second bug the fuzzer found while the first one was being patched - same parser, same unauthenticated reachability, same result of taking down Redfish, KVM and SoL together. Its real value to an operator is as evidence about the code: bmcweb's multipart handling was not hardened, it was patched twice under fuzzing pressure, and the 2026 disclosures below show the same pattern repeating in the HTTP/2 and Expect-header paths. Treat bmcweb version currency as a standing fleet metric rather than a per-CVE chase.","attack_vector":"Unauthenticated HTTP(S) request to bmcweb on the management interface. No credentials, no host access.","remediation":"Same patch train as CVE-2022-2809 - bmcweb 2.13 and later, arriving as a BMC firmware flash per node, out-of-band, ODM-lagged. There is no separate action for this one. The operator-level control that actually pays: know the bmcweb version on every node in your fleet and put a floor on it in your acceptance criteria for ODM firmware drops.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3409","https://github.com/openbmc/bmcweb/security/advisories/GHSA-g3qc-375m-h66j"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-34425","cve":"CVE-2022-34425","aliases":[],"title":"Dell Enterprise SONiC OS (SSH cryptographic key): A cryptographic key weakness in SONiC's SSH implementation lets an unauthenticated remote attacker exploit…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (SSH cryptographic key)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"A cryptographic key weakness in SONiC's SSH implementation lets an unauthenticated remote attacker exploit the switch. Shared or predictable SSH host keys across a product line mean an attacker can impersonate any switch to your automation, harvesting the credentials your config-management pushes. CVE-2025-38741 is the same class recurring in SONiC 4.5.0, which tells you it is a build-pipeline problem, not a one-off.","attack_vector":"Unauthenticated, remote — anyone able to interpose on or reach the switch's SSH service.","remediation":"NOS image upgrade plus reboot, and then **regenerate the switch's SSH host keys** — the upgrade alone does not replace a key that was already weak or shared. Update your automation's known_hosts afterwards. Verify host-key uniqueness across the fleet as a standing check.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34425","https://nvd.nist.gov/vuln/detail/CVE-2025-38741"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-35729","cve":"CVE-2022-35729","aliases":["INTEL-SA-00737"],"title":"Intel OpenBMC firmware (before version 0.72) - network-facing service: An unauthenticated caller reads out of bounds and crashes the BMC. Intel shipped this in the same SA-00737…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel OpenBMC firmware (before version 0.72) - network-facing service","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated caller reads out of bounds and crashes the BMC. Intel shipped this in the same SA-00737 bundle as the OpenBMC IPMI authentication bypass, which is the useful context: Intel-badged OpenBMC below 0.72 carried both a pre-auth takeover and a pre-auth crash at the same time. Losing the BMC on Intel-platform GPU nodes means losing remote power and console; an attacker who wants an operator blind while working on hosts starts here.","attack_vector":"Unauthenticated, over the network, against the BMC management interface.","remediation":"Fixed in Intel OpenBMC 0.72 and later - a BMC firmware update per node, out-of-band, through Intel's platform update packages. If you are on an affected Intel platform you are almost certainly also exposed to CVE-2021-39296 from the same advisory, so treat this as one campaign rather than two. Config-only first move remains the same: ACL the management network and disable IPMI over LAN.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-35729","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00737.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-3602","cve":"CVE-2022-3602","aliases":[],"title":"OpenSSL 3.0: X.509 email-address punycode buffer overflow (4-byte stack overflow)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSL 3.0","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"X.509 email-address punycode buffer overflow (4-byte stack overflow)","attack_vector":"Unauthenticated network (TLS peers/clients)","remediation":"Package update + restart consuming services; no reboot. Low real risk for a neocloud control plane but audit any mTLS front door","references":["https://access.redhat.com/security/cve/CVE-2022-3602"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-37122","cve":"CVE-2022-37122","aliases":["ZSL-2022-5709"],"title":"Carel pCOWeb HVAC BACnet gateway 2.1.0 (logdownload.cgi): Unauthenticated arbitrary file read off the gateway that bridges your BACnet field bus to IP. On its own that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Carel pCOWeb HVAC BACnet gateway 2.1.0 (logdownload.cgi)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated arbitrary file read off the gateway that bridges your BACnet field bus to IP. On its own that is information disclosure, but on this class of device the files worth reading are the ones that turn into control: stored web credentials, the BACnet device map, SNMP community strings, and the config that tells an attacker exactly which object instance is the CRAH fan-speed command and which is the chilled-water setpoint. This is the reconnaissance step that makes a subsequent thermal attack precise instead of a guess. Carel pCOWeb cards sit inside a very long list of OEM cooling products (chillers, CRAC/CRAH units, close-coupled cooling), so operators frequently have these in the hall without knowing the Carel name appears anywhere in their asset list.","attack_vector":"Unauthenticated HTTP GET on the facility network. No credentials, no user interaction. Reachability is the whole question: if your mechanical VLAN is flat with anything an attacker can phish into, this is a one-request win. These gateways also turn up on remote-access boxes that the mechanical contractor installed for support, which is how they end up internet-reachable.","remediation":"Carel firmware updates for pCOWeb are distributed through the OEM that embedded the card, not directly, so the practical path is: identify which of your cooling units carry a pCOWeb card (check the cooling vendor's BOM, not your CMDB), then ask that OEM for a firmware level that closes it. Expect a field technician and a maintenance window on live cooling, which most operators will not schedule for a file-read bug - so plan on segmentation as the actual control. Deny the gateway's web port from everything except the BMS supervisor, and log all access to it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-37122","https://www.zeroscience.mk/en/vulnerabilities/ZSL-2022-5709.php","https://packetstormsecurity.com/files/167684/"],"status":"curated"},{"id":"CVE-2022-3786","cve":"CVE-2022-3786","aliases":[],"title":"OpenSSL 3.0: X.509 email-address variable-length buffer overflow (DoS)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSL 3.0","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"X.509 email-address variable-length buffer overflow (DoS)","attack_vector":"Unauthenticated network","remediation":"Package update + service restart; no reboot","references":["https://access.redhat.com/security/cve/CVE-2022-3786"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2022-39278","cve":"CVE-2022-39278","aliases":[],"title":"Istio: Crafted message DoSes istiod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted message DoSes istiod","attack_vector":"Any pod on the cluster network","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-39278"],"status":"curated"},{"id":"CVE-2022-40242","cve":"CVE-2022-40242","aliases":[],"title":"AMI MegaRAC: Default credentials for the `sysadmin` account, shell access to the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Default credentials for the `sysadmin` account, shell access to the BMC","attack_vector":"Network / SSH to BMC","remediation":"Same intake-time credential rotation; add a fleet scan asserting no default BMC accounts remain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40242"],"status":"curated"},{"id":"CVE-2022-42276","cve":"CVE-2022-42276","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The SmiFlash SMM handler lets a privileged local user read, write and erase the SPI flash directly - a direct…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"The SmiFlash SMM handler lets a privileged local user read, write and erase the SPI flash directly - a direct write primitive to the platform firmware, with scope extending to other components. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42276","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42277","cve":"CVE-2022-42277","aliases":[],"title":"NVIDIA DGX Station - SBIOS / SMM firmware: The same SmiFlash read/write/erase primitive on DGX Station, giving a privileged local user arbitrary control…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Station - SBIOS / SMM firmware","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"The same SmiFlash read/write/erase primitive on DGX Station, giving a privileged local user arbitrary control of the platform flash. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42277","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-43377","cve":"CVE-2022-43377","aliases":["CVE-2022-43376","CVE-2022-43378","SEVD-2022-312-01"],"title":"Schneider Electric APC NetBotz 4 environmental appliances (355/450/455/550/570, V4.7.0 and prior): No rate limiting on authentication, so an attacker brute-forces the account and takes over the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric APC NetBotz 4 environmental appliances (355/450/455/550/570, V4.7.0 and prior)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"No rate limiting on authentication, so an attacker brute-forces the account and takes over the appliance, with stored XSS and clickjacking issues alongside it that let them hit an administrator's browser session. NetBotz appliances are the environmental eyes of the room - temperature, humidity, airflow, leak detection under liquid-cooled racks, door sensors, and camera pods. Control of one means the attacker decides what the operations team sees. During a cooling attack that is decisive: the temperature curve on the wall display stays flat while inlet temperatures climb toward the GPU thermal-shutdown threshold, and the first real signal is accelerators dropping off the fabric. The camera pods are a second concern - a NetBotz with camera modules is a video feed into the hall and the cage aisles, which is both a surveillance-evasion tool for someone about to walk in and a privacy exposure. Leak detection matters specifically for direct-liquid-cooled GPU racks, where a suppressed leak alarm is a route to real hardware destruction.","attack_vector":"Network access to the appliance's web interface, unauthenticated for the brute-force. NetBotz units sit on the facility monitoring VLAN, are polled by DCIM, and are frequently reachable from the corporate network because facilities staff want the camera feed. Weak or default passwords on these appliances are extremely common because they are installed once and never revisited.","remediation":"Firmware update to a fixed NetBotz 4 release per Schneider's SEVD-2022-312-01 - straightforward appliance firmware, no cooling impact, so schedule it. Then do the part that actually closes the brute-force risk regardless of version: set a strong unique password per appliance (they are usually all identical across a site), disable unused accounts, and put the appliance behind an ACL permitting only the DCIM collector and a jump host. Treat environmental telemetry integrity as a control objective in its own right - if the only temperature data you have comes from appliances on the same VLAN as everything else, you have no independent way to detect a cooling attack, so consider a second, separately-networked temperature source for critical GPU rows.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43377","https://nvd.nist.gov/vuln/detail/CVE-2022-43376","https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2022-312-01"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-43945","cve":"CVE-2022-43945","aliases":[],"title":"Linux nfsd (NFS server): NFSD buffer overflow - a client can force the send buffer to overflow the page array","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux nfsd (NFS server)","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"NFSD buffer overflow - a client can force the send buffer to overflow the page array","attack_vector":"Network (remote)","remediation":"Data-plane: kernel patch + rolling reboot of every NFS server - tenant-visible","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43945"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-46463","cve":"CVE-2022-46463","aliases":[],"title":"Harbor: Public and private image repositories accessible without authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":7.5,"severity":"high","kev":false,"impact":"Public and private image repositories accessible without authentication","attack_vector":"Unauthenticated network","remediation":"Upgrade Harbor and explicitly disable anonymous pull; audit whether tenant images were exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-46463"],"status":"curated"},{"id":"CVE-2023-0200","cve":"CVE-2023-0200","aliases":[],"title":"NVIDIA DGX-2 - SBIOS / SMM firmware: An out-of-bounds access in the OFBD SMM handler against a preconditioned heap reaches SMM code execution with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-2 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An out-of-bounds access in the OFBD SMM handler against a preconditioned heap reaches SMM code execution with a changed scope. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0200","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-788"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-0202","cve":"CVE-2023-0202","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The GenericSio and LegacySmmSredir SMM APIs allow arbitrary modification of SMRAM, which is full System…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The GenericSio and LegacySmmSredir SMM APIs allow arbitrary modification of SMRAM, which is full System Management Mode compromise. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0202","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-123"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-0206","cve":"CVE-2023-0206","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: The NVME SMM API allows arbitrary SMRAM modification, again reaching full SMM compromise. This is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The NVME SMM API allows arbitrary SMRAM modification, again reaching full SMM compromise. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0206","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-0207","cve":"CVE-2023-0207","aliases":[],"title":"NVIDIA DGX-2 - SBIOS / SMM firmware: Privileged code can modify the ServerSetup NVRAM variable at runtime, altering platform configuration below…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-2 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Privileged code can modify the ServerSetup NVRAM variable at runtime, altering platform configuration below the OS. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0207","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-732"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20578","cve":"CVE-2023-20578","aliases":[],"title":"AMD SMM communications buffer - TOCTOU (AMD-SB-3003): MULTI-TENANT ISOLATION: A time-of-check-to-time-of-use race on the SMM communications buffer lets ring-0 code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SMM communications buffer - TOCTOU (AMD-SB-3003)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A time-of-check-to-time-of-use race on the SMM communications buffer lets ring-0 code swap the buffer contents between validation and use, reaching arbitrary code execution in System Management Mode. Disclosed in the same August 2024 AMD server bulletin as SinkClose, and the same practical outcome: host root becomes ring -2, which is a compromise you cannot clean by reimaging.","attack_vector":"Local, requires ring-0 plus access to the BIOS menu or a UEFI shell - a meaningfully higher bar than plain root, but well within reach of anyone with physical or BMC-mediated console access to the node.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because the extra prerequisite is console/UEFI-shell access, BMC hardening is a genuine compensating control: restrict who can reach the virtual console and who can reboot a node into firmware setup.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20578","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-24574","cve":"CVE-2023-24574","aliases":[],"title":"Dell Enterprise SONiC OS (authentication component): FABRIC DOS: uncontrolled resource consumption in SONiC's authentication component, reachable by an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (authentication component)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: uncontrolled resource consumption in SONiC's authentication component, reachable by an unauthenticated remote attacker. Exhausting the authentication path is a good denial of service because it locks *you* out of the switch at the same time it stays up forwarding — you lose the ability to respond.","attack_vector":"Unauthenticated, remote to the switch management services on Enterprise SONiC 3.5.3, 4.0.0, 4.0.1, 4.0.2.","remediation":"NOS image upgrade plus reboot. Interim: rate-limit and ACL the management interface so only your jump hosts can reach the authentication endpoints — a live config change that also preserves your own access.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-24574"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-25191","cve":"CVE-2023-25191","aliases":[],"title":"AMI MegaRAC SPx (Redfish): Password disclosure through Redfish","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (Redfish)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Password disclosure through Redfish; credentials frequently reused across a whole fleet of identically-provisioned nodes","attack_vector":"Network / Redfish","remediation":"Firmware update to SPx_12-update-7.00 / SPx_13-update-5.00 plus a fleet-wide BMC credential rotation","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25191"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25506","cve":"CVE-2023-25506","aliases":[],"title":"NVIDIA DGX-1 - SBIOS / SMM firmware: An out-of-bounds access in the Ofbd handler in the AMI SBIOS reaches SMM code execution against a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-1 - SBIOS / SMM firmware","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An out-of-bounds access in the Ofbd handler in the AMI SBIOS reaches SMM code execution against a preconditioned heap, with scope extending to other components. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5458. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25506","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-788"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25521","cve":"CVE-2023-25521","aliases":[],"title":"DGX A100 / A800 SBIOS: Code execution + privesc in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 / A800 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Code execution + privesc in SBIOS","attack_vector":"Local operator / compromised host OS","remediation":"Flash SBIOS to 1.21 via DGX firmware update container; full node drain + power cycle","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5461/5461.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-250"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25522","cve":"CVE-2023-25522","aliases":[],"title":"DGX A100 / A800 SBIOS: DoS / data tampering / info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 / A800 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS / data tampering / info disclosure","attack_vector":"Local operator / compromised host OS","remediation":"Flash SBIOS 1.21; node drain + power cycle","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5461/5461.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25525","cve":"CVE-2023-25525","aliases":[],"title":"Cumulus Linux (switch OS): Cross-tenant info disclosure (VxLAN IPv6 mis-forwarding on SVI)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux (switch OS)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Cross-tenant info disclosure (VxLAN IPv6 mis-forwarding on SVI)","attack_vector":"Network-adjacent unauthenticated on the fabric","remediation":"Upgrade Cumulus Linux to 5.6.0+; rolling switch upgrade, fabric redundancy required","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5480/5480.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-284"]},{"id":"CVE-2023-27532","cve":"CVE-2023-27532","aliases":[],"title":"Veeam Backup & Replication: Encrypted credentials in the configuration database can be obtained","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Veeam Backup & Replication","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"[KEV] Encrypted credentials in the configuration database can be obtained -> access to backup infrastructure hosts","attack_vector":"Network (remote)","remediation":"Control-plane: patch + rotate every credential the backup server held","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27532"],"status":"curated"},{"id":"CVE-2023-28432","cve":"CVE-2023-28432","aliases":[],"title":"MinIO: Cluster returns all env vars incl. MINIO_SECRET_KEY and MINIO_ROOT_PASSWORD","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MinIO","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"[KEV] Cluster returns all env vars incl. MINIO_SECRET_KEY and MINIO_ROOT_PASSWORD","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + ROTATE root password and every derived tenant key","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28432"],"status":"curated"},{"id":"CVE-2023-28840","cve":"CVE-2023-28840","aliases":[],"title":"Docker / moby: Swarm overlay-network encryption silently not applied","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Swarm overlay-network encryption silently not applied; cluster traffic sent in the clear","attack_vector":"Anyone with access to the underlay network between nodes","remediation":"Upgrade moby; if using Swarm overlay encryption assume traffic was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28840"],"status":"curated"},{"id":"CVE-2023-29552","cve":"CVE-2023-29552","aliases":["SLP reflective amplification","AMI-SA-2023004"],"title":"AMI MegaRAC SPx 12 (Service Location Protocol service): Every BMC running SLP is a free DDoS cannon pointed at the internet, with an amplification factor reported…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 (Service Location Protocol service)","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"Every BMC running SLP is a free DDoS cannon pointed at the internet, with an amplification factor reported above 2000x. The damage is usually reputational and contractual before it is technical: your management subnet starts sourcing multi-gigabit reflected floods, upstream transit providers null-route you, and the abuse tickets land on the operator. Secondary effect is that the BMC's own control plane saturates, so you lose out-of-band management of the affected nodes exactly when you need it. CISA has this in the Known Exploited Vulnerabilities catalog - it is being used in the wild, not theoretically.","attack_vector":"Unauthenticated UDP to the SLP service on the BMC. Dangerous specifically where BMC management interfaces have been given routable or internet-reachable addresses, which happens more often than operators admit in leased colo, in early-stage neocloud builds, and on remote edge racks reached over a public IP.","remediation":"Config-only, no flash, no reboot - and it is one of the few in this cluster you can fix today. Disable SLP on the BMC (AMI ships it disabled on SPx_12 fixed builds and SPx_13 does not support it at all), and block UDP/427 inbound at the edge. Then do the structural fix: get BMC interfaces off any publicly routable address and behind a VPN or bastion. Verify by scanning your own management ranges from outside, because operators routinely discover BMCs they did not know were exposed.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023004.pdf","https://www.bitsight.com/blog/new-high-severity-vulnerability-cve-2023-29552-discovered-service-location-protocol-slp","https://www.cisa.gov/known-exploited-vulnerabilities-catalog","https://nvd.nist.gov/vuln/detail/CVE-2023-29552"],"status":"curated"},{"id":"CVE-2023-31032","cve":"CVE-2023-31032","aliases":[],"title":"DGX A100 SBIOS: Improper crypto config / UI-layer restriction in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper crypto config / UI-layer restriction in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+; node power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31032","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-627"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31035","cve":"CVE-2023-31035","aliases":[],"title":"DGX A100 SBIOS: Improper input validation in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper input validation in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31035","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31036","cve":"CVE-2023-31036","aliases":[],"title":"Triton Inference Server: RCE / privesc / data tampering via model-load path traversal (`--model-control explicit`)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"RCE / privesc / data tampering via model-load path traversal (`--model-control explicit`)","attack_vector":"Authenticated user of the inference endpoint; malicious model repo","remediation":"Upgrade Triton to 2.40+; rebuild inference serving images; disable explicit model control on multi-tenant endpoints","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5509/5509.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-23"]},{"id":"CVE-2023-3127","cve":"CVE-2023-3127","aliases":["ICSA-23-192-02"],"title":"Software House iSTAR Ultra, Ultra LT, Ultra G2 and Edge G2 door controllers: An unauthenticated user can log into the controller with administrator rights. No exploitation, no chain…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Software House iSTAR Ultra, Ultra LT, Ultra G2 and Edge G2 door controllers","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated user can log into the controller with administrator rights. No exploitation, no chain - just log in as admin on the panel that governs the doors into the hall and the cages inside it. Administrator on an iSTAR panel means unlocking doors, enrolling credentials, changing door schedules, and editing or clearing the local event log so the entry leaves no record. Whoever walks in can pull drives containing model weights and customer data, attach a console to a running node, plug into the out-of-band management switch and reach every BMC in the row, or leave a hardware implant behind. For a bare-metal GPU provider this is a direct breach of the physical isolation guarantee sold to tenants, and because the log can be cleared from the same session, you may never be able to prove it did or did not happen. This is the fourth distinct critical or high finding on the iSTAR platform in this database's window, which is itself the finding: treat the platform as needing continuous advisory tracking rather than set-and-forget.","attack_vector":"Unauthenticated network access to the controller on the physical-security VLAN. Nothing else is required. That VLAN typically also carries CCTV, intercom and the security integrator's remote-support path, any of which is a route in from a wider network.","remediation":"Firmware update per Johnson Controls' advisory for each affected iSTAR model - a security-integrator engagement with doors in local fallback during the flash, so it needs staff at affected doors for the window. Because the flaw grants administrator without authentication, assume any panel reachable during the exposure window may have had credentials added or the event log cleared: audit the panel's credential list and the head-end's cardholder database against a known-good baseline after patching. Structurally, put the physical-security VLAN behind a firewall with an explicit allow-list from the head-end only, and stop treating that VLAN as trusted because it is 'the security network'.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-192-02","https://nvd.nist.gov/vuln/detail/CVE-2023-3127","https://www.johnsoncontrols.com/cyber-solutions/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31315","cve":"CVE-2023-31315","aliases":[],"title":"AMD CPU (Sinkclose): Sinkclose: SMM lock bypass - ring-0 attacker gains SMM execution, enabling firmware-persistent implants that…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Sinkclose)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Sinkclose: SMM lock bypass - ring-0 attacker gains SMM execution, enabling firmware-persistent implants that survive reimaging","attack_vector":"Local ring-0 (post-kernel-escape), i.e. the second stage after any kernel privesc above","remediation":"AGESA/BIOS firmware update + reboot; requires a full firmware rollout across the fleet, not a package update. Some older EPYC SKUs never got fixes","references":["https://access.redhat.com/security/cve/CVE-2023-31315"],"status":"curated","fleet":{"ubiquity":"very common - AMD EPYC is a standard GPU-server host CPU and the host side of MI300 platforms; the flaw reaches back to 2006-era silicon","remediation_pain":"microcode+reboot / firmware-flash via an AGESA/BIOS update per node; AMD initially declined to patch some older Zen parts, leaving unpatchable-mitigate-only nodes in mixed fleets","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"A tenant (or escaped container) with kernel access reaches System Management Mode, the most privileged mode on the box, and can plant an SMM implant invisible to OS and hypervisor that survives a disk wipe."}},{"id":"CVE-2023-31320","cve":"CVE-2023-31320","aliases":[],"title":"AMD Radeon Graphics display driver - input validation: Improper input validation in the Radeon display driver lets an attacker corrupt the display and deny service.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Radeon Graphics display driver - input validation","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper input validation in the Radeon display driver lets an attacker corrupt the display and deny service. On headless datacenter Instinct nodes the display pipeline is barely exercised, so real exposure is low - included for completeness of the AMD GPU driver picture rather than because it should move up your queue.","attack_vector":"Local, via the display driver path. Effectively inert on headless compute nodes.","remediation":"Update the AMD graphics driver and reload or reboot. Low priority on headless GPU fleets.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31320","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-33411","cve":"CVE-2023-33411","aliases":[],"title":"Supermicro BMC web server on X11 and M11 based boards with firmware up to 3.17.02: An unauthenticated attacker reads files out of the BMC's filesystem. In practice that is the credential…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web server on X11 and M11 based boards with firmware up to 3.17.02","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated attacker reads files out of the BMC's filesystem. In practice that is the credential store, configuration, and keys - which is how this becomes the first link in a chain rather than an information-disclosure footnote: the traversal hands over the credentials needed to exploit CVE-2023-33412 and CVE-2023-33413 on the same box, and because BMC credentials are typically identical across a fleet, one node's disclosure unlocks all of them. Directory traversal in the HTTP server, reachable with no authentication whatsoever.","attack_vector":"Anything routable to the BMC's HTTP interface, unauthenticated. No credential, no host foothold, no tenant access needed - only a network path to the out-of-band management VLAN.","remediation":"Firmware flash to BMC 3.17.02 or later per board, from Supermicro's December 2023 advisory. Because this one needs no credentials, it should be at the front of the queue for any X11 fleet, and any node that was ever exposed to an untrusted network should have its BMC credentials treated as compromised and rotated - to unique per-node values, not another shared password. Network isolation buys time but does not help if your management VLAN is flat and reachable from tenant hosts.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-33411","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2023/33xxx/CVE-2023-33411.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-36835","cve":"CVE-2023-36835","aliases":[],"title":"Juniper Junos OS PFE on QFX10000 Series (VXLAN tunnel routing): FABRIC DOS: a specific *valid* IP packet that needs to be routed over a VXLAN tunnel wedges the Packet…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS PFE on QFX10000 Series (VXLAN tunnel routing)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: a specific *valid* IP packet that needs to be routed over a VXLAN tunnel wedges the Packet Forwarding Engine on a QFX10000. Because the trigger is a legitimate packet rather than a malformed one, no input filter catches it, and QFX10000 sits in the spine or super-spine role where a wedge partitions the fabric rather than dropping one rack.","attack_vector":"A network-based attacker — or, given the trigger is valid traffic, an unlucky workload — sending the specific packet into a VXLAN-routed path.","remediation":"Junos upgrade plus reboot on affected QFX10000 devices, staged so redundant spines are never both down. No filtering workaround, since the triggering packet is valid.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36835"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-38950","cve":"CVE-2023-38950","aliases":[],"title":"ZKTeco BioTime v8.5.5 (iclock API path traversal): Unauthenticated arbitrary file read on the BioTime server via the iclock API - and this one is in CISA's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ZKTeco BioTime v8.5.5 (iclock API path traversal)","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"Unauthenticated arbitrary file read on the BioTime server via the iclock API - and this one is in CISA's Known Exploited Vulnerabilities catalog, meaning it is being used against real targets, not theorised about. BioTime is the time-and-attendance and access management server that pairs with ZKTeco terminals, so the files worth reading include its configuration, database credentials and the stored data on who badges in where and when. Combined with the companion issues in the same disclosure set (an unauthenticated administrator password reset via a hidden API, and authenticated arbitrary file write) an attacker moves from file read to full control of the access-management server, and from there to door control. The KEV listing should drive urgency: if you have BioTime anywhere in the facility, treat it as a live target. Personnel movement data is also a targeting asset in its own right - it tells an attacker when the hall is unstaffed.","attack_vector":"Unauthenticated HTTP to the BioTime server's iclock API. BioTime is commonly published to the corporate network for HR and facilities use, and internet exposure is not rare because the product is sold on remote attendance management. Active exploitation is confirmed, so assume internet-reachable instances are already being scanned.","remediation":"Patch immediately to BioTime 9.0.1 (build 20240617.19506) or later per the vendor - and because this is KEV-listed with a known-exploited history, patching is not the end of the work: assume compromise on any instance that was internet-reachable, rotate the database and administrator credentials, audit the cardholder and access-rule data against a known-good baseline, and review door-open events for the exposure window. Then remove BioTime from internet and general corporate reachability entirely. If it is running on a server that also holds anything else, isolate it.","references":["https://www.cisa.gov/known-exploited-vulnerabilities-catalog","https://nvd.nist.gov/vuln/detail/CVE-2023-38950","https://nvd.nist.gov/vuln/detail/CVE-2023-38949"],"status":"curated"},{"id":"CVE-2023-39538","cve":"CVE-2023-39538","aliases":["LogoFAIL"],"title":"UEFI image parsers, AMI AptioV: Unrestricted upload of a crafted BMP logo parsed by the BIOS at boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI image parsers, AMI AptioV","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unrestricted upload of a crafted BMP logo parsed by the BIOS at boot; leads to code execution in DXE, before Secure Boot is enforced. Persistent, invisible to the OS","attack_vector":"Local write access to the ESP","remediation":"BIOS/UEFI firmware update per platform — the slowest update in the stack, gated on the ODM shipping an AMI rebase. No dbx-style shortcut exists because the flaw is in the firmware, not a signed binary","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39538"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-39539","cve":"CVE-2023-39539","aliases":["LogoFAIL"],"title":"UEFI image parsers, AMI AptioV: Second LogoFAIL image-parser flaw in AMI AptioV BIOS","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI image parsers, AMI AptioV","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Second LogoFAIL image-parser flaw in AMI AptioV BIOS; same pre-Secure-Boot code execution primitive","attack_vector":"Local, ESP write","remediation":"BIOS firmware update; on GPU nodes this means a full BIOS flash cycle with the node drained","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-39539"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-4346","cve":"CVE-2023-4346","aliases":[],"title":"KNX devices using KNX Connection Authorization Option 1 (BCU key): An attacker sets the BCU key on KNX devices and permanently locks legitimate operators out - the device…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"KNX devices using KNX Connection Authorization Option 1 (BCU key)","year":"2023","cvss_score":7.5,"severity":"high","kev":true,"impact":"An attacker sets the BCU key on KNX devices and permanently locks legitimate operators out - the device cannot be reset or reprogrammed without vendor intervention or physical replacement. This is in CISA's Known Exploited Vulnerabilities catalog, and it has been used in the wild to brick building-automation installations. Where KNX drives lighting, HVAC zone control or blinds in a facility, this is a denial of service you cannot recover from with a reboot or a firmware push: the devices are electronically bricked and the fix is a truck roll with replacement hardware, potentially across every affected device on the bus. In a datacenter context the exposure is usually peripheral rather than core cooling - KNX is more common in European commercial buildings than in purpose-built halls - but it appears in mixed-use and converted buildings, and in office and ancillary space attached to a datacenter. The operator-facing point is the recovery profile: unlike almost everything else on this list, there is no software remedy after the fact.","attack_vector":"Any device that can send telegrams on the KNX bus, or on a KNXnet/IP segment that routes onto it. KNX has no meaningful authentication in its classic form, so this requires only reachability. That means the building network, a KNX/IP router bridging segments, or physical access to the twisted-pair bus in an accessible space.","remediation":"Prevention only - once devices are locked, recovery means vendor unlock procedures where they exist or hardware replacement. Set the BCU key yourself to a known value on every device during commissioning so an attacker cannot claim it, and record it securely. Isolate KNXnet/IP strictly: no route from tenant, corporate or internet networks to the KNX segment, and audit every KNX/IP router for whether it is bridging more than it should. Where the installation supports it, migrate to KNX Secure (KNX Data Secure / IP Secure), which adds authentication and is the actual fix - a device-by-device upgrade project. Given the KEV listing, treat any internet-reachable KNXnet/IP interface as an emergency.","references":["https://www.cisa.gov/known-exploited-vulnerabilities-catalog","https://nvd.nist.gov/vuln/detail/CVE-2023-4346"],"status":"curated"},{"id":"CVE-2023-44199","cve":"CVE-2023-44199","aliases":[],"title":"Juniper Junos OS Packet Forwarding Engine (MX Series): FABRIC DOS: improper handling of unusual conditions in the Packet Forwarding Engine lets an unauthenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS Packet Forwarding Engine (MX Series)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: improper handling of unusual conditions in the Packet Forwarding Engine lets an unauthenticated network attacker deny service. A PFE-level failure is worse than a control-plane one — it stops the data plane, so traffic stops even if the routing engine stays up.","attack_vector":"Unauthenticated, network-based, against the PFE on Junos MX platforms.","remediation":"Junos upgrade plus reboot. MX platforms are usually the cluster's edge/border routers, so plan around a redundant pair — patch one side, fail over, patch the other.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-44199"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-44487","cve":"CVE-2023-44487","aliases":[],"title":"Envoy: \"HTTP/2 Rapid Reset\": stream-cancellation flood exhausts server resources. Exploited in the wild at massive…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"\"HTTP/2 Rapid Reset\": stream-cancellation flood exhausts server resources. Exploited in the wild at massive scale in late 2023","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy, Istio, Traefik, ingress-nginx and the Go runtime in every Go-based controller. Broad, multi-component rollout; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-44487"],"status":"curated"},{"id":"CVE-2023-4486","cve":"CVE-2023-4486","aliases":["ICSA-23-341-03"],"title":"Johnson Controls Metasys NAE55 / SNE / SNC network engines and Facility Explorer F4-SNC (before 11.0.6 / 12.0.4): Sending invalid credentials to the login endpoint knocks the engine over. These…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Johnson Controls Metasys NAE55 / SNE / SNC network engines and Facility Explorer F4-SNC (before 11.0.6 / 12.0.4)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Sending invalid credentials to the login endpoint knocks the engine over. These engines are not dashboards - they are the supervisory controllers that execute the control logic for air handlers, CRAHs and chilled-water plant, and coordinate the field controllers under them. Denial of service on an NAE or SNE means the hall loses coordinated thermal control and falls back to whatever the individual field controllers do standalone, which is typically last-known-value or a crude failsafe. With a 40-140 kW rack density and no supervisory loop closing on actual return-air temperature, you are running blind during exactly the period a competent attacker would also be driving load or blocking alarms. It is trivially repeatable, needs no valid credentials, and can be held indefinitely - so it is a sustained availability attack on the hall rather than a one-shot.","attack_vector":"Unauthenticated TCP to the engine's login endpoint on the facility network. Metasys engines are field-mounted in mechanical rooms and IDF closets and sit on the building VLAN; they are frequently reachable from anywhere on that VLAN with no ACL, because the assumption is that only the ADS talks to them. In practice the mechanical contractor's laptop, the fire-alarm integrator's gateway and the landlord's network all live there too.","remediation":"Firmware update on each engine to 11.0.6 / 12.0.4 or later. That is a controller flash, done by a Johnson Controls technician or certified integrator, engine by engine, with each engine offline during the flash - meaning a real maintenance window on live cooling that many operators will not schedule during peak season. Realistic interim control is an ACL that permits the engine's login port only from the ADS, which removes the attack surface without touching firmware. In a leased colo you cannot flash the landlord's engines: require the firmware version in writing and make DoS resilience of the mechanical control layer an explicit item in the SLA conversation, because a cooling outage caused by the landlord's unpatched engine still kills your training run.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-341-03","https://nvd.nist.gov/vuln/detail/CVE-2023-4486","https://www.johnsoncontrols.com/cyber-solutions/security-advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45232","cve":"CVE-2023-45232","aliases":["PixieFail","AMI-SA-2024001"],"title":"AMI AptioV UEFI BIOS (EDK II network stack, IPv6): An infinite loop when the firmware parses unknown options in an IPv6 Destination Options header. A single…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (EDK II network stack, IPv6)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An infinite loop when the firmware parses unknown options in an IPv6 Destination Options header. A single crafted packet hangs the node in pre-boot firmware - it never reaches the OS, never reports in, and cannot be recovered by a normal reboot because it will hang again on the next boot as long as the attacker keeps sending. On a GPU cluster this is a targeted capacity-denial tool: hold a set of nodes out of the scheduler indefinitely, and because the node is stuck below the OS your host-level monitoring shows nothing but silence.","attack_vector":"Network-reachable, unauthenticated, no interaction, low complexity - one packet during the node's network boot window. Anything that can put IPv6 traffic on the provisioning or boot segment qualifies, including a compromised neighbouring node.","remediation":"BIOS update with the patched EDK II network package - firmware flash plus reboot per node, vendor-rebase-gated. The cheap and immediately available mitigation is the same as for the rest of the PixieFail family: disable network/PXE boot in BIOS where it is not needed, and where it is, put the provisioning network behind strict segmentation so no untrusted host can send packets into the boot window. BIOS setup change plus one reboot, versus a firmware flash campaign.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024001.pdf","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h","https://nvd.nist.gov/vuln/detail/CVE-2023-45232"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45233","cve":"CVE-2023-45233","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (IPv6 Destination Options header, PadN option parsing): Same shape as the unknown-option hang but reached through the PadN option, which is trivially craftable. One…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IPv6 Destination Options header, PadN option parsing)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Same shape as the unknown-option hang but reached through the PadN option, which is trivially craftable. One packet parks a node in firmware forever. In a netboot-driven cluster this is a cheap way to deny an operator their fleet during a reprovisioning window, and the failure looks like a hardware fault rather than an attack.","attack_vector":"Unauthenticated, on-link attacker sending crafted IPv6 packets to nodes during network boot.","remediation":"Firmware flash from the server OEM, one reboot per node. No runtime fix. Practical interim control is to disable IPv6 network boot in the UEFI setup (a config change, deployable via the OEM's remote BIOS-settings tooling without a flash) and to keep the provisioning VLAN reachable only from the deployment controllers.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45233","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45236","cve":"CVE-2023-45236","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (TCP initial sequence number generation): The firmware's TCP initial sequence numbers are predictable, so an off-path attacker can inject into or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (TCP initial sequence number generation)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The firmware's TCP initial sequence numbers are predictable, so an off-path attacker can inject into or hijack the boot-time TCP session. In practice that means substituting the payload the node is downloading - the HTTP-boot image, the kernel, the initrd - without ever being on the wire. For a bare-metal GPU cloud that HTTP-boots tenant images, this is a supply-chain swap at provisioning time that no post-boot integrity check will notice if the swapped image is what gets measured.","attack_vector":"Off-path attacker who can guess the ISN - no need to sit on the provisioning segment at all, which makes this materially worse than the on-link PixieFail bugs. Unauthenticated, pre-OS.","remediation":"OEM BIOS update; the fix replaces the ISN generator, so there is no configuration toggle that helps. Flash + reboot per node. Compensating control while you wait: use HTTPS boot with proper certificate validation rather than plain HTTP/TFTP, and verify signatures on the downloaded image inside the boot flow rather than relying on transport integrity.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45236","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45237","cve":"CVE-2023-45237","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (PseudoRandom number generation used by the network stack): The weak PRNG behind the previous issue - the firmware's randomness source is not random enough for anything…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (PseudoRandom number generation used by the network stack)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"The weak PRNG behind the previous issue - the firmware's randomness source is not random enough for anything security-relevant it feeds, including sequence numbers and transaction identifiers used during netboot. The operator-visible consequence is that boot-time network exchanges are spoofable by an attacker who does not need to see them.","attack_vector":"Off-path or on-path attacker predicting firmware-generated values during network boot. Unauthenticated, pre-OS.","remediation":"Firmware flash via the server OEM. Reboot per node. There is nothing to configure - the entropy source is compiled in. Until patched, assume boot-time network exchanges are forgeable and lean on cryptographic verification of the boot payload (signed images, Secure Boot with your own keys) rather than on network trust.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45237","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-4607","cve":"CVE-2023-4607","aliases":["LEN-140960"],"title":"Lenovo XClarity Controller (XCC) - permission API: An authenticated XCC user can change the permissions of any user through a crafted API command - including…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - permission API","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"An authenticated XCC user can change the permissions of any user through a crafted API command - including their own. Fixed in the same advisory as the password-change flaw and with the same practical result: whatever limited BMC account you issued becomes an administrative one, and the attacker inherits out-of-band power control, Virtual Media boot, console access and firmware update rights on the node. The privilege model in this XCC generation should be treated as advisory rather than enforcing until patched.","attack_vector":"Any authenticated XCC account, at any privilege level, reaching the XCC over the out-of-band management VLAN.","remediation":"Flash XCC to the per-model version in LEN-140960 - out-of-band, per-node, no host reboot, no job drain. Because the fix is per-SKU, treat it as one campaign covering both this and CVE-2023-4606. Until patched, the only real control is reducing the number of XCC accounts that exist at all, since privilege tiers are not a boundary here.","references":["https://support.lenovo.com/us/en/product_security/LEN-140960","https://nvd.nist.gov/vuln/detail/CVE-2023-4607"],"status":"curated"},{"id":"CVE-2023-49298","cve":"CVE-2023-49298","aliases":[],"title":"OpenZFS: Block-cloning path can replace file contents with zero bytes, potentially disabling security mechanisms","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenZFS","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Block-cloning path can replace file contents with zero bytes, potentially disabling security mechanisms","attack_vector":"Network (remote)","remediation":"Data-plane: module upgrade, node drain + reboot; scrub and verify datasets","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49298"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-49933","cve":"CVE-2023-49933","aliases":[],"title":"Slurm: Improper message-integrity enforcement allows RPC traffic modification","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper message-integrity enforcement allows RPC traffic modification","attack_vector":"Anyone on the cluster management network","remediation":"Upgrade Slurm; isolate the Slurm control network","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49933"],"status":"curated"},{"id":"CVE-2023-50272","cve":"CVE-2023-50272","aliases":["HPESBHF04584"],"title":"HPE iLO 5 / iLO 6 (authentication bypass): Authentication bypass on the iLO itself, remotely, with no credentials. That is the whole out-of-band plane…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 / iLO 6 (authentication bypass)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Authentication bypass on the iLO itself, remotely, with no credentials. That is the whole out-of-band plane on a ProLiant or Apollo node: power control, Virtual Media to boot an attacker-supplied image, remote console into the tenant's session, and a firmware-level foothold that persists through host reimaging. The CVSS vector marks scope as changed, which reflects exactly that - getting the BMC gets you more than the BMC. Affects iLO 5 from v2.63 up to (not including) v3.00, and iLO 6 from v1.05 up to v1.55, so it spans both the Gen10/Gen10 Plus and Gen11 fleets.","attack_vector":"Anything routable to the iLO address on the out-of-band management VLAN, unauthenticated. Attack complexity is rated high, so it is not a trivial one-shot, but it requires no account and no host access - the exposure is defined purely by who can reach the iLO.","remediation":"Flash iLO 5 to v3.00 or later, iLO 6 to v1.55 or later. Out-of-band, per-node, via the iLO web UI, iLOrest, Redfish or OneView - no host reboot and no drain of running jobs; the iLO resets itself and OOB access is unavailable for a couple of minutes. Interim config-only control: restrict the iLO management network to an explicit allowlist of jump hosts, since there is no per-feature toggle that closes an authentication bypass.","references":["https://support.hpe.com/hpesc/public/docDisplay?docLocale=en_US&docId=hpesbhf04584en_us","https://nvd.nist.gov/vuln/detail/CVE-2023-50272"],"status":"curated"},{"id":"CVE-2023-51232","cve":"CVE-2023-51232","aliases":[],"title":"Dagster (webserver): Directory traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dagster (webserver)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal → sensitive file disclosure","attack_vector":"Unauthenticated network to the Dagster webserver","remediation":"Upgrade past 1.5.11","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-51232"],"status":"curated"},{"id":"CVE-2023-52454","cve":"CVE-2023-52454","aliases":["nvmet-tcp invalid H2C PDU length panic","NVMe/TCP DATAL kernel NULL deref"],"title":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c: FABRIC DOS: A host sending an H2CData command with a DATAL inconsistent with the packet size drives a NULL…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF TCP target, drivers/nvme/target/tcp.c","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: A host sending an H2CData command with a DATAL inconsistent with the packet size drives a NULL pointer dereference in nvmet_tcp_build_pdu_iovec() and panics the target kernel. The PDU length was also never checked against the MAXH2CDATA value the target itself advertised during connection setup. On a storage node serving many GPU tenants, one malformed PDU from one tenant takes the node down and every attached volume with it - and since NVMe/TCP is unauthenticated by default, the attacker does not need to be a tenant at all, just reachable.","attack_vector":"Connect to the NVMe/TCP target and send an H2CData PDU whose DATAL does not match the actual packet size, or which exceeds the negotiated MAXH2CDATA. Trivial to construct, instantly fatal to the target, and repeatable after every reboot until patched.","remediation":"Host reboot / kernel upgrade on nvmet-tcp targets. Interim: restrict port 4420 to known initiator addresses and enable in-band authentication (on a patched kernel) so an arbitrary peer cannot reach the PDU parser. If a storage node is serving production tenants and cannot be rebooted immediately, the firewall restriction is the meaningful control - this is remotely triggerable with no state.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2023/CVE-2023-52454.json","https://nvd.nist.gov/vuln/detail/CVE-2023-52454"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52883","cve":"CVE-2023-52883","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix possible null pointer dereference","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52883","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-6021","cve":"CVE-2023-6021","aliases":[],"title":"Ray (log API): LFI — read any file on the head node, unauthenticated","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (log API)","year":"2023","cvss_score":7.5,"severity":"high","kev":false,"impact":"LFI — read any file on the head node, unauthenticated","attack_vector":"Unauthenticated network to the dashboard","remediation":"Upgrade to 2.8.1+; head-node secrets and cloud credentials are in scope","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6021"],"status":"curated"},{"id":"CVE-2024-0096","cve":"CVE-2024-0096","aliases":[],"title":"ChatRTX: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0096","https://github.com/NVIDIA/product-security/tree/main/2024/5533"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-269"]},{"id":"CVE-2024-0097","cve":"CVE-2024-0097","aliases":[],"title":"ChatRTX: Local privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Local privesc","attack_vector":"Local Windows user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0097","https://github.com/NVIDIA/product-security/tree/main/2024/5533"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:H/I:H/A:N","cwe":["CWE-269"]},{"id":"CVE-2024-0101","cve":"CVE-2024-0101","aliases":[],"title":"Mellanox OS / MetroX / Onyx / Skyway: Switch/gateway DoS (improper input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox OS / MetroX / Onyx / Skyway","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Switch/gateway DoS (improper input validation)","attack_vector":"Network-adjacent unauthenticated on the fabric","remediation":"Upgrade switch OS image; rolling switch reload","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0101","https://github.com/NVIDIA/product-security/tree/main/2024/5559"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-693"]},{"id":"CVE-2024-0112","cve":"CVE-2024-0112","aliases":[],"title":"IGX Orin / Jetson AGX Orin: Privesc / DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"IGX Orin / Jetson AGX Orin","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Privesc / DoS","attack_vector":"Local attacker on the device","remediation":"Flash IGX/Jetson firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0112","https://github.com/NVIDIA/product-security/tree/main/2025/5611"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-20"]},{"id":"CVE-2024-0113","cve":"CVE-2024-0113","aliases":[],"title":"NVIDIA Mellanox OS, ONYX, Skyway, MetroX-2/MetroX-3 XC: A crafted URI causes CGI path traversal in the switch web interface, reaching privilege escalation and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Mellanox OS, ONYX, Skyway, MetroX-2/MetroX-3 XC","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A crafted URI causes CGI path traversal in the switch web interface, reaching privilege escalation and information disclosure on the switch itself. These are the InfiniBand and Ethernet switches carrying your east-west training traffic and, on MetroX, your inter-site links.","attack_vector":"Network, requires the switch web management interface to be reachable and a user to follow a crafted URI. Switch management planes are usually far more reachable inside the datacenter than operators assume.","remediation":"Upgrade the switch OS per bulletin 5563. Cost: a switch OS upgrade means a reboot and a link flap - on a fat-tree fabric plan it leaf-by-leaf with ECMP draining, or you take collective jobs down. Disable the web interface entirely and manage via CLI/gNMI if you can.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0113","https://github.com/NVIDIA/product-security/tree/main/2024/5563"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-35"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0762","cve":"CVE-2024-0762","aliases":["UEFIcanhazbufferoverflow","CVE-2024-1598"],"title":"Phoenix SecureCore (TPM configuration / SetupUtility, unsafe UEFI variable handling in SMM): A buffer overflow in how SecureCore handles a TPM-configuration UEFI variable inside SMM. An attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Phoenix SecureCore (TPM configuration / SetupUtility, unsafe UEFI variable handling in SMM)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A buffer overflow in how SecureCore handles a TPM-configuration UEFI variable inside SMM. An attacker overwrites adjacent SMM memory, escalates to ring -2, and installs a bootkit that persists through OS reinstall and disk replacement. Eclypsium's point in naming it was breadth: the same SecureCore code ships across Alder Lake, Coffee Lake, Comet Lake, Ice Lake, Jasper Lake, Kaby Lake, Meteor Lake, Raptor Lake, Rocket Lake and Tiger Lake, so hundreds of PC and server models from multiple OEMs inherit it from one IBV defect. That inheritance pattern is the real lesson for a fleet operator - your firmware exposure is set by an IBV you have no contract with.","attack_vector":"Local admin/root on the host OS writing the vulnerable UEFI variable, then triggering the SMM path. NVD scores it AV:L/PR:L; Phoenix's own CNA scoring assumes higher privilege and higher complexity, which is why the two scores differ (7.8 NVD vs 7.5 Phoenix).","remediation":"BIOS update from your server or system OEM built on the fixed Phoenix SecureCore version - Phoenix lists per-platform fixed versions and OEMs shipped on their own schedules through mid-to-late 2024. Firmware flash plus one reboot per node. No config workaround: the TPM configuration variable is part of normal platform setup and cannot be disabled. Audit by silicon generation rather than by this CVE alone - Phoenix filed the Gemini Lake instance of the same defect as a separate advisory (CVE-2024-1598, fixed in SecureCore for Gemini Lake 4.1.0.567), and OEM release notes commonly cite only one of the two, so a fleet can be patched for the headline case and still exposed on Gemini Lake-based management, edge or storage nodes in the same racks. Where a node cannot be patched promptly, restrict who gets administrative access to the host OS, since that is the entry condition.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0762","https://eclypsium.com/blog/ueficanhazbufferoverflow-widespread-impact-from-vulnerability-in-popular-pc-and-server-firmware/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-10188","cve":"CVE-2024-10188","aliases":[],"title":"LiteLLM: Unauthenticated DoS via `ast.literal_eval` on user input","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LiteLLM","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated DoS via `ast.literal_eval` on user input","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10188"],"status":"curated"},{"id":"CVE-2024-10403","cve":"CVE-2024-10403","aliases":[],"title":"Brocade Fabric OS (firmware download credential capture): Fabric OS captures the SFTP/FTP server password used for a firmware download. The credential to your firmware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade Fabric OS (firmware download credential capture)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Fabric OS captures the SFTP/FTP server password used for a firmware download. The credential to your firmware distribution server ends up recoverable from the switch — and that server holds the images you install on every switch in the fabric, so it is a step from 'read a password' to 'supply the next firmware image'. Companion CVE-2023-3489 logs the same password in clear text into SupportSave bundles, which then get emailed to vendor support.","attack_vector":"An attacker with access to the switch or to a SupportSave bundle taken from it.","remediation":"Upgrade Fabric OS past 8.2.3e2 / 9.2.0c / 9.2.1a as applicable — firmware install plus reboot. Immediately: rotate the firmware-server credential, use a single-purpose account with read-only access to the image share, and scrub existing SupportSave archives before sharing them with support.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10403","https://nvd.nist.gov/vuln/detail/CVE-2023-3489"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-11322","cve":"CVE-2024-11322","aliases":[],"title":"CyberPower PowerPanel Business 4.11.0 - Service Watchdog on TCP/2003: An unauthenticated attacker can repeatedly restart the ppbd.exe process via the watchdog service, keeping the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Business 4.11.0 - Service Watchdog on TCP/2003","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unauthenticated attacker can repeatedly restart the ppbd.exe process via the watchdog service, keeping the power-management daemon permanently down. The consequence is not a crash you notice - it is that the software which would have gracefully shut down your fleet during a utility event is not running when the event happens. This is a denial of the safety mechanism, and its cost only materialises during the incident it was supposed to soften.","attack_vector":"Unauthenticated, to TCP/2003 on the host running PowerPanel Business.","remediation":"Upgrade PowerPanel Business, and firewall TCP/2003 to only the hosts that legitimately need it. Also worth building: an alert on the power-management daemon being down, since the failure mode here is silence.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11322"],"status":"curated"},{"id":"CVE-2024-1558","cve":"CVE-2024-1558","aliases":[],"title":"MLflow (`_create_model_version`): Path traversal in model-version creation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (`_create_model_version`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Path traversal in model-version creation","attack_vector":"Authenticated or unauthenticated model registration","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1558"],"status":"curated"},{"id":"CVE-2024-1561","cve":"CVE-2024-1561","aliases":[],"title":"Gradio (`/component_server`): Arbitrary method invocation on components","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio (`/component_server`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary method invocation on components → local file read","attack_vector":"Unauthenticated network to any exposed Gradio demo","remediation":"Upgrade to 4.19.2+. Tenant-launched Gradio demos on GPU nodes routinely get public share links — provider should block or gate `share=True` egress","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1561"],"status":"curated","fleet":{"ubiquity":"Very common - Gradio is the default demo/eval UI shipped in AI containers and on Hugging Face Spaces; frequently exposed with `share=True` from a GPU node","remediation_pain":"**Image rebuild** (`daemon-restart` per app) - upgrade to Gradio 4.13.0+ in every image that bundles it, which in practice is most inference/demo images","pain_class":"daemon-restart","why_fleet_wide":"`/component_server` invokes arbitrary `Component` methods, so `move_resource_to_block_cache()` reads any file on the host - API keys and cloud credentials in env/files - from an internet-exposed demo running on a GPU node"}},{"id":"CVE-2024-1728","cve":"CVE-2024-1728","aliases":[],"title":"Gradio: Local file inclusion via improper input validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Local file inclusion via improper input validation","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1728"],"status":"curated"},{"id":"CVE-2024-21595","cve":"CVE-2024-21595","aliases":[],"title":"Juniper Junos OS Packet Forwarding Engine (VXLAN + ICMP): FABRIC DOS: a high rate of specific ICMP traffic to a device with VXLAN configured deadlocks the Packet…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS Packet Forwarding Engine (VXLAN + ICMP)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: a high rate of specific ICMP traffic to a device with VXLAN configured deadlocks the Packet Forwarding Engine and leaves the switch unresponsive — recovery requires a manual restart, so it does not self-heal. On a VXLAN/EVPN GPU fabric this is a tenant-reachable way to require hands-on intervention on a leaf, and ICMP is not something most operators filter inside the fabric.","attack_vector":"Unauthenticated, network-based — an attacker able to send ICMP at rate toward a VXLAN-configured Junos device. Any tenant workload qualifies.","remediation":"Junos upgrade plus reboot. Immediate mitigation is a control-plane policer rate-limiting ICMP toward the device — a live config change, no downtime, and it converts a manual-restart outage into a throttled nuisance. Related Junos VXLAN PFE issues: CVE-2023-36835 (QFX10000, PFE wedge on a valid IP packet routed over a VXLAN tunnel) and CVE-2022-22171.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21595","https://nvd.nist.gov/vuln/detail/CVE-2023-36835","https://nvd.nist.gov/vuln/detail/CVE-2022-22171"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-24974","cve":"CVE-2024-24974","aliases":[],"title":"OpenVPN: The interactive service pipe is reachable remotely","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The interactive service pipe is reachable remotely -> interact with the privileged OpenVPN service","attack_vector":"Network (remote)","remediation":"Control-plane: patch; restrict the service named pipe","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24974"],"status":"curated"},{"id":"CVE-2024-26147","cve":"CVE-2024-26147","aliases":[],"title":"Helm: Uninitialized variable panic parsing index and plugin YAML","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Uninitialized variable panic parsing index and plugin YAML","attack_vector":"Malicious chart repo index","remediation":"Upgrade Helm to 3.14.2+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26147"],"status":"curated"},{"id":"CVE-2024-27318","cve":"CVE-2024-27318","aliases":[],"title":"ONNX: Directory traversal in `external_data` — bypass of the 1.13 fix","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ONNX","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Directory traversal in `external_data` — bypass of the 1.13 fix","attack_vector":"Customer-supplied ONNX model","remediation":"Upgrade past 1.15.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27318"],"status":"curated"},{"id":"CVE-2024-28028","cve":"CVE-2024-28028","aliases":[],"title":"Intel Neural Compressor: Unauthenticated input-validation failure leading to escalation of privilege in Neural Compressor. Another…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Neural Compressor","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated input-validation failure leading to escalation of privilege in Neural Compressor. Another reason not to leave the optimisation service on an open cluster network.","attack_vector":"Unauthenticated, network-reachable where the service is exposed.","remediation":"Upgrade to v3.0 or later and gate the service behind auth and network policy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28028","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated"},{"id":"CVE-2024-28869","cve":"CVE-2024-28869","aliases":[],"title":"Traefik: GET with a Content-Length header hangs the endpoint indefinitely","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"GET with a Content-Length header hangs the endpoint indefinitely; ingress DoS","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28869"],"status":"curated"},{"id":"CVE-2024-29966","cve":"CVE-2024-29966","aliases":["CVE-2024-29960","CVE-2024-29965"],"title":"Brocade SANnav OVA appliance image, before v2.3.1 and v2.3.0a: Three defects that together mean every SANnav OVA deployment shares the same secrets. The documentation ships…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Brocade SANnav OVA appliance image, before v2.3.1 and v2.3.0a","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Three defects that together mean every SANnav OVA deployment shares the same secrets. The documentation ships what is effectively the appliance's root password; every VM built from the official OVA carries identical SSH host keys, so SSH to SANnav is trivially machine-in-the-middle-able; and appliance backups are created world-readable, so a local user can copy a backup, restore it onto their own appliance and recover the passwords of every switch in the fabric. For an operator this is the cheapest possible path from 'someone got a shell somewhere near the management network' to 'attacker holds admin on every FC switch', and therefore to zoning changes that expose one tenant's LUNs to another.","attack_vector":"The root password and SSH keys are usable by anyone with network access to the appliance; the world-readable backup requires only an unprivileged local account on the SANnav host.","remediation":"Upgrade to SANnav 2.3.1 / 2.3.0a, then do the cleanup the upgrade does not do for you: change the appliance root password, regenerate the SSH host keys so your instance is no longer keyed identically to every other deployment, fix permissions on existing backup files and move them off the appliance, and rotate every switch credential that a stolen backup would have exposed. Management-plane work only - no switch firmware flash, no fabric outage.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29966","https://nvd.nist.gov/vuln/detail/CVE-2024-29960","https://nvd.nist.gov/vuln/detail/CVE-2024-29965"],"status":"curated"},{"id":"CVE-2024-31142","cve":"CVE-2024-31142","aliases":["XSA-455"],"title":"Xen (x86 speculation): Incorrect logic for BTC/SRSO mitigations - guests are unprotected against branch-type confusion despite…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 speculation)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Incorrect logic for BTC/SRSO mitigations - guests are unprotected against branch-type confusion despite mitigations appearing enabled","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + host reboot. Worth flagging: \"mitigation reported as on\" was false, so audit rather than trust the sysfs status","references":["https://xenbits.xen.org/xsa/advisory-455.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-31916","cve":"CVE-2024-31916","aliases":["IBM X-Force 290026"],"title":"IBM OpenBMC bmcweb HTTPS server (FW1050.00 - FW1050.10): Certain URIs on IBM's OpenBMC-derived bmcweb return their content to callers who never authenticated. The…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC bmcweb HTTPS server (FW1050.00 - FW1050.10)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Certain URIs on IBM's OpenBMC-derived bmcweb return their content to callers who never authenticated. The Redfish tree is where BMC-side inventory lives - serial numbers, firmware versions, sensor and account metadata - so an unauthenticated reader on the management VLAN gets a precise map of the fleet: which nodes run which firmware, and therefore which nodes are still vulnerable to everything else in this cluster. It is reconnaissance rather than control, but it is the reconnaissance that makes a targeted BMC campaign cheap.","attack_vector":"Unauthenticated HTTPS to the BMC's Redfish/web endpoint. Any host that can route to the management network.","remediation":"Fixed in IBM firmware after FW1050.10; delivery is an OpenPower/Power system firmware update, which on IBM hardware is a supported in-band update path rather than a raw SPI flash, but still a per-node reboot-class operation with a maintenance window. Config-only first move: confirm no BMC in the fleet answers HTTPS from outside your management VLAN, and treat any Redfish data reachable pre-auth as public.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31916","https://www.ibm.com/support/pages/node/7158679","https://exchange.xforce.ibmcloud.com/vulnerabilities/290026"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-34997","cve":"CVE-2024-34997","aliases":[],"title":"joblib (`NumpyArrayWrapper.read_array`): Deserialization vulnerability in joblib 1.4.2","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"joblib (`NumpyArrayWrapper.read_array`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Deserialization vulnerability in joblib 1.4.2","attack_vector":"Customer-supplied joblib artifact","remediation":"Contested as intended pickle behavior; the real control is format policy, not a version bump","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-34997"],"status":"curated"},{"id":"CVE-2024-35124","cve":"CVE-2024-35124","aliases":["IBM X-Force 290674"],"title":"IBM OpenBMC default password and session management (FW1020, FW1030, FW1050): The combination of a shipped default password and how sessions are managed lets an attacker reach BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC default password and session management (FW1020, FW1030, FW1050)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"The combination of a shipped default password and how sessions are managed lets an attacker reach BMC administrator. Full administrative control of the BMC means power state, virtual media, console, firmware update - i.e. the ability to install a persistent implant and to reimage or brick nodes. Default-credential findings are unglamorous and they are also how BMC fleets actually get owned: an operator who racked 400 nodes and changed the OS credentials but not the BMC's is the modal case.","attack_vector":"Network access to the BMC plus, per the CVSS vector, some user interaction and higher attack complexity. Practically: an attacker on the management network against a BMC whose default credential was never rotated or where a stale session can be reused.","remediation":"Fixed in IBM firmware past FW1050.10 / FW1030.50 / FW1020.60 - a per-node system firmware update with a maintenance window. The action that matters more and costs nothing: audit every BMC in the fleet for the vendor default credential right now, rotate to per-node unique passwords, and make credential rotation part of node provisioning rather than a one-time sweep. On a rented bare-metal fleet, also rotate BMC credentials between tenants.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35124","https://www.ibm.com/support/pages/node/7163195","https://exchange.xforce.ibmcloud.com/vulnerabilities/290674"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-35178","cve":"CVE-2024-35178","aliases":[],"title":"Jupyter Server (Windows): Unauthenticated attackers can leak the NTLM hash of the host","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server (Windows)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated attackers can leak the NTLM hash of the host","attack_vector":"Unauthenticated network to the notebook server","remediation":"Upgrade; Windows GPU hosts only","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35178"],"status":"curated"},{"id":"CVE-2024-36347","cve":"CVE-2024-36347","aliases":[],"title":"AMD CPU (EntrySign): Improper signature verification in the AMD CPU microcode patch loader - a ring-0 attacker can load arbitrary…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (EntrySign)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper signature verification in the AMD CPU microcode patch loader - a ring-0 attacker can load arbitrary microcode, defeating SEV-SNP attestation","attack_vector":"Local ring-0 / compromised host; breaks confidential-VM guarantees","remediation":"AGESA/BIOS firmware update + reboot across the fleet; SEV-SNP attestation reports from unpatched hosts cannot be trusted","references":["https://access.redhat.com/security/cve/CVE-2024-36347"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-36354","cve":"CVE-2024-36354","aliases":["BadRAM (SMM variant)"],"title":"AMD - DIMM SPD address aliasing bypassing SMM isolation (AMD-SB-3014): MULTI-TENANT ISOLATION: The BadRAM SPD-aliasing technique aimed at System Management Mode rather than at SEV.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD - DIMM SPD address aliasing bypassing SMM isolation (AMD-SB-3014)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The BadRAM SPD-aliasing technique aimed at System Management Mode rather than at SEV. By lying about a DIMM's size in its serial-presence-detect chip, an attacker creates physical address aliases that let ring-0 code reach into SMRAM and execute at SMM - the level above the hypervisor. Where the SEV-facing BadRAM breaks confidential VMs, this one breaks the platform outright, and it lands beneath every detection tool you run.","attack_vector":"Either brief physical access to the DIMM's SPD chip (the published rig is a ~$10 microcontroller), or - crucially - **ring-0 on a host that has non-compliant DIMMs with unlocked SPD, which needs no physical access at all**. That second path is what makes this a real datacenter concern rather than an evil-maid curiosity.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. The fix adds a boot-time alias-detection scan. Beyond patching, change procurement: specify SPD-lockable DIMMs and verify the lock is actually set, because on non-compliant modules a remote ring-0 attacker reaches this without ever entering your building. Patch this together with the SEV-facing BadRAM CVE - they ship in the same firmware wave.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36354","https://badram.eu/","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3014.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-36432","cve":"CVE-2024-36432","aliases":[],"title":"BIOS firmware on Supermicro X11DPG-HGX2, X11PDG-QT, X11PDG-OT and X11PDG-SN before version 4.4: An arbitrary memory write primitive inside platform firmware, on the boards that carry HGX GPU…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BIOS firmware on Supermicro X11DPG-HGX2, X11PDG-QT, X11PDG-OT and X11PDG-SN before version 4.4","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An arbitrary memory write primitive inside platform firmware, on the boards that carry HGX GPU baseboards. What an attacker gets is the ability to corrupt or take over firmware-privileged execution - which on x86 means SMM and the pre-boot environment, below the hypervisor, below the host kernel, and outside anything the operator's security tooling can see. Scope is marked changed in the CVSS vector, meaning the compromise crosses a security boundary. On a GPU host this is the persistence layer under a very expensive, very heavily shared machine. The X11DPG-HGX2 is the head node board for NVIDIA HGX-2 baseboards, so this is literally GPU-platform BIOS rather than generic server BIOS.","attack_vector":"Local, high-privilege access to the host - root or equivalent on the node's operating system, with high attack complexity. The realistic actor is a tenant on leased bare metal, or an attacker who already has host root and wants to convert it into something that survives the node being wiped and re-let.","remediation":"BIOS flash to version 4.4 or later from Supermicro's July 2024 BIOS advisory. BIOS updates on these boards require a host reboot and, on some SKUs, a BMC-mediated update, so this is a per-node maintenance window on machines that are usually running multi-day training jobs - schedule it into a drain cycle rather than expecting an ad-hoc window. There is no config-only mitigation: the write primitive is in the firmware itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36432","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/36xxx/CVE-2024-36432.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-36434","cve":"CVE-2024-36434","aliases":[],"title":"Supermicro BIOS SMM callout (X11DPH-T / X11DPH-Tq): Execution in System Management Mode, the most privileged execution context on the machine - above the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BIOS SMM callout (X11DPH-T / X11DPH-Tq)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Execution in System Management Mode, the most privileged execution context on the machine - above the hypervisor, invisible to the OS, and able to write the platform's own firmware. An SMM implant is the deepest persistent foothold available on an x86 node: it survives OS reinstall, disk replacement, and hypervisor redeployment, and no host-based tooling can enumerate it. For a bare-metal operator this is the finding that breaks the tenant handoff guarantee outright, because you cannot prove a returned node is clean. And X11DPH-i before version 4.4 - SMM code calling out to memory the attacker controls, which is the classic route from ring 0 into ring -2.","attack_vector":"Local, host-side, high privilege - root on the node's OS, triggering the SMI that reaches the vulnerable callout. Attack complexity is rated high, so it is a targeted attack rather than an opportunistic one, but a bare-metal tenant has unlimited time and full access to attempt it.","remediation":"BIOS flash to 4.4 or later from Supermicro's July 2024 BIOS advisory, per board. Same image covers the arbitrary-write issues on the neighbouring X11DPH SKUs, so batch them. There is no configuration change that mitigates an SMM callout. If you rent bare metal on these boards, the additional operational control worth adding is a firmware measurement taken at node return and compared against a known-good baseline, since a patched BIOS does not tell you whether the node was implanted before you patched it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36434","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/36xxx/CVE-2024-36434.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-37125","cve":"CVE-2024-37125","aliases":[],"title":"Dell SmartFabric OS10 (uncontrolled resource consumption): FABRIC DOS: a remote unauthenticated host can exhaust resources on an OS10 switch across 10.5.3.x through…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (uncontrolled resource consumption)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: a remote unauthenticated host can exhaust resources on an OS10 switch across 10.5.3.x through 10.5.6.x. No credentials, no adjacency requirement beyond IP reachability — so any tenant workload that can address the switch can attempt it.","attack_vector":"Unauthenticated remote host with IP reachability to the switch.","remediation":"OS10 upgrade plus reload. Interim: control-plane policing and management ACLs restricting who can address the switch at all — live config changes that are worth having permanently.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37125"],"status":"curated"},{"id":"CVE-2024-37775","cve":"CVE-2024-37775","aliases":[],"title":"Sunbird DCIM dcTrack v9.1.2 - ticket location RBAC: Incorrect access control lets an attacker create or update tickets against locations they should not have…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Sunbird DCIM dcTrack v9.1.2 - ticket location RBAC","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Incorrect access control lets an attacker create or update tickets against locations they should not have access to, bypassing the RBAC check. In a multi-tenant colo or a shared cage environment, that is a tenant-boundary problem inside the facility workflow system - work orders touching another customer's rack.","attack_vector":"Any authenticated dcTrack user.","remediation":"Upgrade past 9.1.2. If you run dcTrack with per-customer location scoping as a tenant-isolation control, treat that control as having been ineffective for the affected period and review the ticket history.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37775"],"status":"curated"},{"id":"CVE-2024-38486","cve":"CVE-2024-38486","aliases":[],"title":"Dell SmartFabric OS10 (command injection): Command injection in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x. Listed as a distinct entry because its…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (command injection)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Command injection in SmartFabric OS10 10.5.5.4-10.5.5.10 and 10.5.6.x. Listed as a distinct entry because its affected-version window is narrower than the later OS10 command-injection batch, so a fleet on 10.5.5.x needs this specific fix even if it has applied a 10.6.x-targeted advisory elsewhere.","attack_vector":"An attacker able to supply input to the affected OS10 command path.","remediation":"OS10 upgrade plus switch reload. Verify the fixed-version table against your exact running build — Dell's OS10 advisories have heavily overlapping but non-identical version ranges and it is easy to conclude you are patched when you are not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38486","https://nvd.nist.gov/vuln/detail/CVE-2024-39577"],"status":"curated"},{"id":"CVE-2024-38813","cve":"CVE-2024-38813","aliases":[],"title":"VMware vCenter: Privilege escalation to root on vCenter via a crafted network packet","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware vCenter","year":"2024","cvss_score":7.5,"severity":"high","kev":true,"impact":"Privilege escalation to root on vCenter via a crafted network packet [KEV]","attack_vector":"Network access to the management plane","remediation":"vCenter patch + service restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38813"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-39321","cve":"CVE-2024-39321","aliases":[],"title":"Traefik: IP allow-lists bypassed via HTTP/3 early data in QUIC 0-RTT with spoofed addresses","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"IP allow-lists bypassed via HTTP/3 early data in QUIC 0-RTT with spoofed addresses","attack_vector":"Unauthenticated network","remediation":"Rolling Traefik upgrade; disable 0-RTT","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39321"],"status":"curated"},{"id":"CVE-2024-39719","cve":"CVE-2024-39719","aliases":[],"title":"Ollama: File-existence disclosure via `api/create`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"File-existence disclosure via `api/create`","attack_vector":"Unauthenticated network","remediation":"Upgrade; enumerates provider host paths","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39719"],"status":"curated"},{"id":"CVE-2024-39722","cve":"CVE-2024-39722","aliases":[],"title":"Ollama: Path traversal in `api/push` discloses server filesystem layout","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Path traversal in `api/push` discloses server filesystem layout","attack_vector":"Unauthenticated network","remediation":"Upgrade past 0.1.46","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39722"],"status":"curated"},{"id":"CVE-2024-40634","cve":"CVE-2024-40634","aliases":[],"title":"Argo CD: Large JSON payload to /api/webhook DoSes the API server from an unauthenticated position","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Large JSON payload to /api/webhook DoSes the API server from an unauthenticated position","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40634"],"status":"curated"},{"id":"CVE-2024-40992","cve":"CVE-2024-40992","aliases":["RDMA/rxe UD responder length check regression"],"title":"Linux kernel - RDMA/rxe unreliable datagram responder, drivers/infiniband/sw/rxe/rxe_resp.c: FABRIC DOS: The IB architecture says a UD request packet with an invalid length must be silently…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe unreliable datagram responder, drivers/infiniband/sw/rxe/rxe_resp.c","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: The IB architecture says a UD request packet with an invalid length must be silently dropped, but a regression made rxe return a malformed-WQE error that pushes the UD queue pair into the ERROR state instead. A remote sender only has to transmit one oversized UD packet to permanently break the victim's UD queue pair. UD queue pairs carry management traffic and are used by MPI and by connection establishment, so killing them takes the node out of collective communication - the job stalls rather than fails cleanly, which is the expensive failure mode on a large training run.","attack_vector":"Send a UD packet whose payload is larger than the receiver's posted receive buffer. Unauthenticated, connectionless by definition (UD), reachable from anywhere on the fabric that can address the victim's QP. No exploit primitive needed - just a packet that is too big.","remediation":"Host reboot / kernel upgrade. Where rxe is not needed - the common case on clusters with real RNICs - unload and blacklist rdma_rxe instead (config change, no downtime). This one is a stability item as much as a security item: it also fires accidentally under MTU mismatches, so fixing it removes a class of unexplained job hangs.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2024/CVE-2024-40992.json","https://nvd.nist.gov/vuln/detail/CVE-2024-40992"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42145","cve":"CVE-2024-42145","aliases":["IB/core implement a limit on UMAD receive list"],"title":"Linux kernel InfiniBand core (ib_umad): ib_umad kept received management datagrams on an unbounded list. Any node on the fabric that sends MADs…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand core (ib_umad)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"ib_umad kept received management datagrams on an unbounded list. Any node on the fabric that sends MADs faster than userspace drains them exhausts host memory and takes the node down. This is a fabric-wide, unauthenticated availability attack against every host running a subnet manager, an IB diagnostic agent, or UFM's host-side collectors - one malicious or misbehaving endpoint degrades the management plane for the entire cluster.","attack_vector":"Any unauthenticated node attached to the InfiniBand subnet, sending a flood of management datagrams. No credentials anywhere.","remediation":"Upgrade the host kernel to 6.10 or a stable backport (4.19.318, 5.4.280, 5.10.222, 5.15.163, 6.1.98, 6.6.39, 6.9.9) - the fix caps the list at 200k entries. Rolling reboot, prioritizing the nodes that run OpenSM/UFM agents. No firmware flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42145","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-42145.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45436","cve":"CVE-2024-45436","aliases":[],"title":"Ollama (`extractFromZipFile`): Zip-slip: archive members extracted outside the parent directory","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (`extractFromZipFile`)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Zip-slip: archive members extracted outside the parent directory","attack_vector":"Customer-supplied model archive","remediation":"Upgrade past 0.1.47","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45436"],"status":"curated"},{"id":"CVE-2024-45776","cve":"CVE-2024-45776","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (gettext / message catalogue): Integer overflow reading a crafted translation catalogue gives both an out-of-bounds read and write. Language…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (gettext / message catalogue)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Integer overflow reading a crafted translation catalogue gives both an out-of-bounds read and write. Language files are unsigned data on the boot partition, which makes them an easy carrier for a bootkit on a node an attacker has held once.","attack_vector":"Attacker-supplied .mo file in GRUB's locale directory - local root or previous tenant.","remediation":"grub2 package update + reboot. Removing locale files from a server image is a cheap surface reduction.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45776","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45777","cve":"CVE-2024-45777","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (gettext / message catalogue): Second integer overflow in the same translation path, producing a heap out-of-bounds write and pre-boot code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (gettext / message catalogue)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Second integer overflow in the same translation path, producing a heap out-of-bounds write and pre-boot code execution.","attack_vector":"Attacker-supplied locale catalogue on the boot partition.","remediation":"grub2 package update + reboot; covered by the same distro update as the rest of the batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45777","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45782","cve":"CVE-2024-45782","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (HFS filesystem parser): An unbounded strcpy of the HFS volume name overflows a fixed buffer. About as direct a memory-corruption…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HFS filesystem parser)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"An unbounded strcpy of the HFS volume name overflows a fixed buffer. About as direct a memory-corruption primitive as this codebase contains, and it fires on nothing more than attaching a crafted volume.","attack_vector":"Attacker-supplied HFS volume, including one presented over BMC virtual media.","remediation":"grub2 package update + reboot. Strip the HFS module if you never boot Apple-formatted media, which on a GPU fleet is always.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45782","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45807","cve":"CVE-2024-45807","aliases":[],"title":"Envoy: Stream-management bugs in the default oghttp HTTP/2 codec","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Stream-management bugs in the default oghttp HTTP/2 codec","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45807"],"status":"curated"},{"id":"CVE-2024-4956","cve":"CVE-2024-4956","aliases":[],"title":"Sonatype Nexus Repository 3: Unauthenticated path traversal","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Sonatype Nexus Repository 3","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated path traversal -> read arbitrary system files from the artifact host","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade to 3.68.1+; rotate anything readable on that host","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-4956"],"status":"curated"},{"id":"CVE-2024-50608","cve":"CVE-2024-50608","aliases":[],"title":"Fluent Bit: Prometheus Remote Write input crashes on a Content-Length: 0 packet","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fluent Bit","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Prometheus Remote Write input crashes on a Content-Length: 0 packet","attack_vector":"Network (remote)","remediation":"Data-plane: DaemonSet image bump across the fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50608"],"status":"curated"},{"id":"CVE-2024-50609","cve":"CVE-2024-50609","aliases":[],"title":"Fluent Bit: OpenTelemetry input plugin crashes on a Content-Length: 0 packet","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Fluent Bit","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"OpenTelemetry input plugin crashes on a Content-Length: 0 packet -> log-pipeline outage","attack_vector":"Network (remote)","remediation":"Data-plane: DaemonSet image bump across the fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50609"],"status":"curated"},{"id":"CVE-2024-53270","cve":"CVE-2024-53270","aliases":[],"title":"Envoy: Load-shed path assumes an active request exists","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Load-shed path assumes an active request exists; proxy crash","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53270"],"status":"curated"},{"id":"CVE-2024-54084","cve":"CVE-2024-54084","aliases":["AMI-SA-2025003"],"title":"AMI AptioV UEFI BIOS: A time-of-check-to-time-of-use race in the BIOS leading to arbitrary code execution with a changed scope.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"A time-of-check-to-time-of-use race in the BIOS leading to arbitrary code execution with a changed scope. This is the quiet half of the March 2025 AMI advisory - it shipped alongside the CVSS 10.0 MegaRAC authentication bypass that got all the attention, and most operators patched the BMC and forgot the BIOS. Successful exploitation puts attacker code in the firmware boot path, below the OS and below any EDR you run, on a node that will keep passing every host-level integrity check you have.","attack_vector":"Local access with high privileges, high attack complexity. Needs root or kernel code on the host plus the ability to win a timing window during a firmware operation. On bare-metal GPU rentals the tenant holds that privilege by contract; on managed nodes it requires a prior host compromise.","remediation":"BIOS update to BKC_5.38 or later - firmware flash plus a full host reboot, per node, gated on your server vendor rebasing. Check specifically whether your March 2025 remediation covered the BIOS: many fleets flashed only the MegaRAC fix for CVE-2024-54085 from the same advisory and left this one open. No config-only mitigation exists for a TOCTOU in firmware.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025003.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-54084"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-55567","cve":"CVE-2024-55567","aliases":["INSYDE-SA-2024018"],"title":"Insyde InsydeH2O (UsbCoreDxe SMM module): Another SMM callout in the USB core driver - improper input validation lets SMM be redirected into…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (UsbCoreDxe SMM module)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Another SMM callout in the USB core driver - improper input validation lets SMM be redirected into attacker-controlled code outside SMRAM, giving ring -2 execution. Notable mainly because it is the same driver family Insyde has now patched repeatedly (2021 through 2024), which tells an operator something useful: assume the USB stack in your BIOS will need patching again, and build the flash cadence to match rather than treating each one as a one-off.","attack_vector":"Local admin/root on the host OS triggering the vulnerable SMI.","remediation":"OEM BIOS update on Insyde kernel 5.4 / 05.47.01, 5.5 / 05.55.01, 5.6 / 05.62.01, 5.7 / 05.71.01 or later. Firmware flash, one reboot per node. Partial config workaround: disable USB legacy/emulation support in BIOS on headless GPU nodes, which shrinks the reachable surface without a flash - but confirm on your platform that it actually unloads the SMM module.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-55567","https://www.insyde.com/security-pledge/sa-2024018/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-7409","cve":"CVE-2024-7409","aliases":[],"title":"QEMU (NBD server): Improper synchronisation during socket closure - DoS of the QEMU NBD server","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (NBD server)","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper synchronisation during socket closure - DoS of the QEMU NBD server","attack_vector":"Unauthenticated network to the NBD port; tenant VM guest","remediation":"QEMU update + restart; do not expose NBD to tenant networks","references":["https://access.redhat.com/security/cve/CVE-2024-7409"],"status":"curated"},{"id":"CVE-2025-0312","cve":"CVE-2025-0312","aliases":[],"title":"Ollama (GGUF import): Crafted GGUF causes DoS on model create","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF import)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted GGUF causes DoS on model create","attack_vector":"Customer-supplied GGUF model file","remediation":"Upgrade past 0.3.14; GGUF parsing is unhardened C-adjacent code reachable by any model upload","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0312"],"status":"curated"},{"id":"CVE-2025-0624","cve":"CVE-2025-0624","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (network config file search): grub_net_search_config_file copies a network-controlled variable with strcpy into a fixed buffer. This is the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (network config file search)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"grub_net_search_config_file copies a network-controlled variable with strcpy into a fixed buffer. This is the highest-priority entry in the 2025 batch for anyone running netboot: an attacker on the provisioning segment corrupts GRUB's memory during PXE and takes the node before the OS exists.","attack_vector":"Anyone who can respond on the network boot path - rogue DHCP server, compromised provisioning host, or a tenant that has been given L2 access to the provisioning VLAN by mistake.","remediation":"grub2 package update + reboot, AND rebuild/replace the netboot GRUB binary served over TFTP/HTTP - the served image is the actual attack surface here and patching running nodes does not touch it. Segment the provisioning network away from tenant traffic.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0624","https://access.redhat.com/security/cve/CVE-2025-0624"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-1137","cve":"CVE-2025-1137","aliases":[],"title":"IBM Storage Scale (command input neutralization): An authenticated user can execute privileged commands due to improper input neutralization. On a shared…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale (command input neutralization)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An authenticated user can execute privileged commands due to improper input neutralization. On a shared storage cluster the set of 'authenticated users' is much wider than the set of storage admins — it typically includes every tenant service account with a filesystem role.","attack_vector":"Authenticated user on Storage Scale 5.2.2.0 or 5.2.2.1 under certain configurations.","remediation":"Upgrade Storage Scale past 5.2.2.1 — rolling node upgrade. Review which accounts hold command-line access to the storage cluster in the meantime; most fleets have accumulated more than they intended.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1137"],"status":"curated"},{"id":"CVE-2025-14847","cve":"CVE-2025-14847","aliases":[],"title":"MongoDB Server: Mismatched Zlib compressed header lengths","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MongoDB Server","year":"2025","cvss_score":7.5,"severity":"high","kev":true,"impact":"[KEV] Mismatched Zlib compressed header lengths -> unauthenticated read of uninitialized heap memory","attack_vector":"Network (remote)","remediation":"Control-plane: patch the metadata store; block unauthenticated wire-protocol reach","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14847"],"status":"curated"},{"id":"CVE-2025-23320","cve":"CVE-2025-23320","aliases":[],"title":"NVIDIA Triton (Python backend): Information disclosure from the Python backend","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"NVIDIA Triton (Python backend)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Information disclosure from the Python backend","attack_vector":"Unauthenticated network","remediation":"Patch; leaks the shared-memory key that enables the write primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23320"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-209"]},{"id":"CVE-2025-23321","cve":"CVE-2025-23321","aliases":[],"title":"NVIDIA Triton Inference Server: An invalid request triggers a divide-by-zero and kills the server process. On a shared inference tier this is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An invalid request triggers a divide-by-zero and kills the server process. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23321","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-369"]},{"id":"CVE-2025-23322","cve":"CVE-2025-23322","aliases":[],"title":"NVIDIA Triton Inference Server: Cancelling a stream before it is processed causes a double free and crashes the server. On a shared inference…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Cancelling a stream before it is processed causes a double free and crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23322","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-415"]},{"id":"CVE-2025-23323","cve":"CVE-2025-23323","aliases":[],"title":"NVIDIA Triton Inference Server: An integer overflow on an invalid request leads to a segmentation fault. On a shared inference tier this is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An integer overflow on an invalid request leads to a segmentation fault. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23323","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"]},{"id":"CVE-2025-23324","cve":"CVE-2025-23324","aliases":[],"title":"NVIDIA Triton Inference Server: A second integer-overflow-to-segfault path on invalid requests. On a shared inference tier this is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"A second integer-overflow-to-segfault path on invalid requests. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23324","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"]},{"id":"CVE-2025-23325","cve":"CVE-2025-23325","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input drives uncontrolled recursion and exhausts the stack. On a shared inference tier this is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input drives uncontrolled recursion and exhausts the stack. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23325","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-674"]},{"id":"CVE-2025-23326","cve":"CVE-2025-23326","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input causes an integer overflow and crashes the server. On a shared inference tier this is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input causes an integer overflow and crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23326","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-680"]},{"id":"CVE-2025-23327","cve":"CVE-2025-23327","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input causes an integer overflow that reaches data tampering as well as denial of service - one of…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input causes an integer overflow that reaches data tampering as well as denial of service - one of the few in this set with an integrity impact, so served results can be corrupted rather than just interrupted. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23327","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-190"]},{"id":"CVE-2025-23328","cve":"CVE-2025-23328","aliases":[],"title":"NVIDIA Triton Inference Server: Crafted input causes an out-of-bounds write in the server. On a shared inference tier this is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted input causes an out-of-bounds write in the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5691. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23328","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-787"]},{"id":"CVE-2025-23329","cve":"CVE-2025-23329","aliases":[],"title":"NVIDIA Triton Inference Server: MULTI-TENANT ISOLATION: an attacker who can locate and reach the Python backend's shared memory region…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: an attacker who can locate and reach the Python backend's shared memory region corrupts it directly. Anything co-located in that region belongs to other inference requests. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5691. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23329","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-284"]},{"id":"CVE-2025-23331","cve":"CVE-2025-23331","aliases":[],"title":"NVIDIA Triton Inference Server: An invalid request drives an excessive memory allocation and a segmentation fault. On a shared inference tier…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"An invalid request drives an excessive memory allocation and a segmentation fault. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23331","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-789"]},{"id":"CVE-2025-24357","cve":"CVE-2025-24357","aliases":[],"title":"vLLM (weight loading): `hf_model_weights_iterator` uses `torch.load` without `weights_only`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (weight loading)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"`hf_model_weights_iterator` uses `torch.load` without `weights_only` → RCE","attack_vector":"Customer-supplied Hub checkpoint","remediation":"Upgrade; a serving engine loading pickle weights inherits every PyTorch pickle CVE","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24357"],"status":"curated"},{"id":"CVE-2025-2704","cve":"CVE-2025-2704","aliases":[],"title":"OpenVPN: Corrupting and replaying early-handshake packets against a tls-crypt-v2 server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"OpenVPN","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Corrupting and replaying early-handshake packets against a tls-crypt-v2 server -> denial of service","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade the VPN concentrator","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2704"],"status":"curated"},{"id":"CVE-2025-27817","cve":"CVE-2025-27817","aliases":[],"title":"Apache Kafka (client): SASL/OAUTHBEARER endpoint URLs accept file://","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Kafka (client)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"SASL/OAUTHBEARER endpoint URLs accept file:// -> arbitrary file read and SSRF","attack_vector":"Network (remote)","remediation":"Control-plane: client library bump across all internal producers/consumers","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-27817"],"status":"curated"},{"id":"CVE-2025-30202","cve":"CVE-2025-30202","aliases":[],"title":"vLLM (ZeroMQ): DoS and data exposure over ZeroMQ","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (ZeroMQ)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS and data exposure over ZeroMQ","attack_vector":"Network to the internal socket","remediation":"Upgrade to 0.8.5+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-30202"],"status":"curated"},{"id":"CVE-2025-33201","cve":"CVE-2025-33201","aliases":[],"title":"NVIDIA Triton Inference Server: Extra-large payloads trigger an improper-exceptional-condition check and crash the server. On a shared…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Extra-large payloads trigger an improper-exceptional-condition check and crash the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5734. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33201","https://github.com/NVIDIA/product-security/tree/main/2025/5734"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-754"]},{"id":"CVE-2025-33202","cve":"CVE-2025-33202","aliases":[],"title":"NVIDIA Triton Inference Server: Extra-large payloads cause a stack overflow and kill the server. On a shared inference tier this is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Extra-large payloads cause a stack overflow and kill the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5723. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33202","https://github.com/NVIDIA/product-security/tree/main/2025/5723"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-121"]},{"id":"CVE-2025-33211","cve":"CVE-2025-33211","aliases":[],"title":"NVIDIA Triton Inference Server: Improper validation of a specified quantity in input crashes the server. On a shared inference tier this is a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper validation of a specified quantity in input crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5734. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33211","https://github.com/NVIDIA/product-security/tree/main/2025/5734"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-1284"]},{"id":"CVE-2025-33238","cve":"CVE-2025-33238","aliases":[],"title":"Triton Inference Server: Remote DoS via race condition","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Remote DoS via race condition","attack_vector":"Network-adjacent unauthenticated client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33238","https://github.com/NVIDIA/product-security/tree/main/2026/5790"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362"]},{"id":"CVE-2025-33254","cve":"CVE-2025-33254","aliases":[],"title":"Triton Inference Server: Remote DoS via race condition in network processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Remote DoS via race condition in network processing","attack_vector":"Network-adjacent unauthenticated client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33254","https://github.com/NVIDIA/product-security/tree/main/2026/5790"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362"]},{"id":"CVE-2025-33255","cve":"CVE-2025-33255","aliases":[],"title":"TensorRT-LLM: Privesc via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Privesc via unsafe deserialization","attack_vector":"Malicious engine/model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33255","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-38590","cve":"CVE-2025-38590","aliases":["net/mlx5e remove skb secpath if xfrm state is not found"],"title":"Linux kernel mlx5_core IPsec RX offload: When hardware reports an xfrm state ID for a decrypted packet whose state has already been freed, the secpath…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core IPsec RX offload","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"When hardware reports an xfrm state ID for a decrypted packet whose state has already been freed, the secpath extension is left attached with length zero and the policy check reads sp->xvec[-1], faulting the kernel. Any remote IPsec peer can crash a node doing hardware IPsec offload on ConnectX - relevant if you encrypt tenant traffic in flight across the fabric.","attack_vector":"Remote IPsec peer, unauthenticated with respect to this bug - the peer just needs to be in an SA that gets torn down while packets are in flight.","remediation":"Upgrade the host kernel to 6.17 or a stable backport (6.6.102, 6.12.42, 6.15.10, 6.16.1). Rolling reboot of nodes doing IPsec offload. Interim: move IPsec off hardware offload to software xfrm (config change, CPU cost, no reboot).","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38590","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-38590.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38741","cve":"CVE-2025-38741","aliases":[],"title":"Dell Enterprise SONiC OS 4.5.0 (SSH cryptographic key): The SSH cryptographic-key weakness recurring in Enterprise SONiC 4.5.0, three years after the same class was…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS 4.5.0 (SSH cryptographic key)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"The SSH cryptographic-key weakness recurring in Enterprise SONiC 4.5.0, three years after the same class was fixed in 4.0.x. For an operator the lesson is that SONiC host-key uniqueness is not something to assume from a version number — check it directly on every switch you deploy.","attack_vector":"Unauthenticated, remote against the switch's SSH service.","remediation":"NOS image upgrade plus reboot, then regenerate host keys and refresh your automation's trust store. Consider adding a fleet-wide host-key uniqueness assertion to your provisioning tests.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38741"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-48956","cve":"CVE-2025-48956","aliases":[],"title":"vLLM (HTTP GET): Single HTTP GET crashes the server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (HTTP GET)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Single HTTP GET crashes the server","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade to 0.10.1.1+. Trivially exploitable against any internet-exposed tenant endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48956"],"status":"curated"},{"id":"CVE-2025-54502","cve":"CVE-2025-54502","aliases":["APCB SMM driver LocateProtocol misuse"],"title":"AMD Platform Configuration Blob (APCB) SMM driver, EPYC and Instinct MI300A/MI300C: Incorrect use of the UEFI LocateProtocol boot service in the APCB SMM driver lets ring 0 escalate to SMM and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"AMD Platform Configuration Blob (APCB) SMM driver, EPYC and Instinct MI300A/MI300C","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Incorrect use of the UEFI LocateProtocol boot service in the APCB SMM driver lets ring 0 escalate to SMM and execute arbitrary code. What makes this one worth flagging for AI operators specifically is the affected list: alongside EPYC 7002 through 9005 it explicitly includes AMD Instinct MI300A and the MI300C-based EPYC 9V64H. Those are accelerator nodes, frequently sold bare metal or with device passthrough, where the customer legitimately has ring 0. Ring -2 code on an MI300A node persists across every tenant that follows.","attack_vector":"Privileged local attacker at ring 0 - on a bare-metal MI300A rental that is the customer by design.","remediation":"Platform Initialization firmware per AMD-SB-7054: MI300A 1.0.0.C (OEM release 2025-12-11), MI300C 1.0.0.3 (2025-12-10), TurinPI 1.0.0.9 for EPYC 9005 (2025-12-31), GenoaPI 1.0.0.H, MilanPI 1.0.0.J, RomePI 1.0.0.P. BIOS flash and reboot per node - on an Instinct fleet that means evicting whatever is training on it. Given the bare-metal exposure, prioritize Instinct nodes over general-purpose EPYC and add a firmware measurement to the between-tenants checklist.","references":["https://www.amd.com/en/resources/product-security/bulletin/AMD-SB-7054.html","https://nvd.nist.gov/vuln/detail/CVE-2025-54502"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-55551","cve":"CVE-2025-55551","aliases":[],"title":"PyTorch (`torch.linalg.lu`): DoS on slice operation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (`torch.linalg.lu`)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS on slice operation","attack_vector":"Tenant-controlled tensor shapes; matters for shared-GPU multi-tenant serving","remediation":"Tenant-owned code path. Provider action is per-tenant GPU/process isolation so a crash does not take out co-tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55551"],"status":"curated"},{"id":"CVE-2025-55558","cve":"CVE-2025-55558","aliases":[],"title":"PyTorch (KV/conv path buffer overflow): Buffer overflow when a model combines Conv2d + hardshrink + view","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (KV/conv path buffer overflow)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Buffer overflow when a model combines Conv2d + hardshrink + view","attack_vector":"Customer-supplied model graph","remediation":"Tenant-owned; provider isolates blast radius per container","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55558"],"status":"curated"},{"id":"CVE-2025-5777","cve":"CVE-2025-5777","aliases":[],"title":"Citrix NetScaler ADC/Gateway: \"CitrixBleed 2\" - insufficient input validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2025","cvss_score":7.5,"severity":"high","kev":true,"impact":"[KEV] \"CitrixBleed 2\" - insufficient input validation -> memory overread of session material","attack_vector":"Network (remote)","remediation":"Control-plane: patch and kill all sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-5777"],"status":"curated"},{"id":"CVE-2025-58187","cve":"CVE-2025-58187","aliases":[],"title":"Go crypto/x509 (Tailscale, Go infra): Name-constraint checking scales non-linearly with certificate size","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Go crypto/x509 (Tailscale, Go infra)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Name-constraint checking scales non-linearly with certificate size -> CPU exhaustion when validating chains","attack_vector":"Network (remote)","remediation":"Control-plane: Go toolchain bump and rebuild of every TLS-validating service","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-58187"],"status":"curated"},{"id":"CVE-2025-59425","cve":"CVE-2025-59425","aliases":[],"title":"vLLM (API key comparison): Timing attack recovers the API key","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (API key comparison)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Timing attack recovers the API key","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade to 0.11.0+; do not rely on vLLM's own API key as the tenant auth boundary — front it with a gateway","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-59425"],"status":"curated"},{"id":"CVE-2025-59538","cve":"CVE-2025-59538","aliases":[],"title":"Argo CD: Azure DevOps webhook credentials mishandled, allowing unauthorised webhook use","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Azure DevOps webhook credentials mishandled, allowing unauthorised webhook use","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade; rotate webhook credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-59538"],"status":"curated"},{"id":"CVE-2025-65002","cve":"CVE-2025-65002","aliases":[],"title":"Fujitsu / Fsas Technologies iRMC S6 BMC (M5-generation servers): A length-boundary bug in BMC authentication - a username of exactly the wrong length gets handled incorrectly…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Fujitsu / Fsas Technologies iRMC S6 BMC (M5-generation servers)","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"A length-boundary bug in BMC authentication - a username of exactly the wrong length gets handled incorrectly and access control does not apply as intended. What the operator loses is the guarantee that their Redfish and web management surface enforces the roles they configured. This is worth carrying in a GPU-fleet database less for Fujitsu's market share than for the pattern: BMC authentication logic keeps failing on boundary conditions in string handling, and operators who assume Redfish role enforcement is sound are relying on code with a long history of exactly this class of defect. Servers before firmware 1.37S - Redfish and web UI access control mishandles the case where a username is exactly 16 characters long.","attack_vector":"Network reachability to the iRMC's Redfish or web interface. Exploitation depends on the username in play hitting the 16-character boundary, which an attacker can arrange when they control account creation or can guess an existing account name of that length.","remediation":"Firmware flash of the iRMC to 1.37S or later - a per-node out-of-band BMC update. A config-only interim step that genuinely helps: audit BMC account names and eliminate any that are exactly 16 characters, which removes the triggering condition without waiting for a flash window. Fujitsu's PSIRT publishes a readable PDF advisory, which puts them ahead of most of the ODM vendors in this database on advisory accessibility.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-65002","https://security.ts.fujitsu.com/ProductSecurity/content/FsasTech-PSIRT-FTI-ISS-2025-082610-Security-Notice.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-7342","cve":"CVE-2025-7342","aliases":[],"title":"Kubernetes Image Builder: Nutanix/OVA Windows images use default credentials unless overridden","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes Image Builder","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"Nutanix/OVA Windows images use default credentials unless overridden","attack_vector":"Unauthenticated network reaching an affected node","remediation":"Rebuild affected Windows node images","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-13460","cve":"CVE-2026-13460","aliases":[],"title":"IBM Storage Scale GUI (hardcoded inter-node token): A hardcoded token in the Storage Scale GUI source, used for inter-node cluster communication and REST access.…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale GUI (hardcoded inter-node token)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A hardcoded token in the Storage Scale GUI source, used for inter-node cluster communication and REST access. A hardcoded credential in a storage cluster's management path is the same shipped-secret problem as default BMC passwords: it is identical on every deployment, it is in a source tree anyone can read, and rotating it is not something the product expects you to do.","attack_vector":"Anyone who can reach the Storage Scale GUI/REST endpoint and knows the token — which, once published, is everyone.","remediation":"Upgrade Storage Scale past the affected 5.2.3.x / 6.0.x levels. GUI-layer upgrade plus service restart; the filesystem stays up. Immediately restrict the GUI/REST endpoint to a management network — a firewall change, applied live, that matters more than the patch timing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13460"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-15977","cve":"CVE-2026-15977","aliases":[],"title":"SGLang (`/server_info`): Endpoint returns API keys and SSL keyfile paths","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/server_info`)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Endpoint returns API keys and SSL keyfile paths","attack_vector":"Network to the serving port with only `--admin-*` partly configured","remediation":"Upgrade; rotate any keys exposed on a previously reachable instance","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15977"],"status":"curated"},{"id":"CVE-2026-15978","cve":"CVE-2026-15978","aliases":[],"title":"SGLang (weight exfiltration): Two endpoints allow a remote attacker to pull model weights when no API key is set","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (weight exfiltration)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Two endpoints allow a remote attacker to pull model weights when no API key is set","attack_vector":"Unauthenticated network to the serving port","remediation":"Upgrade; for a neocloud hosting customer models, this is direct tenant-IP exfiltration","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15978"],"status":"curated"},{"id":"CVE-2026-1669","cve":"CVE-2026-1669","aliases":[],"title":"Keras (HDF5 external links): Arbitrary local file read during model load","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (HDF5 external links)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary local file read during model load","attack_vector":"Customer-supplied HDF5 model — reads provider or co-tenant files visible to the process","remediation":"Upgrade; better, disallow HDF5. Note the process' service-account tokens and mounted secrets are in scope","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1669"],"status":"curated"},{"id":"CVE-2026-23440","cve":"CVE-2026-23440","aliases":["net/mlx5e race condition during IPsec ESN update"],"title":"Linux kernel mlx5_core IPsec full offload (ESN handling): The extended-sequence-number wrap event can be processed twice because the arm flag is re-set too late while…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core IPsec full offload (ESN handling)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The extended-sequence-number wrap event can be processed twice because the arm flag is re-set too late while the xfrm state lock is dropped and retaken. The driver then programs invalid ESN state, anti-replay fails, and all IPsec traffic on that SA halts. Operationally this is a stall of encrypted east-west traffic, not a memory-safety bug - but on an encrypted fabric it looks like a hard partition.","attack_vector":"Remote and unauthenticated in effect: an IPsec peer driving enough traffic to wrap the sequence number reaches the race. No credentials on the host.","remediation":"Upgrade the host kernel to 7.0 or a stable backport (6.6.130, 6.12.78, 6.18.20, 6.19.10). Rolling reboot of IPsec-offload nodes. Interim: shorten SA rekey intervals so ESN wrap is not reached, a config change on the IKE daemon.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23440","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-23440.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24146","cve":"CVE-2026-24146","aliases":[],"title":"Triton Inference Server: DoS via memory exhaustion on malformed input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via memory exhaustion on malformed input","attack_vector":"Any client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images; add request limits","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24146","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-789"]},{"id":"CVE-2026-24158","cve":"CVE-2026-24158","aliases":[],"title":"NVIDIA Triton Inference Server: A large compressed payload - a decompression bomb against the HTTP endpoint - exhausts the server. Trivially…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A large compressed payload - a decompression bomb against the HTTP endpoint - exhausts the server. Trivially weaponised by any client. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5790. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24158","https://github.com/NVIDIA/product-security/tree/main/2026/5790"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-789"]},{"id":"CVE-2026-24163","cve":"CVE-2026-24163","aliases":[],"title":"TensorRT-LLM: RCE via insecure config deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"RCE via insecure config deserialization","attack_vector":"Malicious model config","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24163","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24173","cve":"CVE-2026-24173","aliases":[],"title":"NVIDIA Triton Inference Server: A malformed request crashes the server outright. On a shared inference tier this is a noisy-neighbour weapon…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A malformed request crashes the server outright. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5816. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24173","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"]},{"id":"CVE-2026-24174","cve":"CVE-2026-24174","aliases":[],"title":"NVIDIA Triton Inference Server: A second malformed-request crash path. On a shared inference tier this is a noisy-neighbour weapon: one…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A second malformed-request crash path. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5816. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24174","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-681"]},{"id":"CVE-2026-24175","cve":"CVE-2026-24175","aliases":[],"title":"NVIDIA Triton Inference Server: A malformed request header crashes the server before the body is even parsed. On a shared inference tier this…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"A malformed request header crashes the server before the body is even parsed. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5816. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24175","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-248"]},{"id":"CVE-2026-24184","cve":"CVE-2026-24184","aliases":[],"title":"NVIDIA Cumulus Linux - LLDP daemon: Crafted LLDP frames overflow a buffer in the LLDP daemon, reaching code execution on the switch from an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Cumulus Linux - LLDP daemon","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Crafted LLDP frames overflow a buffer in the LLDP daemon, reaching code execution on the switch from an unauthenticated attacker on an adjacent network. LLDP is processed from every connected port by default, so any compromised host in the rack can reach it.","attack_vector":"Adjacent network, unauthenticated. A single compromised server NIC sends LLDP frames to the leaf it is plugged into. This is a host-to-fabric escalation path.","remediation":"Upgrade Cumulus Linux per bulletin 5817. Cost: switch reboot and link flap, sequenced leaf-by-leaf. Interim control: disable LLDP receive on host-facing ports if your tooling does not depend on it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24184","https://github.com/NVIDIA/product-security/tree/main/2026/5817"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24209","cve":"CVE-2026-24209","aliases":[],"title":"Triton Inference Server: Arbitrary file access via unsafe path operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary file access via unsafe path operations","attack_vector":"Network client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24209","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-22"]},{"id":"CVE-2026-24210","cve":"CVE-2026-24210","aliases":[],"title":"Triton Inference Server: DoS / memory corruption (integer overflow in buffer allocation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS / memory corruption (integer overflow in buffer allocation)","attack_vector":"Network client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24210","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"]},{"id":"CVE-2026-24212","cve":"CVE-2026-24212","aliases":[],"title":"Isaac Launchable: Info disclosure (unencrypted sensitive data in transit)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Isaac Launchable","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Info disclosure (unencrypted sensitive data in transit)","attack_vector":"Network observer","remediation":"Upgrade the package; enforce TLS","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24212","https://github.com/NVIDIA/product-security/tree/main/2026/5830"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-319"]},{"id":"CVE-2026-24255","cve":"CVE-2026-24255","aliases":[],"title":"NVIDIA Dynamo: MITM via improper TLS certificate verification","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"MITM via improper TLS certificate verification","attack_vector":"Network attacker on the inter-node path","remediation":"Bump Dynamo; redeploy; verify TLS config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24255","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:H/A:N","cwe":["CWE-1023"]},{"id":"CVE-2026-24264","cve":"CVE-2026-24264","aliases":[],"title":"Triton Inference Server: DoS via improper exception handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via improper exception handling","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24264","https://github.com/NVIDIA/product-security/tree/main/2026/5848"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-409"]},{"id":"CVE-2026-32666","cve":"CVE-2026-32666","aliases":["CVE-2026-24060","CVE-2026-25086","ICSA-26-078-08"],"title":"Automated Logic WebCTRL / i-Vu server and controllers, BACnet transport trust: This is the vendor formally conceding the structural problem: WebCTRL inherits BACnet's total absence of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Automated Logic WebCTRL / i-Vu server and controllers, BACnet transport trust","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"This is the vendor formally conceding the structural problem: WebCTRL inherits BACnet's total absence of network-layer authentication and adds no validation of its own, so an attacker on the BACnet segment can spoof packets to the WebCTRL server or to any Automated Logic controller and have them processed as legitimate. The companion issues are just as bad in practice - service traffic including file contents crosses the wire unencrypted and is trivially readable with Wireshark's BACnet dissector, and under some conditions an attacker can bind the WebCTRL service port and impersonate the server without ever injecting code. Operationally this means write commands to setpoints, fan speeds and schedules can be forged, and the operator has no cryptographic way to tell a real command from a fake one. In a GPU hall the practical consequence is that thermal control is only as trustworthy as the physical and VLAN boundary around the BACnet network, which for most operators is much weaker than they assume.","attack_vector":"Any host that can put packets on the BACnet/IP segment. No credentials exist to steal because none are used. This includes the mechanical contractor's laptop, a compromised BMS workstation, a rogue device in an unlocked mechanical room, and - in a leased colo - anything the landlord has on the shared building network. Also reachable through a BACnet router that bridges IP to MS/TP.","remediation":"Partly unpatchable by design. The plaintext and port-binding issues have fixes in current WebCTRL releases and you should take them, but the underlying spoofing exposure is a protocol property: BACnet/IP has no authentication and Automated Logic explicitly says it does not add validation. The only real control is segmentation and physical security of the BACnet segment - dedicated VLAN, no routing to tenant/corporate/internet, port security or 802.1X on the switch ports that carry it, and locked mechanical rooms. Where the vendor supports BACnet Secure Connect (BACnet/SC), moving to it is the actual fix and it is a controller-by-controller project with a contractor, so budget it as a capital line rather than a patch. Leased site: name this in the contract - require BACnet segment isolation with evidence, because you cannot fix someone else's protocol.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-078-08","https://nvd.nist.gov/vuln/detail/CVE-2026-32666","https://nvd.nist.gov/vuln/detail/CVE-2026-24060","https://www.automatedlogic.com/en/company/security-commitment/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2026-32981","cve":"CVE-2026-32981","aliases":[],"title":"Ray Dashboard: Path traversal in the dashboard static-file handler (port 8265)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray Dashboard","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Path traversal in the dashboard static-file handler (port 8265)","attack_vector":"Unauthenticated network","remediation":"Upgrade past 2.8.1","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-32981"],"status":"curated"},{"id":"CVE-2026-33554","cve":"CVE-2026-33554","aliases":[],"title":"GNU FreeIPMI's ipmi-oem tool before version 1.6.17: The direction of trust is what makes this operator-relevant: the vulnerability is in the management station…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GNU FreeIPMI's ipmi-oem tool before version 1.6.17","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The direction of trust is what makes this operator-relevant: the vulnerability is in the management station, not the managed node. A malicious or compromised BMC sends a crafted response and corrupts memory in the tool running on your management host. In a fleet that means one compromised BMC can attack the machine that polls every other BMC - so the blast radius of a single bad node extends to your entire out-of-band control plane, including whatever credentials that host holds for the rest of the fleet. Exploitable buffer overflows in the parsing of IPMI response messages, i.e. the client trusts what the BMC sends back.","attack_vector":"Requires the operator's own tooling to talk to a hostile IPMI responder. That happens when a BMC has already been compromised, when a node of unknown provenance is brought into the fleet, or when an attacker on the management VLAN can spoof or intercept IPMI responses - which the weak IPMI session security elsewhere in this list makes plausible.","remediation":"Package update of FreeIPMI to 1.6.17 or later on every management host, monitoring collector and provisioning box that runs ipmi-oem. This is an ordinary distribution package update, so rollout is cheap - no firmware flash, no node reboot, no maintenance window - which makes it one of the few items here you can just fix. The architectural follow-up worth doing: run fleet IPMI polling from a host that holds no other credentials, so a compromise of the poller does not hand over the whole management plane.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33554","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/33xxx/CVE-2026-33554.json"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2026-33697","cve":"CVE-2026-33697","aliases":["CoCoS aTLS relay"],"title":"Cocos AI - attested TLS (aTLS) on AMD SEV-SNP and Intel TDX: MULTI-TENANT ISOLATION: the attested-TLS implementation is vulnerable to a relay attack in which an attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cocos AI - attested TLS (aTLS) on AMD SEV-SNP and Intel TDX","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the attested-TLS implementation is vulnerable to a relay attack in which an attacker extracts the ephemeral TLS private key used during the intra-handshake attestation, letting them relay a genuine attestation report from a real confidential VM while terminating the session themselves. Affects both the SEV-SNP and TDX deployment targets. The point of confidential AI is that the client can prove it is talking to a specific attested enclave; a relay attack means that proof is worth nothing while the transport still looks correct.","attack_vector":"A network attacker positioned between the client and the confidential workload. No credentials needed - the attack is against the binding between the attestation and the TLS session, not against either one alone.","remediation":"Upgrade Cocos past v0.8.2. The broader operator lesson is worth more than the patch: any attested-TLS design that does not cryptographically bind the attestation report to the exact TLS key in use is relayable, so if you build or buy a confidential-AI stack, make binding an explicit acceptance criterion. Cost: application upgrade and redeploy; no firmware or driver change.","references":["https://github.com/ultravioletrs/cocos/security/advisories/GHSA-vfgg-mvxx-mgg7","https://nvd.nist.gov/vuln/detail/CVE-2026-33697"],"status":"curated"},{"id":"CVE-2026-34486","cve":"CVE-2026-34486","aliases":[],"title":"Apache Tomcat: Missing encryption of sensitive data introduced by the CVE-2026-29146 fix","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Tomcat","year":"2026","cvss_score":7.5,"severity":"high","kev":true,"impact":"[KEV] Missing encryption of sensitive data introduced by the CVE-2026-29146 fix","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade Tomcat under every Java control-plane/console service","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-34486"],"status":"curated"},{"id":"CVE-2026-41523","cve":"CVE-2026-41523","aliases":[],"title":"vLLM (activation function loading): Assert-based security check bypass, unauthenticated","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (activation function loading)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Assert-based security check bypass, unauthenticated","attack_vector":"Unauthenticated network","remediation":"Upgrade to 0.22.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41523"],"status":"curated"},{"id":"CVE-2026-46114","cve":"CVE-2026-46114","aliases":["RDMA/rxe ATOMIC_WRITE zero-length disclosure","Soft-RoCE skb tailroom leak"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c: TENANT ISOLATION: atomic_write_reply() dereferences 8 bytes at the payload address unconditionally, while the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) responder, drivers/infiniband/sw/rxe/rxe_resp.c","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"TENANT ISOLATION: atomic_write_reply() dereferences 8 bytes at the payload address unconditionally, while the rkey check accepted an ATOMIC_WRITE with a RETH length of zero. A remote initiator sending a zero-length ATOMIC_WRITE makes the responder read 8 bytes past the logical end of the packet into the socket buffer's tailroom and then write those bytes into the attacker's own memory region. That is a clean remote read primitive: four bytes of uninitialised kernel memory disclosed to the attacker per probe, repeatable at will, ideal for defeating KASLR or harvesting kernel pointers before a heavier exploit. The IB specification defines ATOMIC_WRITE as exactly 8 bytes, so anything else was always protocol-invalid.","attack_vector":"Remote initiator on an rxe connection sets the RETH length to 0 on an ATOMIC_WRITE and reads back the responder's reply, which now contains kernel tailroom bytes. Repeat to accumulate a memory-disclosure oracle. No local privilege and no authentication beyond reaching the Soft-RoCE endpoint.","remediation":"Host reboot / kernel upgrade. Same family control as the other rxe findings: blacklist and unload rdma_rxe where Soft-RoCE is not intentionally deployed, which is a config change with no downtime and removes this along with the rest. Where rxe is in use, upgrade the kernel and reboot on rolling drain, and firewall UDP/4791 to known peers in the meantime.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46114.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46114"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46133","cve":"CVE-2026-46133","aliases":["RDMA/rxe unknown opcode ICRC out-of-bounds","Soft-RoCE rxe_opcode table gap"],"title":"Linux kernel - RDMA/rxe (Soft-RoCE) ICRC processing, drivers/infiniband/sw/rxe: FABRIC DOS: The follow-up to CVE-2026-46043, and the reason to check you have both. The rxe_opcode[] table…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel - RDMA/rxe (Soft-RoCE) ICRC processing, drivers/infiniband/sw/rxe","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: The follow-up to CVE-2026-46043, and the reason to check you have both. The rxe_opcode[] table has 256 entries but only defined IB opcodes are populated; an undefined opcode such as 0xff reads a zero-initialised entry, so the length check added by the previous fix degenerates to a comparison against zero and stops constraining the packet length. rxe_icrc_hdr() then computes length minus the BTH size, which underflows, producing an out-of-bounds read. One unauthenticated UDP packet still panics the node. The defect predates the earlier fix and reaches back to the original Soft-RoCE driver, so any kernel with rxe loaded has carried it for years.","attack_vector":"A single UDP datagram to port 4791 carrying an opcode not defined in the IB specification. No connection state, no authentication, no prior contact with the target. Trivially scriptable and trivially fleet-wide.","remediation":"Host reboot / kernel upgrade to a kernel carrying this fix specifically - patching only CVE-2026-46043 leaves you exposed. As with the rest of the rxe family, the zero-cost control is to blacklist and unload rdma_rxe where Soft-RoCE is not in use (config change, no downtime), which is the right answer on essentially every production GPU node with real RDMA hardware. Where rxe must stay, restrict UDP/4791 at the host firewall and switch ACLs to known peers while the kernel rollout proceeds.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-46133.json","https://nvd.nist.gov/vuln/detail/CVE-2026-46133"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-47471","cve":"CVE-2026-47471","aliases":[],"title":"TensorRT-LLM: RCE potential (OOB write in tensor manipulation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"RCE potential (OOB write in tensor manipulation)","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47471","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-122"]},{"id":"CVE-2026-47476","cve":"CVE-2026-47476","aliases":[],"title":"Triton Inference Server: DoS via resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via resource exhaustion","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47476","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"]},{"id":"CVE-2026-47477","cve":"CVE-2026-47477","aliases":[],"title":"Triton Inference Server: DoS via stack overflow in recursive processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via stack overflow in recursive processing","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47477","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-121"]},{"id":"CVE-2026-47478","cve":"CVE-2026-47478","aliases":[],"title":"Triton Inference Server: DoS via memory leak (improper resource cleanup)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via memory leak (improper resource cleanup)","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47478","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-910"]},{"id":"CVE-2026-47479","cve":"CVE-2026-47479","aliases":[],"title":"Triton Inference Server: DoS via allocation-request resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via allocation-request resource exhaustion","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47479","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"]},{"id":"CVE-2026-47480","cve":"CVE-2026-47480","aliases":[],"title":"Triton Inference Server: DoS via unchecked exception","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via unchecked exception","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47480","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-248"]},{"id":"CVE-2026-47482","cve":"CVE-2026-47482","aliases":[],"title":"Triton Inference Server: DoS via resource leak in error paths","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"DoS via resource leak in error paths","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47482","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-401"]},{"id":"CVE-2026-47612","cve":"CVE-2026-47612","aliases":[],"title":"NVIDIA Dynamo: Arbitrary file access via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Arbitrary file access via path traversal","attack_vector":"Network client","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47612","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-22"]},{"id":"CVE-2026-47613","cve":"CVE-2026-47613","aliases":[],"title":"NVIDIA Dynamo: SSRF via URL handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF via URL handling","attack_vector":"Tenant-supplied URL","remediation":"Bump Dynamo; add egress network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47613","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"]},{"id":"CVE-2026-47614","cve":"CVE-2026-47614","aliases":[],"title":"NVIDIA Dynamo: SSRF in remote resource loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in remote resource loading","attack_vector":"Tenant-supplied model URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47614","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"]},{"id":"CVE-2026-47615","cve":"CVE-2026-47615","aliases":[],"title":"NVIDIA Dynamo: SSRF in model fetching","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in model fetching","attack_vector":"Tenant-supplied model URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47615","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"]},{"id":"CVE-2026-47616","cve":"CVE-2026-47616","aliases":[],"title":"NVIDIA Dynamo: SSRF in data retrieval","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in data retrieval","attack_vector":"Tenant-supplied URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47616","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"]},{"id":"CVE-2026-47617","cve":"CVE-2026-47617","aliases":[],"title":"NVIDIA Dynamo: SSRF in API communication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in API communication","attack_vector":"Tenant-supplied URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47617","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"]},{"id":"CVE-2026-47618","cve":"CVE-2026-47618","aliases":[],"title":"NVIDIA Dynamo: SSRF in external service communication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"SSRF in external service communication","attack_vector":"Tenant-supplied URI","remediation":"Bump Dynamo; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47618","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-918"]},{"id":"CVE-2026-47628","cve":"CVE-2026-47628","aliases":[],"title":"NVIDIA Triton Inference Server: Allocation of resources without limits lets an unauthenticated caller exhaust the server. On a shared…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Allocation of resources without limits lets an unauthenticated caller exhaust the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47628","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-770"]},{"id":"CVE-2026-47629","cve":"CVE-2026-47629","aliases":[],"title":"NVIDIA Triton Inference Server: Improper input validation crashes the server. On a shared inference tier this is a noisy-neighbour weapon…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Improper input validation crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47629","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"]},{"id":"CVE-2026-50031","cve":"CVE-2026-50031","aliases":[],"title":"GNU FreeIPMI ipmi-oem before 1.6.18: Same shape as its predecessor and the same fleet consequence: a hostile BMC response corrupts memory in the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GNU FreeIPMI ipmi-oem before 1.6.18","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Same shape as its predecessor and the same fleet consequence: a hostile BMC response corrupts memory in the tool on your management host. The operationally important detail is that this is the second round - an operator who updated to 1.6.17 believing FreeIPMI's OEM response parsing was fixed is still exposed. Treat the ipmi-oem response parser as untrusted code handling untrusted input until proven otherwise, rather than assuming a given release closed the class. A further set of response-message buffer overflows found after the 1.6.17 fix, meaning the first round of hardening did not cover the whole parser.","attack_vector":"Your management tooling querying a BMC that returns crafted responses - a compromised controller, a node brought in from an untrusted source, or an attacker able to interpose on IPMI traffic across the management VLAN.","remediation":"Package update to FreeIPMI 1.6.18 or later everywhere ipmi-oem runs. Cheap: distribution package update, no reboot, no firmware. Given that OEM extension parsing has now produced two rounds of overflows, the stronger move for most GPU operators is to stop using ipmi-oem for routine fleet polling at all - vendor-specific OEM IPMI commands are rarely load-bearing, and dropping them removes this parser from your control plane entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-50031","https://lists.gnu.org/archive/html/info-gnu/2026-06/msg00000.html","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/50xxx/CVE-2026-50031.json"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2026-57231","cve":"CVE-2026-57231","aliases":[],"title":"Podman: Image env var with a key and no value causes Podman to pass the host's value of that variable into the…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Image env var with a key and no value causes Podman to pass the host's value of that variable into the container; host secret leakage","attack_vector":"Malicious image","remediation":"Upgrade Podman to 5.8.4+; sanitise the environment of any host that runs untrusted images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-57231"],"status":"curated"},{"id":"CVE-2026-5757","cve":"CVE-2026-5757","aliases":[],"title":"Ollama (quantization engine): Unauthenticated remote information disclosure — reads and exfiltrates model data","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (quantization engine)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated remote information disclosure — reads and exfiltrates model data","attack_vector":"Unauthenticated network to the quantization endpoint","remediation":"Upgrade; direct cross-tenant model-weight exposure if Ollama is shared","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-5757"],"status":"curated"},{"id":"CVE-2026-62430","cve":"CVE-2026-62430","aliases":["XSA-503"],"title":"Xen (vRTC): Out-of-bounds read in vRTC emulation - hypervisor memory disclosure to a guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (vRTC)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Out-of-bounds read in vRTC emulation - hypervisor memory disclosure to a guest","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-503.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-65315","cve":"CVE-2026-65315","aliases":[],"title":"Ollama (GGUF metadata parser): Uncontrolled memory allocation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama (GGUF metadata parser)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Uncontrolled memory allocation → remote crash","attack_vector":"Customer-supplied GGUF","remediation":"Upgrade; GGUF parser hardening is still incomplete three years on","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65315"],"status":"curated"},{"id":"CVE-2026-69111","cve":"CVE-2026-69111","aliases":[],"title":"Milvus: Unauthenticated DoS terminating service components","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Milvus","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"Unauthenticated DoS terminating service components","attack_vector":"Unauthenticated network","remediation":"Upgrade past 2.6.22","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-69111"],"status":"curated"},{"id":"NCVD-2018-001-rocev2-lossless-ethernet-fabric","cve":null,"aliases":["PFC deadlock","PFC storm","head-of-line blocking","congestion spreading","Revisiting Network Support for RDMA (SIGCOMM 2018)"],"title":"RoCEv2 lossless Ethernet fabric - IEEE 802.1Qbb Priority Flow Control: FABRIC DOS: RoCE requires a lossless network, which in practice means Priority Flow Control, and PFC brings…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RoCEv2 lossless Ethernet fabric - IEEE 802.1Qbb Priority Flow Control","year":"2018","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: RoCE requires a lossless network, which in practice means Priority Flow Control, and PFC brings head-of-line blocking, congestion spreading, and occasional deadlocks as documented properties rather than as bugs. PFC pause frames propagate backwards hop by hop, so a single misbehaving or malicious endpoint that refuses to drain its receive queue can pause its upstream switch port, which pauses its upstream, and so on until a large fraction of an AI training fabric stops forwarding. A cyclic buffer dependency produces a deadlock that does not clear on its own. This is the highest-blast-radius item in this slice: one tenant NIC can wedge a whole rail of a GPU cluster and take every job on it down at once.","attack_vector":"An endpoint on the lossless priority class stops consuming traffic - by design, or through a driver hang, or via a hardware fault - and its NIC emits sustained PFC pause frames. Congestion spreads to unrelated flows sharing the same priority queue on intermediate switches, including flows between tenants who have nothing to do with the source. Deadlock arises when routing plus link failures create a cycle of buffer dependencies, and it persists until an operator intervenes. Triggering it requires no protocol violation whatsoever, which is why it also happens accidentally.","remediation":"Config change plus, on many platforms, a switch reload. Enable PFC watchdog on every switch (Cisco NX-OS, Arista EOS, NVIDIA Cumulus, SONiC all ship one) so a stuck queue is drained and the port error-disabled instead of pausing the fabric - this is a config push, but on some platforms the underlying buffer/QoS profile change needs a switch reload, so schedule per-rail. Keep the lossless class narrow (one priority, storage and RDMA separated), tune ECN/DCQCN thresholds to mark before PFC ever fires, and prefer deadlock-free routing (up-down or edge-disjoint) on Clos fabrics. Longer term, move to RoCE implementations that tolerate loss (the IRN design line, and DCQCN-plus-selective-retransmit NIC firmware) so PFC can be disabled entirely - firmware flash plus a driver upgrade fleet-wide, and a redesign of the QoS profile, so treat as a project.","references":["https://arxiv.org/abs/1806.08159","https://arxiv.org/abs/2207.10898"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2021-005-infiniband-roce-communication-ma","cve":null,"aliases":["ReDMArk DoS","RDMA connection-manager resource exhaustion","RNIC queue-pair exhaustion"],"title":"InfiniBand/RoCE Communication Manager (CM) and RNIC connection-state resources: FABRIC DOS: RNICs hold per-connection state in a fixed, small on-chip resource pool. ReDMArk demonstrated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand/RoCE Communication Manager (CM) and RNIC connection-state resources","year":"2021","cvss_score":7.5,"severity":"high","kev":false,"impact":"FABRIC DOS: RNICs hold per-connection state in a fixed, small on-chip resource pool. ReDMArk demonstrated that an unauthenticated peer can drive a victim RNIC to exhaustion simply by opening connections through the RDMA Connection Manager, or by sending malformed CM datagrams that leave half-open state behind. Once the pool is full the victim node can no longer accept new queue pairs, which on a GPU cluster means NCCL rendezvous fails, jobs hang at the allreduce, and the node effectively drops out of the scheduler while still appearing healthy to a ping-based health check.","attack_vector":"Any node that can reach the victim's CM service (UD QP1 on InfiniBand, UDP/4791 plus the CM port on RoCE) opens connections in a loop, or sends CM REQ messages and never completes the handshake. Because CM traffic is unauthenticated and is processed before any application-level identity exists, no tenant credentials are needed. In a shared cluster a single misbehaving or malicious container with RDMA device access is enough.","remediation":"No patch. Config change: rate-limit CM traffic at the switch or in the host's RDMA CM (per-source connection caps), enforce per-tenant P_Key partitions so a tenant can only reach nodes in its own job, and use SR-IOV VF resource limits so one VF cannot consume the PF's whole QP pool. Vendor firmware on newer ConnectX/BlueField parts adds per-VF resource quotas - that is a firmware flash plus a host reboot, rolling. Cheapest immediate mitigation is scheduler-side: do not co-schedule untrusted tenants onto nodes sharing an RNIC.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://netsec.ethz.ch/publications/papers/sec21summer-redmark.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2024-006-nvidia-gpudirect-rdma-gpu-bar1-w","cve":null,"aliases":["GPUDirect RDMA BAR1 exposure","nvidia-peermem","nvidia_p2p_get_pages","GPU memory registered as an RDMA memory region"],"title":"NVIDIA GPUDirect RDMA - GPU BAR1 window peer-mapped to the RNIC, nvidia-peermem / nvidia_p2p_get_pages: TENANT ISOLATION: GPUDirect RDMA works by exposing GPU memory through a PCIe BAR window…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPUDirect RDMA - GPU BAR1 window peer-mapped to the RNIC, nvidia-peermem / nvidia_p2p_get_pages","year":"2024","cvss_score":7.5,"severity":"high","kev":false,"impact":"TENANT ISOLATION: GPUDirect RDMA works by exposing GPU memory through a PCIe BAR window (BAR1) so a third-party device - the NIC - can DMA into it directly, with nvidia_p2p_get_pages() pinning the range and handing the peer device physical addresses. The practical consequence for a cluster operator is that every weakness in RDMA memory protection now applies to HBM, not just host DRAM: an rkey guessed or injected per the ReDMArk findings resolves to a region backed by GPU memory, and a successful remote read returns model weights or KV cache straight out of the GPU with no CPU involvement and no host-side trace. The GPU's own MMU is not in this path - the RNIC's rkey check is the entire access control. NVIDIA's documentation notes the 64-bit p2pToken is randomised specifically to keep an adversary from guessing it, which is an acknowledgement that guessability is the threat model here.","attack_vector":"Requires an attacker able to reach the victim's RNIC on the fabric and to guess or inject against the memory region that covers GPU memory - the ReDMArk and NeVerMore primitives. A second, local path: any process that can obtain a peer-mapping token or that shares the RDMA device can register GPU memory it should not reach, since the peer-memory client trusts the calling context. Stale mappings are a third: the nvidia_p2p callback must free the page table on deallocation, and a mapping that outlives its buffer leaves the NIC pointed at memory that has been handed to another context.","remediation":"Driver upgrade plus config change. Keep the NVIDIA GPU driver, nvidia-peermem (or the in-tree dma-buf peer path on recent kernels), and the RDMA stack on current versions - a driver upgrade requiring a host reboot on GPU nodes, so batch it with a scheduled drain. Config: never share an RNIC or an IB device node between tenants when GPUDirect is enabled, register the narrowest possible GPU regions with the least permission, use a per-tenant protection domain, and prefer dma-buf-based registration with explicit lifetime over legacy peermem pinning. Where GPUDirect is not actually needed for a workload, disabling it removes the exposure entirely at a bandwidth cost. Combine with the fabric partitioning in the InfiniBand and RoCE entries - the RDMA-layer fixes are what actually protect the GPU memory.","references":["https://docs.nvidia.com/cuda/gpudirect-rdma/","https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger","https://arxiv.org/abs/2202.08080"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2025-006-nvidia-bluefield-3-rnic-dpu-micr","cve":null,"aliases":["Noisy Neighbor","RDMA state saturation attack","RDMA pipeline saturation attack","BlueField-3 resource exhaustion","arXiv:2510.12629"],"title":"NVIDIA BlueField-3 RNIC/DPU - microarchitectural state and pipeline resources under containerized multi-tenancy: TENANT ISOLATION: Measured on current NVIDIA BlueField-3 hardware, two families of…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA BlueField-3 RNIC/DPU - microarchitectural state and pipeline resources under containerized multi-tenancy","year":"2025","cvss_score":7.5,"severity":"high","kev":false,"impact":"TENANT ISOLATION: Measured on current NVIDIA BlueField-3 hardware, two families of resource-exhaustion attack from a co-resident container cost the victim container up to 93.9% of its bandwidth, a 1,117x latency increase, and a 115% rise in NIC cache misses. Pipeline saturation additionally produces severe link-level congestion with strong amplification - small verb requests translating into disproportionately large resource consumption, so the attacker spends almost nothing. For a GPU cloud this is a direct, reproducible way for one tenant to make another tenant's distributed training run at a fraction of its purchased throughput, on the exact hardware generation being deployed for AI fabrics today.","attack_vector":"A container with ordinary RDMA access issues verb patterns designed either to saturate RNIC connection/translation state (state saturation) or to congest the NIC's processing pipeline (pipeline saturation). No kernel exploit, no privilege escalation, no fabric spoofing - just legitimate verbs at an adversarial shape. Because the amplification is high, an attacker constrained by its own bandwidth quota can still consume the shared NIC.","remediation":"No patch available; NVIDIA has not published an advisory for this class. Config change: enforce per-container caps on queue pairs, memory regions, and completion queues via the RDMA cgroup controller and per-VF resource limits (applied at container/VF creation, no reboot), and monitor per-container verb rates so abuse is at least attributable. The paper proposes HT-Verbs, a telemetry-driven throttling framework that needs no hardware change but is research code, not a product. Structural fix is a dedicated VF or NIC per tenant. Treat RDMA bandwidth SLAs on shared BlueField-3 ports as unenforceable until you have measured your own configuration.","references":["https://arxiv.org/abs/2510.12629","https://www.usenix.org/conference/nsdi23/presentation/kong"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-001-openbmc-bmcweb-http-1-1-expect-1","cve":null,"aliases":["GHSA-p3gc-68x5-g9w3","CONN-F2"],"title":"OpenBMC bmcweb HTTP/1.1 Expect: 100-continue handling: bmcweb applies a 4 KB body limit to unauthenticated requests, and the Expect: 100-continue path returns…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb HTTP/1.1 Expect: 100-continue handling","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"bmcweb applies a 4 KB body limit to unauthenticated requests, and the Expect: 100-continue path returns before that limit is applied. An unauthenticated client streams roughly 10 MB into BMC memory per connection - a 2,560x amplification over the intended cap - and drives the BMC out of memory. Redfish, the web UI, KVM and serial-over-LAN all die with the process. What makes this one worth tracking specifically: it has a published GitHub advisory but no CVE was ever assigned, so it does not appear in NVD, OSV or the vendor scanners most operators rely on. If your firmware risk process is 'match SBOM against NVD', you will not see it. Intel firmware shipped in January 2026 still contained it.","attack_vector":"Unauthenticated HTTP(S) to bmcweb on the BMC management interface, using a stock HTTP client. No credentials, no host access.","remediation":"Fixed on 2026-04-21 in bmcweb commit 0b2049b0, released in bmcweb 3.0.0; affected versions are 2.18.0 and earlier. Reaching a deployed fleet needs a BMC firmware flash - per node, out-of-band, and dependent on your ODM rebasing to bmcweb 3.0.0, which for most server vendors will lag by quarters. Config-only mitigation and the realistic near-term answer: put the BMC HTTPS port behind an ACL so only management jump hosts can reach it, and do not expose the web interface beyond the management VLAN. Because there is no CVE, add it to your firmware acceptance checklist by bmcweb version rather than expecting a scanner to flag it.","references":["https://github.com/openbmc/bmcweb/security/advisories/GHSA-p3gc-68x5-g9w3","https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-002-openbmc-bmcweb-http-2-content-le","cve":null,"aliases":["H2-F2"],"title":"OpenBMC bmcweb HTTP/2 Content-Length handling: bmcweb passes the client-supplied Content-Length straight into a reserve() call. One HEADERS frame declaring…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb HTTP/2 Content-Length handling","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"bmcweb passes the client-supplied Content-Length straight into a reserve() call. One HEADERS frame declaring a nonsense length like 999999999999 throws bad_alloc and takes down the single-threaded event loop - the entire management interface, from one frame, with no body sent and no credentials. Notable for the disclosure hygiene as much as the bug: it was patched the same day as the Expect-header issue but received no advisory of its own and no CVE, so nothing downstream records that affected firmware is vulnerable. An operator auditing by advisory count will undercount bmcweb's exposure.","attack_vector":"Unauthenticated HTTP/2 to bmcweb on the management interface. A single frame is enough; no payload, no session.","remediation":"Fixed 2026-04-21 in bmcweb commit 62526bb0, same release train as the Expect fix (bmcweb 3.0.0). Since it shares a release with the advisory-carrying bug, one BMC firmware flash covers both - per node, out-of-band, ODM-rebase-lagged. Until then, config-only: restrict who can reach the BMC HTTPS port, and disable h2 if your tooling permits. Verify by bmcweb version, not by scanner output, since no CVE exists to match against.","references":["https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-003-openbmc-bmcweb-http-2-body-buffe","cve":null,"aliases":["H2-F1"],"title":"OpenBMC bmcweb HTTP/2 body buffering (HttpBody::reader, nghttp2 flow control): The HTTP/2 code path in bmcweb appends incoming DATA frames with no size check at all, and nghttp2's default…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb HTTP/2 body buffering (HttpBody::reader, nghttp2 flow control)","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"The HTTP/2 code path in bmcweb appends incoming DATA frames with no size check at all, and nghttp2's default window auto-replenishment means the data flows before authentication happens. Same outcome as the Expect bug but worse, because it bypasses the body limit entirely rather than merely leaking past it, and because h2 is negotiated by default over ALPN so it is the path a modern client takes automatically. The researcher rates it the most severe of the four findings. Reported as unfixed on master at time of disclosure, with no advisory and no CVE - so this is a known, public, unpatched pre-auth DoS in the daemon that owns every out-of-band control path on your fleet.","attack_vector":"Unauthenticated HTTP/2 over TLS to bmcweb on the management interface. ALPN negotiates h2 by default, so no unusual client is needed.","remediation":"No fix as of disclosure; proposed Gerrit patches were posted to the OpenBMC mailing list. This is a config-only situation until upstream lands a fix and your ODM rebases - so ACL the BMC HTTPS port to management jump hosts only, and if your tooling can live on HTTP/1.1, disabling h2 in ALPN on the BMC removes the path. Assume no firmware you can buy today contains a fix. When one exists it arrives as a per-node out-of-band BMC firmware flash with the usual ODM lag.","references":["https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-015-nvme-over-fabrics-discovery-cont","cve":null,"aliases":["NVMe-oF discovery controller unauthenticated by default","NVMe/TCP no auth by default","nvmet_host_allowed discovery bypass"],"title":"NVMe-over-Fabrics discovery controller - Linux kernel nvmet (drivers/nvme/target/discovery.c), NVMe/TCP and NVMe/RDMA: TENANT ISOLATION: The NVMe-oF discovery controller answers any peer that can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe-over-Fabrics discovery controller - Linux kernel nvmet (drivers/nvme/target/discovery.c), NVMe/TCP and NVMe/RDMA","year":"2026","cvss_score":7.5,"severity":"high","kev":false,"impact":"TENANT ISOLATION: The NVMe-oF discovery controller answers any peer that can reach the port, by design. In Linux nvmet this is explicit in the source: nvmet_host_allowed() returns true unconditionally for the discovery subsystem, so every Get Log Page against the discovery controller is served pre-authentication over TCP, RDMA, or FC. The response enumerates every subsystem NQN, transport type, and address on that target. For a neocloud running disaggregated NVMe behind a tenant-reachable fabric, that hands an attacker the complete map of which customer's storage lives where and what NQN to spoof to reach it - the reconnaissance step that makes NQN spoofing practical. NVMe/TCP additionally ships with no authentication and no transport encryption unless in-band DH-HMAC-CHAP and TLS are explicitly configured, which is not the default in most deployments.","attack_vector":"The attacker points nvme discover at any reachable target IP on port 4420 (or the RDMA equivalent) with an arbitrary Host NQN and receives the full discovery log page. No credentials, no prior connection, no exploit. From there they connect to a named subsystem asserting a permitted Host NQN. The same unauthenticated surface is what makes memory-disclosure bugs in the discovery path (see CVE-2026-64320) remotely reachable pre-auth.","remediation":"Config change, no reboot, do it now: enable in-band DH-HMAC-CHAP on the target and require it for I/O subsystems; enable TLS for NVMe/TCP where the kernel version supports it; bind the discovery controller to a management interface unreachable from tenant networks rather than to the tenant storage fabric; and use per-subsystem allowed_hosts lists rather than the permissive default. All of this is nvmet configfs or SPDK RPC and applies at runtime, but initiators must be reconfigured in lockstep, so plan a rolling reattach per tenant. Firewall NVMe/TCP 4420 to known initiator addresses as an immediate stopgap.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2026/CVE-2026-64320.json","https://arxiv.org/abs/2202.08080","https://arxiv.org/abs/1903.09355"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2020-13817","cve":"CVE-2020-13817","aliases":[],"title":"ntpd (transmit timestamp prediction): A remote attacker who can predict transmit timestamps can crash ntpd or, worse, change the system time.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ntpd (transmit timestamp prediction)","year":"2020","cvss_score":7.4,"severity":"high","kev":false,"impact":"A remote attacker who can predict transmit timestamps can crash ntpd or, worse, change the system time. Changing the time on a cluster node is a more interesting attack than crashing it: certificates become valid or invalid, log timestamps stop lining up with reality, and scheduled jobs fire at the wrong moment.","attack_vector":"Remote attacker able to predict the client's transmit timestamps for outgoing packets.","remediation":"Upgrade ntp to 4.2.8p14 or later, or move to chrony/NTS. Package upgrade and service restart. Monitor for step changes in system time as a detection control — most fleets do not alert on this and should.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13817"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2021-1228","cve":"CVE-2021-1228","aliases":[],"title":"Cisco Nexus 9000 in ACI mode (fabric infrastructure VLAN): TENANT ISOLATION: a device plugged into a normal front-panel port can talk its way onto the ACI…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus 9000 in ACI mode (fabric infrastructure VLAN)","year":"2021","cvss_score":7.4,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a device plugged into a normal front-panel port can talk its way onto the ACI infrastructure VLAN — the fabric's own control plane. From there an attacker sees and can influence the fabric's internal signalling rather than one tenant's EPG. In a multi-tenant ACI build this is the boundary that separates 'a tenant' from 'the fabric operator'.","attack_vector":"Unauthenticated, adjacent — physical or logical access to a leaf front-panel port. Any tenant with a bare-metal node, or anyone who can plug into a rack, is in position.","remediation":"ACI software upgrade across the APIC cluster and the leaf/spine switches — a staged fabric upgrade, not a single reload, and Cisco's recommended sequence takes hours on a large pod. Interim mitigation is strict port-level admission control and disabling unused ports, both live config changes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1228","https://nvd.nist.gov/vuln/detail/CVE-2019-1890"],"status":"curated"},{"id":"CVE-2021-3493","cve":"CVE-2021-3493","aliases":[],"title":"Linux kernel (overlayfs, Ubuntu patch): OverlayFS file-capability privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (overlayfs, Ubuntu patch)","year":"2021","cvss_score":7.4,"severity":"high","kev":true,"impact":"OverlayFS file-capability privilege escalation; trivially weaponised inside unprivileged user namespaces [KEV]","attack_vector":"Any tenant process in a container","remediation":"Ubuntu-specific patch delta. Livepatchable on Ubuntu Pro Livepatch; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-3493"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-21656","cve":"CVE-2022-21656","aliases":[],"title":"Envoy: Type-confusion in default certificate validation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"Type-confusion in default certificate validation","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21656"],"status":"curated"},{"id":"CVE-2022-21953","cve":"CVE-2022-21953","aliases":[],"title":"Rancher: Missing authorization allows an authenticated user to create a shell pod with kubectl access in the local…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"Missing authorization allows an authenticated user to create a shell pod with kubectl access in the local cluster; management-plane takeover","attack_vector":"Any authenticated Rancher user","remediation":"Upgrade Rancher","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21953"],"status":"curated"},{"id":"CVE-2022-28734","cve":"CVE-2022-28734","aliases":[],"title":"GRUB2 (HTTP chunked transfer): Out-of-bounds write handling chunked HTTP responses during HTTP boot. An attacker who can answer or MITM the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HTTP chunked transfer)","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"Out-of-bounds write handling chunked HTTP responses during HTTP boot. An attacker who can answer or MITM the HTTP boot request executes code in the bootloader on a node that has not booted an OS yet - the cleanest possible foothold on a bare-metal fleet.","attack_vector":"Network position on the provisioning path during HTTP boot: rogue DHCP, DNS spoofing, or a compromised provisioning server.","remediation":"grub2 package update + reboot, and replace the served netboot binary. If you HTTP-boot, move to HTTPS boot with a pinned CA and isolate the provisioning VLAN - the transport was the actual weakness here.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28734","https://access.redhat.com/security/cve/CVE-2022-28734"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-31671","cve":"CVE-2022-31671","aliases":[],"title":"Harbor registry: P2P preheat execution logs readable/updatable by any authenticated user via job ID enumeration","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Harbor registry","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"P2P preheat execution logs readable/updatable by any authenticated user via job ID enumeration","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; job logs often contain registry credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31671"],"status":"curated"},{"id":"CVE-2022-47630","cve":"CVE-2022-47630","aliases":["TFV-10","TF-A X.509 parser out-of-bounds read"],"title":"Arm Trusted Firmware-A through v2.8, X.509 certificate parser used by Trusted Board Boot (get_ext, auth_nvctr): The code that validates the secure boot certificate chain reads out of bounds on a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-A through v2.8, X.509 certificate parser used by Trusted Board Boot (get_ext, auth_nvctr)","year":"2022","cvss_score":7.4,"severity":"high","kev":false,"impact":"The code that validates the secure boot certificate chain reads out of bounds on a malformed certificate. So the component whose entire job is to decide whether firmware is trustworthy can be driven off the rails by the untrusted input it is inspecting - dangerous read side effects and leakage of microarchitectural state. It undermines the chain of trust at the exact point where a bare-metal operator is trying to prove to the next tenant that the box is clean.","attack_vector":"An attacker able to place a crafted certificate in the boot chain: control of the firmware image or the firmware-update path. On bare-metal GPU rental, a prior tenant with flash write access. Not remote.","remediation":"Upgrade TF-A past v2.8 with the TFV-10 fix and have the OEM re-issue the platform firmware. Flash + reboot + drain per node. There is no runtime mitigation - the parser runs before anything you control. Pair it with the operational control that actually helps: measure and attest boot firmware between tenants instead of trusting the parser.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-10.html","https://nvd.nist.gov/vuln/detail/CVE-2022-47630"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0286","cve":"CVE-2023-0286","aliases":[],"title":"OpenSSL: X.400 address type confusion in X.509 GeneralName - crash or possible memory disclosure during cert…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSL","year":"2023","cvss_score":7.4,"severity":"high","kev":false,"impact":"X.400 address type confusion in X.509 GeneralName - crash or possible memory disclosure during cert verification","attack_vector":"Unauthenticated network","remediation":"Package update + service restart","references":["https://access.redhat.com/security/cve/CVE-2023-0286"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2023-1829","cve":"CVE-2023-1829","aliases":[],"title":"Linux kernel (net/sched tcindex): Use-after-free in the tcindex traffic-control filter - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched tcindex)","year":"2023","cvss_score":7.4,"severity":"high","kev":false,"impact":"Use-after-free in the tcindex traffic-control filter - local root","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot. Blacklist `cls_tcindex`","references":["https://access.redhat.com/security/cve/CVE-2023-1829"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-40548","cve":"CVE-2023-40548","aliases":["shim 15.8 batch"],"title":"shim (verify_sbat_section): Integer overflow leading to heap overflow while verifying the SBAT section on 32-bit systems. The irony…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (verify_sbat_section)","year":"2023","cvss_score":7.4,"severity":"high","kev":false,"impact":"Integer overflow leading to heap overflow while verifying the SBAT section on 32-bit systems. The irony matters operationally: SBAT is the mechanism meant to revoke vulnerable bootloaders, and parsing it is itself exploitable.","attack_vector":"Crafted binary presented to shim on a 32-bit EFI implementation. Rare on server-class GPU hardware, common on 32-bit-EFI edge and embedded boxes.","remediation":"shim package update + reboot. Low priority on x86-64 server fleets, real on any 32-bit-EFI hardware you still operate.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40548","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46817","cve":"CVE-2024-46817","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.4,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Stop amdgpu_dm initialize when stream nums greater than 6","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-20111","cve":"CVE-2025-20111","aliases":[],"title":"Cisco Nexus 3000/9000 (health monitoring diagnostics): FABRIC DOS: the health monitoring diagnostics subsystem on Nexus 3000 and 9000 switches in standalone NX-OS…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus 3000/9000 (health monitoring diagnostics)","year":"2025","cvss_score":7.4,"severity":"high","kev":false,"impact":"FABRIC DOS: the health monitoring diagnostics subsystem on Nexus 3000 and 9000 switches in standalone NX-OS mode can be driven into a denial of service by an unauthenticated attacker. The irony is worth noting for operators — the subsystem whose job is to tell you the switch is healthy is itself the thing that takes it down.","attack_vector":"Unauthenticated, remote — traffic reaching the affected diagnostics handling on the switch.","remediation":"NX-OS upgrade plus reload per switch, staged across MLAG/ECMP pairs. Restrict management-plane reachability with CoPP in the meantime — a live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20111"],"status":"curated"},{"id":"CVE-2026-18556","cve":"CVE-2026-18556","aliases":[],"title":"N-able N-central: Authentication bypass using an alternate path or channel on the RMM server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"N-able N-central","year":"2026","cvss_score":7.4,"severity":"high","kev":true,"impact":"[KEV] Authentication bypass using an alternate path or channel on the RMM server","attack_vector":"Network (remote)","remediation":"Control-plane: patch; the RMM has agent reach into every managed host","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-18556"],"status":"curated"},{"id":"CVE-2026-47473","cve":"CVE-2026-47473","aliases":[],"title":"TensorRT-LLM: Memory corruption (improper array bounds checking)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.4,"severity":"high","kev":false,"impact":"Memory corruption (improper array bounds checking)","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47473","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-123"]},{"id":"CVE-2018-3615","cve":"CVE-2018-3615","aliases":["Foreshadow","L1TF-SGX"],"title":"Intel SGX (L1 terminal fault on enclave pages): MULTI-TENANT ISOLATION: Speculative execution lets code outside an enclave read the enclave's data out of the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX (L1 terminal fault on enclave pages)","year":"2018","cvss_score":7.3,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Speculative execution lets code outside an enclave read the enclave's data out of the L1 data cache, breaking the confidentiality guarantee SGX exists to provide. On a confidential-compute offering this is the whole product failing: the host, or a co-resident tenant with the right sibling-thread placement, reads enclave secrets including sealing and attestation material.","attack_vector":"Local code execution on the same physical core as the enclave. In a cloud that includes a co-tenant on a sibling hyperthread, which is why the mitigation story involves disabling SMT.","remediation":"Microcode update plus OS/hypervisor mitigations. Intel microcode for this can be late-loaded at boot by the OS without waiting for an OEM BIOS release, so the practical path is: update the microcode package, reboot, and disable SMT (or enforce core scheduling) on nodes that run untrusted co-tenants. Also re-attest and re-provision any enclave secrets that existed on unpatched hardware - the microcode fix does not un-leak them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3615","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00161.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2020-8595","cve":"CVE-2020-8595","aliases":[],"title":"Istio: Authentication Policy exact-path matching allows unauthorized access to HTTP paths","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2020","cvss_score":7.3,"severity":"high","kev":false,"impact":"Authentication Policy exact-path matching allows unauthorized access to HTTP paths","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade plus sidecar restart; sidecar restart means restarting tenant pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8595"],"status":"curated"},{"id":"CVE-2021-1074","cve":"CVE-2021-1074","aliases":[],"title":"NVIDIA GPU Display Driver for Windows, installer: A race between the installer's integrity check and execution lets an unprivileged local user swap in a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver for Windows, installer","year":"2021","cvss_score":7.3,"severity":"high","kev":false,"impact":"A race between the installer's integrity check and execution lets an unprivileged local user swap in a malicious file and get it run by an administrator's install. The window is short, which is a mitigation and not a fix - driver rollouts are scheduled and repeatable, so an attacker sitting on the box knows exactly when to try.","attack_vector":"An unprivileged local user present on the host while an administrator runs the driver installer.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1074"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1075","cve":"CVE-2021-1075","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Use of a dangling pointer in the escape handler, rated code-execution capable. Guest-agnostic local kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":7.3,"severity":"high","kev":false,"impact":"Use of a dangling pointer in the escape handler, rated code-execution capable. Guest-agnostic local kernel bug - unprivileged code to SYSTEM on the GPU host.","attack_vector":"Any local user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1075"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1085","cve":"CVE-2021-1085","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the guest can write to a shared memory location after the host has validated its…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.3,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the guest can write to a shared memory location after the host has validated its contents - a double-fetch that NVIDIA rates as escalation-capable. The tenant passes the host's checks with clean data, then substitutes their own. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM able to race the host's validation window.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1085"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-21797","cve":"CVE-2022-21797","aliases":[],"title":"joblib: Arbitrary code execution via `eval` on the `pre_dispatch` flag in `Parallel()`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"joblib","year":"2022","cvss_score":7.3,"severity":"high","kev":false,"impact":"Arbitrary code execution via `eval` on the `pre_dispatch` flag in `Parallel()`","attack_vector":"Tenant code passing attacker-influenced strings","remediation":"Upgrade joblib >= 1.2.0 in base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21797"],"status":"curated"},{"id":"CVE-2022-23817","cve":"CVE-2022-23817","aliases":[],"title":"AMD Secure Processor Secure OS - memory buffer checking: MULTI-TENANT ISOLATION: A malicious trusted application can read and write the ASP Secure OS's kernel virtual…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor Secure OS - memory buffer checking","year":"2022","cvss_score":7.3,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A malicious trusted application can read and write the ASP Secure OS's kernel virtual address space because buffer bounds are not checked at the TA boundary. That is full privilege escalation inside the secure processor - the attacker is now the security engine, not a client of it.","attack_vector":"Local, requires the ability to load a malicious trusted application into the ASP (signing-key compromise or a legitimately signed but attacker-controlled TA).","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23817","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31097","cve":"CVE-2022-31097","aliases":[],"title":"Grafana: Stored XSS via Unified Alerting","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2022","cvss_score":7.3,"severity":"high","kev":false,"impact":"Stored XSS via Unified Alerting -> privilege escalation to Grafana admin","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31097"],"status":"curated"},{"id":"CVE-2023-29057","cve":"CVE-2023-29057","aliases":["LEN-118321"],"title":"Lenovo XClarity Controller (XCC) - LDAP/AD authorization: When XCC is configured to authenticate against Active Directory, a user's local XCC account permissions…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - LDAP/AD authorization","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"When XCC is configured to authenticate against Active Directory, a user's local XCC account permissions silently override the permissions their directory group grants - and the local ones can be higher. The result is privilege escalation that your identity provider cannot see: you revoke someone's admin rights in AD, XCC keeps honouring the stale local grant, and they retain out-of-band control of the node. For an operator this is a offboarding and least-privilege failure more than an exploit, which makes it easy to miss - nothing looks broken, and the directory tells you the access is gone.","attack_vector":"A valid XCC user in a deployment where LDAP/AD is configured for authentication and authorisation and the user also has a local XCC account. No exploit code required - the misbehaviour is in the authorisation logic itself.","remediation":"Flash XCC to the version listed for your model in LEN-118321 - out-of-band, per-node, no host reboot and no drain. Alongside the flash, do the config work that actually closes the gap: enumerate local XCC accounts on every node and delete the ones that shadow directory identities, because patching the precedence logic does not remove local accounts that are already there.","references":["https://support.lenovo.com/us/en/product_security/LEN-118321","https://nvd.nist.gov/vuln/detail/CVE-2023-29057"],"status":"curated"},{"id":"CVE-2023-29164","cve":"CVE-2023-29164","aliases":["INTEL-SA-00990","CVE-2023-49144","CVE-2023-35123","INTEL-SA-01078","CVE-2025-20097"],"title":"BMC firmware for Intel Server Boards S2600WF / S2600ST / S2600BP before 02.01.0017 and M50CYP, and OpenBMC firmware for Intel server platforms (Eagle Stream / Birch Stream): Improper access…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BMC firmware for Intel Server Boards S2600WF / S2600ST / S2600BP before 02.01.0017 and M50CYP, and OpenBMC firmware…","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Improper access control in the BMC firmware on Intel's own server boards, with companion out-of-bounds read and uncaught-exception issues in the OpenBMC builds Intel ships for current Eagle Stream and Birch Stream platforms. BMC compromise is node ownership below the OS: virtual media, power control, KVM, and the firmware update path into BIOS and ME. It survives host reimage by construction and carries into the next tenant, and mass BMC access is a hall-level power-off capability. The OpenBMC entries matter because many neocloud whitebox designs run Intel's OpenBMC tree rather than a commercial BMC stack, and those builds are updated far less often.","attack_vector":"Access to the BMC - network access on the out-of-band management path for the network-facing issues, privileged local access for the information-disclosure path, and the host KCS interface where it has not been disabled.","remediation":"BMC firmware update to 02.01.0017 or later on the S2600 boards, and to the fixed OpenBMC branch (egs-1.15-0 / bhs-0.27 or later) on the current platforms - from Intel or from the ODM that built the board (Quanta, Wiwynn, Supermicro, Inventec). BMC updates do not need a host reboot, so this is cheap to roll relative to BIOS work. If you run whitebox nodes on Intel's OpenBMC, you own the update cadence yourself - there is no OEM pushing it to you, and that is the real finding here. Keep the BMC network isolated, credentials unique per node, and KCS disabled.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29164","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00990.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01078.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31008","cve":"CVE-2023-31008","aliases":[],"title":"DGX H100 BMC (IPMI): Privesc","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Privesc","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-20"]},{"id":"CVE-2023-31016","cve":"CVE-2023-31016","aliases":[],"title":"GPU Display Driver (Windows): Arbitrary code exec via uncontrolled search path","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Arbitrary code exec via uncontrolled search path","attack_vector":"Local low-priv user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"]},{"id":"CVE-2023-34239","cve":"CVE-2023-34239","aliases":[],"title":"Gradio: Lack of path filtering","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"Lack of path filtering → arbitrary file read from the host","attack_vector":"Unauthenticated network to the demo","remediation":"Upgrade; a Gradio demo runs with the tenant's full container filesystem access","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34239"],"status":"curated"},{"id":"CVE-2023-54098","cve":"CVE-2023-54098","aliases":[],"title":"Linux i915 GVT-g mediated GPU virtualisation (debugfs teardown): MULTI-TENANT ISOLATION: Companion to the vGPU debugfs cleanup bug: GVT-g destroys debugfs state without…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux i915 GVT-g mediated GPU virtualisation (debugfs teardown)","year":"2023","cvss_score":7.3,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Companion to the vGPU debugfs cleanup bug: GVT-g destroys debugfs state without checking the DRM minor's debugfs root is still valid, crashing the host on teardown. Again on the vGPU lifecycle path, which is the tenant boundary on GVT-g deployments.","attack_vector":"Whoever can drive vGPU create/destroy on the host.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54098","https://git.kernel.org/stable/c/ae9a61511736cc71a99f01e8b7b90f6fb6128ed8"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-21960","cve":"CVE-2024-21960","aliases":[],"title":"AMD Optimizing CPU Libraries (AOCL) - installation directory permissions: AOCL installs with permissive directory permissions, so a low-privileged user can replace library files that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Optimizing CPU Libraries (AOCL) - installation directory permissions","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"AOCL installs with permissive directory permissions, so a low-privileged user can replace library files that privileged processes later load - straightforward privilege escalation to arbitrary code execution. Worth flagging for AI operators specifically: AOCL (BLIS, libFLAME, AOCL-LibM) is exactly what gets installed on AMD nodes to accelerate the CPU side of an ML pipeline, so it is likely present on your hosts and likely loaded by jobs running as someone else.","attack_vector":"Local, low-privileged user who can write into the AOCL installation directory. If tenants share a node and AOCL lives somewhere world-writable, one tenant poisons the next tenant's math library.","remediation":"Update AOCL and correct the directory permissions - this is a filesystem ACL fix plus a package update, no reboot and no firmware. Audit the permissions on every math and ML library directory on shared nodes while you are there; the same mistake recurs across vendor-supplied HPC packages.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21960","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-22415","cve":"CVE-2024-22415","aliases":[],"title":"jupyter-lsp: Unauthenticated file read/write through the LSP extension","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"jupyter-lsp","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Unauthenticated file read/write through the LSP extension","attack_vector":"Network attacker reaching JupyterLab","remediation":"Upgrade; the extension bypasses notebook auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22415"],"status":"curated"},{"id":"CVE-2024-25621","cve":"CVE-2024-25621","aliases":[],"title":"containerd: Overly broad default permissions on containerd-managed directories","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Overly broad default permissions on containerd-managed directories","attack_vector":"Any local process or partially-escaped container on the node","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25621"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-36292","cve":"CVE-2024-36292","aliases":[],"title":"Intel Data Center GPU Flex Series - Windows driver: Improper buffer restrictions in the Flex Series Windows driver let an authenticated user knock the GPU out.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center GPU Flex Series - Windows driver","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Improper buffer restrictions in the Flex Series Windows driver let an authenticated user knock the GPU out. Flex Series is the media/VDI-oriented datacenter part, so the affected hosts are typically multi-session - meaning any logged-in user, not just an administrator.","attack_vector":"Local, authenticated, on a Windows host running the Flex Series driver before 31.0.101.4314.","remediation":"Update the Flex Series Windows driver to 31.0.101.4314 or later. Cost: Windows driver replacement means a node reboot; drain sessions first.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-36292"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36339","cve":"CVE-2024-36339","aliases":[],"title":"AMD Optimizing CPU Libraries (AOCL) - DLL hijacking: A DLL search-order hijack in AOCL lets an attacker get their library loaded by a privileged process, reaching…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Optimizing CPU Libraries (AOCL) - DLL hijacking","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"A DLL search-order hijack in AOCL lets an attacker get their library loaded by a privileged process, reaching arbitrary code execution. Same practical exposure as the AOCL permissions issue: the CPU math libraries underneath your AMD-node ML stack become a privilege-escalation vector.","attack_vector":"Local, requires the ability to place a library where the search order will find it first. Windows-oriented, though the search-path pattern has Linux analogues via LD_LIBRARY_PATH on badly configured hosts.","remediation":"Update AOCL. Check that no shared-node job can influence library search paths for privileged processes - on Linux that means auditing LD_LIBRARY_PATH handling in your job launcher and any setuid tooling. Package update, no reboot, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36339","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-38552","cve":"CVE-2024-38552","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix potential index out of bounds in color transformation function","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38552","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-44933","cve":"CVE-2024-44933","aliases":[],"title":"Linux bnxt_en driver (bnxt_fill_hw_rss_tbl): Memory out-of-bounds in the RSS indirection-table path of the Broadcom NIC driver, giving kernel panic…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (bnxt_fill_hw_rss_tbl)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Memory out-of-bounds in the RSS indirection-table path of the Broadcom NIC driver, giving kernel panic, denial of service, or an information leak out of kernel memory. The information-leak arm is the one that matters for multi-tenancy — kernel memory on a shared host contains other workloads' data.","attack_vector":"Reachable through ring-reservation and RSS configuration paths in the driver; local host privilege or specific traffic conditions depending on the trigger.","remediation":"Kernel/driver upgrade and host reboot. Rolling across the fleet during normal maintenance; no firmware flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44933"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45333","cve":"CVE-2024-45333","aliases":[],"title":"Intel Data Center GPU Flex Series - Windows driver: A further improper access control in the Flex Series Windows driver allowing an authenticated local user to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center GPU Flex Series - Windows driver","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"A further improper access control in the Flex Series Windows driver allowing an authenticated local user to cause denial of service.","attack_vector":"Local, authenticated, on a Windows host running the driver before 31.0.101.4314.","remediation":"Update to 31.0.101.4314 or later - the same release that fixes CVE-2024-36292, so patch both in one window. Cost: node reboot.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-45333"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46811","cve":"CVE-2024-46811","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index may exceed array range within fpu_update_bw_bounding_box","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46811","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50117","cve":"CVE-2024-50117","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amd): A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amd)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd: Guard against bad data for ATIF ACPI method","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50117","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53104","cve":"CVE-2024-53104","aliases":[],"title":"Linux kernel (uvcvideo): Out-of-bounds write parsing UVC_VS_UNDEFINED frames - exploited in the wild","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (uvcvideo)","year":"2024","cvss_score":7.3,"severity":"high","kev":true,"impact":"Out-of-bounds write parsing UVC_VS_UNDEFINED frames - exploited in the wild [KEV]","attack_vector":"Local user with USB/device access","remediation":"Livepatchable; otherwise drain + reboot. Near-zero exposure on headless GPU servers - deprioritise behind the netfilter set, but it will still show on every compliance scan","references":["https://access.redhat.com/security/cve/CVE-2024-53104"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53108","cve":"CVE-2024-53108","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Adjust VSDB parser for replay feature","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53108","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57804","cve":"CVE-2024-57804","aliases":["CVE-2024-57807"],"title":"Linux kernel mpi3mr driver (Broadcom tri-mode 9600-series HBA/RAID) and megaraid_sas driver: Rapidly toggling PHY enable/disable through the SAS transport sysfs interface corrupts the controller's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mpi3mr driver (Broadcom tri-mode 9600-series HBA/RAID) and megaraid_sas driver","year":"2024","cvss_score":7.3,"severity":"high","kev":false,"impact":"Rapidly toggling PHY enable/disable through the SAS transport sysfs interface corrupts the controller's persistent and current configuration pages on Broadcom tri-mode 9600-series adapters. The corruption is in configuration state the controller keeps, not just in kernel memory - so a host-side operation writes bad persistent state into the storage controller, which is precisely the boundary a tenant handoff is supposed to reset. The megaraid_sas sibling is a lock-ordering deadlock between reset and scan paths that hangs the SCSI host. Both are availability and integrity problems at the controller layer: a node whose controller config pages are corrupt may enumerate drives differently or fail to bring arrays up after the next reboot, and that failure surfaces on the next tenant, not the one who caused it.","attack_vector":"Local root on the bare-metal host, through the SAS transport sysfs PHY controls. On a bare-metal GPU rental this is the tenant themselves - they legitimately have root and the sysfs interface is not namespaced.","remediation":"Kernel/driver update on the host and a reboot. The deeper point for operators is that this is a class you cannot patch away: a bare-metal tenant with root has a large, mostly unaudited surface of sysfs and ioctl interfaces that write persistent state into the storage controller. Between tenants, do not just reimage - re-read and, if your controller tooling supports it, restore the controller configuration to a known-good baseline (storcli/StorCLI config restore or the equivalent), and verify controller firmware version and config page integrity as part of the handoff checklist.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git","https://nvd.nist.gov/vuln/detail/CVE-2024-57804","https://nvd.nist.gov/vuln/detail/CVE-2024-57807"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-10164","cve":"CVE-2025-10164","aliases":[],"title":"SGLang (`/update_weights_from_tensor`): Unsafe deserialization of the `serialized_named_tensors` argument","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"SGLang (`/update_weights_from_tensor`)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Unsafe deserialization of the `serialized_named_tensors` argument","attack_vector":"Network to the SGLang HTTP API","remediation":"Upgrade past 0.4.6; the weight-update endpoint must not be tenant-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-10164"],"status":"curated","fleet":{"ubiquity":"Very common - SGLang is the main vLLM alternative for high-throughput LLM serving on neoclouds; affects 0.4.6 through <0.5.4","remediation_pain":"`daemon-restart` - upgrade to 0.5.4+ and roll every serving replica","pain_class":"daemon-restart","why_fleet_wide":"`update_weights_from_tensor` pickle-deserializes attacker input with no authentication, giving unauthenticated remote code execution on every SGLang serving process reachable on the network"}},{"id":"CVE-2025-22830","cve":"CVE-2025-22830","aliases":["AMI-SA-2025006"],"title":"AMI AptioV UEFI BIOS: A race condition in the BIOS that a skilled local attacker can drive to resource exhaustion, with AMI rating…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"A race condition in the BIOS that a skilled local attacker can drive to resource exhaustion, with AMI rating full confidentiality, integrity and availability impact and a subsequent-system impact as well. The operator-facing outcome is a node that fails to complete firmware operations or wedges in early boot - and given the CIA rating AMI assigns, a successfully-won race is a path to firmware-level compromise rather than just a hang. Winning a race in firmware is exactly the kind of bug that is unreliable in a lab and reliable at fleet scale, where an attacker gets thousands of boot attempts.","attack_vector":"Local access with high privileges and some user interaction, and the attack requires specific conditions to be present - AMI's CVSS v4 vector marks attack requirements as present, meaning the attacker needs the machine in a particular state. Realistically: root on the host plus the ability to trigger a reboot or a firmware operation, which any tenant of a bare-metal node has.","remediation":"BIOS update to AptioV_5.040 or later - firmware flash plus a host reboot per node, gated on your server vendor rebasing the AMI BKC for your SKU. No config-only workaround. Because the trigger involves reboots and firmware operations, one thing you can do without patching is restrict who can initiate BIOS updates and firmware operations out-of-band, and log every one of them.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-22830"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23241","cve":"CVE-2025-23241","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (an integer overflow)…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Kernel-mode flaw in the 800-series Ethernet Linux driver (an integer overflow) reachable by an authenticated user for privilege escalation. Part of the same 2025 batch - patch them as one unit rather than individually.","attack_vector":"Authenticated local user; tenants on nodes exposing VFs or RDMA devices.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 1.17.2 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23241","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01296.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23242","cve":"CVE-2025-23242","aliases":[],"title":"NVIDIA Riva: Unauthorized access to the speech service (insufficient access control)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Riva","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Unauthorized access to the speech service (insufficient access control)","attack_vector":"Network client of the Riva endpoint","remediation":"Upgrade Riva NIM containers; redeploy; put auth in front of the endpoint","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23242","https://github.com/NVIDIA/product-security/tree/main/2025/5625"],"status":"curated","fleet":{"ubiquity":"Common - shipped as a NIM-style microservice, deployed by neoclouds offering managed speech endpoints","remediation_pain":"`daemon-restart` (upgrade to Riva 2.19.0 and redeploy the service)","pain_class":"daemon-restart","why_fleet_wide":"Improper access control in the service auth layer, network-reachable with no user interaction: an unauthenticated caller escalates privileges into the hosting cloud environment and can read other tenants' data"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-284"]},{"id":"CVE-2025-23257","cve":"CVE-2025-23257","aliases":[],"title":"NVIDIA DOCA: Local privesc via insecure file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DOCA","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via insecure file permissions","attack_vector":"Local user on the DPU/host","remediation":"Upgrade DOCA packages; restart DPU services","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23257","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-732"]},{"id":"CVE-2025-23258","cve":"CVE-2025-23258","aliases":[],"title":"NVIDIA DOCA: Local privesc via insecure file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DOCA","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via insecure file permissions","attack_vector":"Local user on the DPU/host","remediation":"Upgrade DOCA packages","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23258","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-732"]},{"id":"CVE-2025-23277","cve":"CVE-2025-23277","aliases":[],"title":"GPU Display Driver: Access-control bypass","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Access-control bypass -> privesc","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23277","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","fleet":{"ubiquity":"Universal - Linux and Windows driver","remediation_pain":"`node-reboot`","pain_class":"node-reboot","why_fleet_wide":"Out-of-bounds memory access in kernel mode from a local tenant: cheap denial of service against a whole GPU node"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-284"]},{"id":"CVE-2025-23344","cve":"CVE-2025-23344","aliases":[],"title":"NVIDIA NVDebug tool: NVDebug allows an actor to run code on the platform host as a non-privileged user, reaching code execution…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVDebug tool","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"NVDebug allows an actor to run code on the platform host as a non-privileged user, reaching code execution and privilege escalation.","attack_vector":"Local, low privileges, user interaction. An unprivileged account on the platform host plus an operator running the tool.","remediation":"Update NVDebug per bulletin 5696. Cost: trivial tool replacement, no drain.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23344","https://github.com/NVIDIA/product-security/tree/main/2025/5696"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-29951","cve":"CVE-2025-29951","aliases":[],"title":"AMD Secure Processor (ASP) bootloader - buffer overflow: MULTI-TENANT ISOLATION: A buffer overflow in the ASP bootloader gives an attacker a memory overwrite in the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor (ASP) bootloader - buffer overflow","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A buffer overflow in the ASP bootloader gives an attacker a memory overwrite in the secure processor's boot path, ending in privilege escalation and arbitrary code execution below the OS. Anything the ASP protects on that node - SEV keys, fTPM, boot measurement - is then attacker-controlled, and the foothold persists across OS reinstalls.","attack_vector":"Local. Needs the ability to influence what the bootloader parses, i.e. SPI ROM write access or a subverted firmware update path.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Lock down SPI ROM writes and platform BIOS update authentication as the interim control; there is no OS-level mitigation for bootloader code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29951","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-30167","cve":"CVE-2025-30167","aliases":[],"title":"Jupyter Core (Windows): Config read from a shared writable path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Core (Windows)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Config read from a shared writable path → code execution as another user","attack_vector":"Co-tenant local user","remediation":"Upgrade to 5.8.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-30167"],"status":"curated"},{"id":"CVE-2025-31133","cve":"CVE-2025-31133","aliases":[],"title":"runc: Insufficient verification of masked-path bind mounts (/dev/null replaced by symlink) enables container escape…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Insufficient verification of masked-path bind mounts (/dev/null replaced by symlink) enables container escape to host root","attack_vector":"Any tenant workload / malicious image with a custom mount config","remediation":"Replace runc on all nodes; running containers stay vulnerable so drain required","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-31133"],"status":"curated","fleet":{"ubiquity":"Universal - all runc ≤1.2.7 / 1.3.2 / 1.4.0-rc.2","remediation_pain":"`node-drain` - binary replace plus recreation of every container; running containers are not retroactively protected","pain_class":"node-drain","why_fleet_wide":"Replacing `/dev/null` in the container with a symlink into host `/proc` makes critical host procfs entries mount writable, yielding container escape to host root from a tenant-controlled image"}},{"id":"CVE-2025-33181","cve":"CVE-2025-33181","aliases":[],"title":"Cumulus Linux / NVOS: Command injection (local)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux / NVOS","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Command injection (local)","attack_vector":"Local switch operator","remediation":"Upgrade Cumulus Linux / NVOS","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33181","https://github.com/NVIDIA/product-security/tree/main/2026/5722"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-77"]},{"id":"CVE-2025-33205","cve":"CVE-2025-33205","aliases":[],"title":"NVIDIA NeMo Framework: A predefined variable pulls in functionality from an untrusted control sphere, reaching code execution. In an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"A predefined variable pulls in functionality from an untrusted control sphere, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5729 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33205","https://github.com/NVIDIA/product-security/tree/main/2025/5729"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-829"]},{"id":"CVE-2025-33212","cve":"CVE-2025-33212","aliases":[],"title":"NVIDIA NeMo Framework: Loading a maliciously crafted model file bypasses the framework's control mechanisms and reaches code…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Loading a maliciously crafted model file bypasses the framework's control mechanisms and reaches code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5736 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33212","https://github.com/NVIDIA/product-security/tree/main/2025/5736"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2025-33228","cve":"CVE-2025-33228","aliases":[],"title":"CUDA Toolkit: Local privesc via command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via command injection","attack_vector":"Local user running toolkit binaries","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33228","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2025-33229","cve":"CVE-2025-33229","aliases":[],"title":"CUDA Toolkit: Code exec via untrusted library load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code exec via untrusted library load","attack_vector":"Local user / malicious image layer","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33229","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"]},{"id":"CVE-2025-33230","cve":"CVE-2025-33230","aliases":[],"title":"CUDA Toolkit: Local privesc via command injection","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Local privesc via command injection","attack_vector":"Local user","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33230","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2025-38508","cve":"CVE-2025-38508","aliases":[],"title":"Linux x86/sev - Secure TSC frequency calculation (TSC_FACTOR): Secure TSC is how an SEV-SNP guest gets a timebase it can trust rather than one the host can manipulate. The…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/sev - Secure TSC frequency calculation (TSC_FACTOR)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Secure TSC is how an SEV-SNP guest gets a timebase it can trust rather than one the host can manipulate. The guest's frequency calculation ignored TSC_FACTOR and so drifted from the real mean TSC frequency. A confidential guest whose clock is wrong makes wrong decisions about timeouts, certificate validity and rate limiting - and clock manipulation is a known lever for defeating the very side-channel defences that confidential computing depends on.","attack_vector":"Affects SEV-SNP guests using Secure TSC. Not an active attack so much as a broken trusted-time guarantee that a host-side adversary can lean on.","remediation":"Fixed in the **guest** kernel, not the host - the hardening lives in the SEV-ES/SNP guest's #VC handler and interrupt entry code. That inverts the usual rollout: you can patch every hypervisor you own and still be exposed, because the protection has to be in the tenant's own VM image. As an operator your job is to ship updated confidential-guest images (or tell tenants which minimum kernel to run) and, where you can, enforce it as an admission requirement. Each guest picks the fix up on its next boot; no host reboot, no firmware update. The fix is in the guest's x86/sev code, so confidential-VM images need updating; host patching alone does not deliver it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38508"],"status":"curated"},{"id":"CVE-2025-40334","cve":"CVE-2025-40334","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate userq buffer virtual address and size","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40334","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-52881","cve":"CVE-2025-52881","aliases":[],"title":"runc: Attacker misdirects runc writes to /proc via racing symlinks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Attacker misdirects runc writes to /proc via racing symlinks; can defeat LSM labelling and escape","attack_vector":"Any tenant workload","remediation":"Replace runc on all nodes; drain required","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52881"],"status":"curated","fleet":{"ubiquity":"Universal - same runc version range","remediation_pain":"`node-drain` - runc upgrade + container recreation across the whole fleet","pain_class":"node-drain","why_fleet_wide":"LSM (AppArmor/SELinux) bypass that makes arbitrary procfs writes easy, turning the other two into reliable host root; AWS, Alibaba and every distro shipped emergency runc rebuilds"}},{"id":"CVE-2025-54386","cve":"CVE-2025-54386","aliases":[],"title":"Traefik: Path traversal in the WASM plugin installation mechanism","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Path traversal in the WASM plugin installation mechanism","attack_vector":"Anyone who can supply a Traefik plugin","remediation":"Rolling Traefik upgrade; restrict plugin sources","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54386"],"status":"curated"},{"id":"CVE-2025-7647","cve":"CVE-2025-7647","aliases":[],"title":"llama-index-core: Predictable hardcoded cache directory","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama-index-core","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Predictable hardcoded cache directory → local hijack","attack_vector":"Co-tenant local user on a shared node","remediation":"Upgrade past 0.12.44; matters on shared bare-metal or shared `/tmp`","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7647"],"status":"curated"},{"id":"CVE-2025-9905","cve":"CVE-2025-9905","aliases":[],"title":"Keras (HDF5 path): Code execution from crafted `.h5`/`.hdf5` model despite safe mode","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (HDF5 path)","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code execution from crafted `.h5`/`.hdf5` model despite safe mode","attack_vector":"Customer-supplied legacy HDF5 model","remediation":"Block `.h5` ingest entirely; the legacy format has no safe loader","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-9905"],"status":"curated"},{"id":"CVE-2025-9906","cve":"CVE-2025-9906","aliases":[],"title":"Keras: Code execution from crafted `.keras` archive despite safe mode","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras","year":"2025","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code execution from crafted `.keras` archive despite safe mode","attack_vector":"Customer-supplied model file","remediation":"Upgrade; same class as above","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-9906"],"status":"curated"},{"id":"CVE-2026-13201","cve":"CVE-2026-13201","aliases":[],"title":"KubeVirt: safepath OpenAtNoFollow resolves via /proc/self/fd, defeating the symlink protection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"safepath OpenAtNoFollow resolves via /proc/self/fd, defeating the symlink protection","attack_vector":"Cluster user with virt-launcher access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13201"],"status":"curated"},{"id":"CVE-2026-15793","cve":"CVE-2026-15793","aliases":[],"title":"BuildKit: git.checkoutbundle=true on a malicious git source yields crafted command invocation on the build host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"git.checkoutbundle=true on a malicious git source yields crafted command invocation on the build host","attack_vector":"Anyone using the raw low-level build API","remediation":"Upgrade BuildKit; block raw LLB API for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15793"],"status":"curated"},{"id":"CVE-2026-2033","cve":"CVE-2026-2033","aliases":[],"title":"MLflow (artifact handler): Directory traversal","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (artifact handler)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Directory traversal → RCE via the artifact handler","attack_vector":"Remote attacker with artifact-write access","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-2033"],"status":"curated"},{"id":"CVE-2026-24156","cve":"CVE-2026-24156","aliases":[],"title":"NVIDIA DALI: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DALI","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious dataset/pipeline","remediation":"Bump DALI; rebuild data-loading images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24156","https://github.com/NVIDIA/product-security/tree/main/2026/5811"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24180","cve":"CVE-2026-24180","aliases":[],"title":"NVIDIA DALI: Code exec via OOB write on malformed tensor dims","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DALI","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code exec via OOB write on malformed tensor dims","attack_vector":"Malicious dataset","remediation":"Bump DALI; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24180","https://github.com/NVIDIA/product-security/tree/main/2026/5814"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-122"]},{"id":"CVE-2026-24181","cve":"CVE-2026-24181","aliases":[],"title":"NVIDIA DALI: Code exec via buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DALI","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Code exec via buffer overflow","attack_vector":"Malicious dataset","remediation":"Bump DALI; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24181","https://github.com/NVIDIA/product-security/tree/main/2026/5814"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-129"]},{"id":"CVE-2026-24206","cve":"CVE-2026-24206","aliases":[],"title":"Triton Inference Server: Weak authentication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Weak authentication -> unauthorized model access","attack_vector":"Network client of the endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24206","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-288"]},{"id":"CVE-2026-24229","cve":"CVE-2026-24229","aliases":[],"title":"TensorRT-LLM: Missing authentication in authorization checks","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Missing authentication in authorization checks","attack_vector":"Network client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24229","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:L","cwe":["CWE-306"]},{"id":"CVE-2026-40110","cve":"CVE-2026-40110","aliases":[],"title":"Jupyter Server (Origin validation): `re.match` used for Origin validation","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Jupyter Server (Origin validation)","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"`re.match` used for Origin validation → prefix-match bypass","attack_vector":"Malicious page visited by the notebook user","remediation":"Upgrade past 2.17.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-40110"],"status":"curated"},{"id":"CVE-2026-46680","cve":"CVE-2026-46680","aliases":[],"title":"containerd: Numeric User directive that fails 32-bit parsing is treated as a username, changing the effective UID","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Numeric User directive that fails 32-bit parsing is treated as a username, changing the effective UID","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46680"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-62290","cve":"CVE-2026-62290","aliases":[],"title":"cert-manager: Challenge resource handling flaw in cert-manager 1.18.0-1.19.5 and 1.20.x","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cert-manager","year":"2026","cvss_score":7.3,"severity":"high","kev":false,"impact":"Challenge resource handling flaw in cert-manager 1.18.0-1.19.5 and 1.20.x","attack_vector":"Cluster user with namespace access who can create Certificates","remediation":"Rolling cert-manager upgrade to 1.19.6/1.20.3+; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-62290"],"status":"curated"},{"id":"CVE-2018-7078","cve":"CVE-2018-7078","aliases":[],"title":"HPE iLO4 / iLO5: Remote code execution on the management controller","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO4 / iLO5","year":"2018","cvss_score":7.2,"severity":"high","kev":false,"impact":"Remote code execution on the management controller","attack_vector":"Network, authenticated","remediation":"iLO firmware update (iLO4 <2.60, iLO5 <1.30)","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7078"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-7105","cve":"CVE-2018-7105","aliases":[],"title":"HPE iLO3/4/5: Arbitrary code execution on the iLO","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO3/4/5","year":"2018","cvss_score":7.2,"severity":"high","kev":false,"impact":"Arbitrary code execution on the iLO","attack_vector":"Network","remediation":"iLO firmware update across three generations simultaneously — the version matrix is the operational cost","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-7105"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-19029","cve":"CVE-2019-19029","aliases":[],"title":"Harbor: SQL injection via user-groups","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":7.2,"severity":"high","kev":false,"impact":"SQL injection via user-groups","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19029"],"status":"curated"},{"id":"CVE-2019-19577","cve":"CVE-2019-19577","aliases":[],"title":"Xen on AMD - x86 HVM pagetable height update: MULTI-TENANT ISOLATION: AMD HVM guest OS users can trigger a data-structure access during a pagetable-height…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD - x86 HVM pagetable height update","year":"2019","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: AMD HVM guest OS users can trigger a data-structure access during a pagetable-height update, causing denial of service or possibly gaining privileges. Privilege escalation out of a guest into the hypervisor is the worst outcome available on a virtualised host - the attacker moves from one tenant's VM to controlling all of them.","attack_vector":"From inside an AMD HVM guest under Xen. Tenant-reachable.","remediation":"Fixed in Xen (XSA-310). Update the hypervisor and reboot the host; no firmware step. Affects Xen through 4.12.x.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19577","https://xenbits.xen.org/xsa/advisory-310.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-12967","cve":"CVE-2020-12967","aliases":["SEVerity","SEVurity"],"title":"AMD SEV / SEV-ES - missing nested page table protection: MULTI-TENANT ISOLATION: SEV and SEV-ES do not protect the nested page tables, so a malicious hypervisor can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES - missing nested page table protection","year":"2020","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: SEV and SEV-ES do not protect the nested page tables, so a malicious hypervisor can remap the guest's physical address space underneath it. The SEVurity and SEVerity research chains turned this into arbitrary code execution inside the encrypted guest: because encryption is keyed by physical address and the host controls the mapping, the host can move ciphertext blocks around to assemble instructions of its choosing inside the victim VM. The guest's memory is encrypted and the attacker still gets code execution in it.","attack_vector":"Requires a malicious administrator who has compromised the hypervisor. Affects SEV and SEV-ES; SEV-SNP's reverse-map table is the architectural answer to exactly this.","remediation":"**Not fixable on SEV or SEV-ES** - the missing protection is architectural, which is why AMD built SEV-SNP with the RMP. The remediation is a platform migration: run confidential workloads on SEV-SNP-capable EPYC (Milan 7003 and later) with SNP actually enabled, not on SEV or SEV-ES. If you are running SEV or SEV-ES today and telling customers their VMs are protected from you, that claim does not hold. Requires new hardware or at minimum a firmware/BIOS enablement pass plus reconfiguring your VMM to launch SNP guests.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12967","https://www.usenix.org/conference/woot20/presentation/werner","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2020-24637","cve":"CVE-2020-24637","aliases":[],"title":"ArubaOS GRUB2 implementation (secure boot): Two flaws in ArubaOS's GRUB2 implementation allow secure boot to be bypassed, leading to remote compromise.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ArubaOS GRUB2 implementation (secure boot)","year":"2020","cvss_score":7.2,"severity":"high","kev":false,"impact":"Two flaws in ArubaOS's GRUB2 implementation allow secure boot to be bypassed, leading to remote compromise. Secure boot on a network device is the control that stops a compromise from becoming permanent; bypassing it means an attacker's image survives reimaging and firmware updates. Same structural problem as the Cisco NX-OS image-verification bypass, on a different vendor.","attack_vector":"An attacker able to influence the boot chain — via administrative access or the documented remote path.","remediation":"ArubaOS upgrade plus reload; this is a bootloader-level fix so it must be applied per device and cannot be worked around in config. Treat any device suspected of pre-patch compromise as needing replacement or a verified out-of-band reflash rather than an in-place upgrade.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24637"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-26122","cve":"CVE-2020-26122","aliases":[],"title":"Inspur NF5266M5 through firmware 3.21.2 and other Inspur M5-generation servers: An attacker with administrative reach installs their own BMC firmware on Inspur server hardware, and from…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Inspur NF5266M5 through firmware 3.21.2 and other Inspur M5-generation servers","year":"2020","cvss_score":7.2,"severity":"high","kev":false,"impact":"An attacker with administrative reach installs their own BMC firmware on Inspur server hardware, and from that point the operator has no reliable way to determine what is running on the controller. The outcomes are the full BMC set - power, console, virtual media, boot control - plus an implant that survives host reimaging and node reprovisioning. Inspur M5 hardware shows up in a lot of Asia-Pacific cloud and HPC estates and in secondhand markets that neoclouds buy from, so this is a real inventory question for anyone assembling capacity opportunistically rather than from a single vendor. The BMC's firmware verification is weak and lacks the checks needed to establish that an image is genuine before flashing it.","attack_vector":"Administrative privilege on the BMC, reachable over the management network. Shared or default BMC credentials on secondhand hardware make this a realistic starting position rather than a theoretical one.","remediation":"Firmware update from Inspur - and obtaining it is the problem. Inspur's English security bulletin path returns a hard 404 and the site's homepage carries no PSIRT link, so there is no reachable vendor advisory channel for a non-Chinese-market operator. Given US export and entity-list constraints on Inspur, many operators will also find vendor support unavailable regardless. The practical position: treat Inspur BMC firmware as unverifiable, isolate these BMCs on a management VLAN with no tenant-reachable route, rotate credentials to per-node unique values, and for secondhand M5 hardware assume the BMC may already carry unknown firmware and plan accordingly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-26122","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2020/26xxx/CVE-2020-26122.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26311","cve":"CVE-2021-26311","aliases":["undeSErVed","SEV memory remapping"],"title":"AMD SEV / SEV-ES - guest address space rearrangement undetected by attestation: MULTI-TENANT ISOLATION: A malicious hypervisor can rearrange memory in the guest's address space and the SEV…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES - guest address space rearrangement undetected by attestation","year":"2021","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A malicious hypervisor can rearrange memory in the guest's address space and the SEV attestation mechanism does not notice, because attestation measures content rather than placement. That lets the host reorder the guest's own encrypted pages to build a different program out of the same measured bytes - the attestation report still validates while the VM executes something the tenant never wrote. This is the failure that makes 'the attestation passed' an insufficient answer on SEV/SEV-ES.","attack_vector":"Malicious hypervisor. Affects SEV and SEV-ES; SEV-SNP addresses it with the reverse-map table enforcing page ownership and mapping.","remediation":"**Not fixable on SEV or SEV-ES** - migrate confidential workloads to SEV-SNP, where the RMP enforces the guest-physical to system-physical mapping. Practically that means EPYC Milan (7003) or newer with SNP enabled in SBIOS and a VMM that launches SNP guests. If you are attesting SEV/SEV-ES guests today, understand that a passing attestation report does not establish the guest is running the code that was measured.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26311","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26344","cve":"CVE-2021-26344","aliases":[],"title":"AMD PSP1 Configuration Block (APCB) parsing: MULTI-TENANT ISOLATION: An out-of-bounds memory write while the platform processes the AMD PSP1 Configuration…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP1 Configuration Block (APCB) parsing","year":"2021","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds memory write while the platform processes the AMD PSP1 Configuration Block. Reaching it requires the ability to modify and re-sign the BIOS image, which is a high bar - but the payoff is memory corruption inside the secure processor's configuration path, i.e. control of the root of trust with a signature that validates.","attack_vector":"Local, and requires both BIOS image modification and the ability to sign the result. That combination points at a signing-key compromise or an insider in the firmware build pipeline rather than a runtime attacker.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. The signing prerequisite means your firmware supply chain is the real control here - who can sign a BIOS for your fleet, and how is that key held?","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26344","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-36784","cve":"CVE-2021-36784","aliases":[],"title":"Rancher: restricted-admin role escalates to full admin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2021","cvss_score":7.2,"severity":"high","kev":false,"impact":"restricted-admin role escalates to full admin","attack_vector":"A restricted-admin user","remediation":"Upgrade Rancher","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36784"],"status":"curated"},{"id":"CVE-2022-30704","cve":"CVE-2022-30704","aliases":["INTEL-SA-00717","CVE-2021-0187"],"title":"Intel TXT SINIT Authenticated Code Module for some Intel processors: Improper initialization in the SINIT ACM - the Intel-signed module that performs the TXT measured launch. A…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TXT SINIT Authenticated Code Module for some Intel processors","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper initialization in the SINIT ACM - the Intel-signed module that performs the TXT measured launch. A privileged local attacker can escalate through it, which means the dynamic root of trust that a TXT launch is supposed to establish can be subverted at the moment it is created. If you use TXT (directly, or via a launch control policy that gates whether a node may join a secure pool), an attacker can make a compromised node produce a passing launch. Nodes admitted to a trusted pool on that basis are not trustworthy, and the compromise is at firmware level so it crosses tenant handoff.","attack_vector":"A privileged local user on the node - local root or SMM-capable code.","remediation":"New SINIT ACM delivered inside a BIOS/platform-firmware update from the OEM (Dell, HPE, Supermicro, Lenovo, Gigabyte, Quanta), plus updating any standalone SINIT binary your tboot/launch stack loads. Host reboot and job drain. If you gate scheduling on TXT launch results, update the launch control policy hashes at the same time or the policy will fail closed and strand capacity.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30704","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00717.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33196","cve":"CVE-2022-33196","aliases":[],"title":"Intel Xeon memory controller configuration (with SGX): MULTI-TENANT ISOLATION: Memory controller configuration registers are left with permissions that let a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon memory controller configuration (with SGX)","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory controller configuration registers are left with permissions that let a privileged local user reach a privilege escalation on SGX-enabled Xeon platforms. Memory controller configuration is the layer that enforces where enclave memory lives, so getting it wrong is structurally worse than a normal ring-0 bug on a confidential-compute host.","attack_vector":"Privileged local access on the host.","remediation":"Platform BIOS/firmware update from the OEM, plus TCB recovery and re-attestation. BIOS means a per-node drain, a reboot, and waiting on OEM packaging - budget quarters, not weeks, on server boards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33196","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00738.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-41804","cve":"CVE-2022-41804","aliases":[],"title":"Intel Xeon processors (SGX/TDX error injection): MULTI-TENANT ISOLATION: Unauthorised error injection against SGX or TDX on affected Xeon parts gives a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon processors (SGX/TDX error injection)","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Unauthorised error injection against SGX or TDX on affected Xeon parts gives a privileged user an escalation. Fault injection against a TEE is the same family as Plundervolt: the host corrupts the protected computation until it gives up its secrets.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update - late-loadable at boot on most distributions, so this one does not wait on an OEM BIOS release - plus a TCB recovery and re-attestation. Reboot required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41804","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00837.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-42278","cve":"CVE-2022-42278","aliases":[],"title":"NVIDIA DGX BMC (AMI-derived management controller): The BMC's SPX REST API lets an authorised attacker read and write arbitrary locations in the IPMI server…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI-derived management controller)","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"The BMC's SPX REST API lets an authorised attacker read and write arbitrary locations in the IPMI server process memory, reaching code execution inside the BMC. A BMC compromise on a DGX gives an attacker power control, virtual media, serial console and a persistent foothold under the host OS on a node holding eight GPUs.","attack_vector":"Network access to the BMC management interface holding credentials at some authorised level. Whether that is 'remote' depends entirely on how genuinely isolated your OOB network is - in practice most fleets have a jump host, a DCIM integration or a monitoring collector that bridges it.","remediation":"Update the DGX BMC firmware bundle per bulletin 5435. Cost: BMC firmware usually updates without draining the GPUs, but the BMC resets and out-of-band access drops for a few minutes. Pair the patch with an actual audit of who can route to the BMC subnet - that control is worth more than the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42278","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42279","cve":"CVE-2022-42279","aliases":[],"title":"DGX servers BMC: OS command injection on BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"OS command injection on BMC","attack_vector":"Authenticated mgmt-LAN admin","remediation":"Flash BMC 2.09.00+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42279","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"]},{"id":"CVE-2022-42289","cve":"CVE-2022-42289","aliases":[],"title":"DGX-2 SBIOS: OS command injection in firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 SBIOS","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"OS command injection in firmware","attack_vector":"Network-adjacent authenticated admin","remediation":"Flash SBIOS/BMC out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42289","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42290","cve":"CVE-2022-42290","aliases":[],"title":"DGX-2 SBIOS: OS command injection in BIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 SBIOS","year":"2022","cvss_score":7.2,"severity":"high","kev":false,"impact":"OS command injection in BIOS","attack_vector":"Network-adjacent authenticated admin","remediation":"Flash SBIOS out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42290","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-78"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25507","cve":"CVE-2023-25507","aliases":[],"title":"NVIDIA DGX BMC (AMI-derived management controller): The DGX-1 BMC's SPX REST API accepts injected shell commands from an authorised caller, giving direct command…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI-derived management controller)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"The DGX-1 BMC's SPX REST API accepts injected shell commands from an authorised caller, giving direct command execution on the management controller. A BMC compromise on a DGX gives an attacker power control, virtual media, serial console and a persistent foothold under the host OS on a node holding eight GPUs.","attack_vector":"Network access to the BMC management interface holding credentials at some authorised level. Whether that is 'remote' depends entirely on how genuinely isolated your OOB network is - in practice most fleets have a jump host, a DCIM integration or a monitoring collector that bridges it.","remediation":"Update the DGX BMC firmware bundle per bulletin 5458. Cost: BMC firmware usually updates without draining the GPUs, but the BMC resets and out-of-band access drops for a few minutes. Pair the patch with an actual audit of who can route to the BMC subnet - that control is worth more than the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25507","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-77"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25549","cve":"CVE-2023-25549","aliases":["SEVD-2023-101-04"],"title":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - network settings endpoint: Code injection through a parameter of the DCE network-settings endpoint gives remote code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric StruxureWare Data Center Expert (V7.9.2 and prior) - network settings endpoint","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Code injection through a parameter of the DCE network-settings endpoint gives remote code execution on the DCIM appliance. Same outcome as the other DCE RCEs: control of the facility-layer aggregation point.","attack_vector":"Authenticated remote access to the DCE administrative interface.","remediation":"Upgrade past V7.9.2 (SEVD-2023-101-04 covers this whole batch - CVE-2023-25547 through -25555 - so patch once). Rotate stored device credentials afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25549"],"status":"curated"},{"id":"CVE-2023-29002","cve":"CVE-2023-29002","aliases":[],"title":"Cilium: Debug mode logs the contents of the cilium-secrets namespace, including TLS private keys","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Debug mode logs the contents of the cilium-secrets namespace, including TLS private keys","attack_vector":"Anyone with log-pipeline read access","remediation":"Disable agent debug mode; rotate the TLS keys in cilium-secrets; scrub logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29002"],"status":"curated"},{"id":"CVE-2023-31037","cve":"CVE-2023-31037","aliases":[],"title":"BlueField-2 / BlueField-3 DPU BMC: Code injection on DPU BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"BlueField-2 / BlueField-3 DPU BMC","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Code injection on DPU BMC -> DPU takeover below the host OS","attack_vector":"Network-adjacent mgmt access to the DPU BMC","remediation":"Flash DPU BMC firmware out-of-band; DPU reset drops tenant networking, schedule drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31037","https://github.com/NVIDIA/product-security/tree/main/2024/5511"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-94"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31313","cve":"CVE-2023-31313","aliases":[],"title":"AMD Power Management Firmware (PMFW) - unintended proxy to the System Management Unit: MULTI-TENANT ISOLATION: The GPU power management firmware acts as an unintended proxy, letting a privileged…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Power Management Firmware (PMFW) - unintended proxy to the System Management Unit","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The GPU power management firmware acts as an unintended proxy, letting a privileged attacker relay malformed messages through to the System Management Unit. The SMU is the always-on microcontroller that governs clocks, voltages and power limits across the platform; getting arbitrary messages to it through the GPU firmware path means the GPU stack becomes a route into platform-level control. It is the confused-deputy pattern applied to firmware: PMFW is trusted by the SMU, so whoever controls PMFW inherits that trust.","attack_vector":"Local, privileged attacker able to send messages to the GPU power management firmware.","remediation":"Fixed in AMD GPU firmware, which on Instinct parts is delivered as a firmware bundle through the ROCm/amdgpu driver package (the PSP loads the signed blobs at driver init) rather than through the server BIOS. Practically: update the AMD GPU driver/firmware package, then **drain the node and reboot** - the firmware is loaded once at driver init, so a reload of the module with no process holding /dev/kfd is the minimum, and a reboot is what you will actually schedule. Some fixes at this layer also require a **GPU VBIOS flash** via AMD's amdvbflash/amdfwtool, which is an offline, per-card operation with real bricking risk - check the AMD bulletin for whether a VBIOS update is called out before assuming a driver package covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31313","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-3260","cve":"CVE-2023-3260","aliases":[],"title":"Dataprobe iBoot PDU: Authenticated OS command injection on the PDU — an attacker who reaches the power controller can cut power to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dataprobe iBoot PDU","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Authenticated OS command injection on the PDU — an attacker who reaches the power controller can cut power to racks, and can pivot from the PDU into the management network","attack_vector":"Network, authenticated","remediation":"PDU firmware update to 1.44.08042023; PDUs are rarely in the patch pipeline at all, so the real cost is building one","references":["https://thehackernews.com/2023/08/multiple-flaws-in-cyberpower-and.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-32666","cve":"CVE-2023-32666","aliases":[],"title":"Intel 4th Gen Xeon on-chip debug and test interface (with SGX or TDX): MULTI-TENANT ISOLATION: The on-chip debug and test interface has improper access control on 4th-generation…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel 4th Gen Xeon on-chip debug and test interface (with SGX or TDX)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The on-chip debug and test interface has improper access control on 4th-generation Xeon when SGX or TDX is enabled, giving a privileged user escalation. Debug interfaces reaching into a TEE is the single worst shape for a confidential-compute claim, because it bypasses the architectural boundary entirely rather than working around it.","attack_vector":"Privileged local access on the host.","remediation":"Microcode/platform firmware update plus TCB recovery. Where the fix lands in microcode it can be late-loaded at boot; where it lands in BIOS, expect the OEM lag. Re-attest afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32666","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00986.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34341","cve":"CVE-2023-34341","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (SPX REST API): Arbitrary read and write into the memory of the BMC's IPMI server process via the SPX REST API. That is a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (SPX REST API)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Arbitrary read and write into the memory of the BMC's IPMI server process via the SPX REST API. That is a full primitive - the attacker reads out credentials, session tokens and keys held in that process, then writes to redirect execution. In fleet terms, one compromised BMC admin credential turns into code execution on the controller and from there into persistent firmware-level control of the node.","attack_vector":"Network-reachable REST API, requires high privileges - an administrative BMC account. The realistic path is credential reuse: fleets provision BMCs from a template and end up with the same admin password across an entire rack or SKU, so a single leaked credential from one node's config, a Redfish scraper, or a decommissioned host escalates to every node sharing it.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. The higher-leverage work is credential hygiene and is config-only: unique per-node BMC admin passwords generated and stored by your secrets manager, no admin credential embedded in provisioning images or monitoring configs, and Redfish accounts scoped to read-only where they only scrape telemetry.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34341"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34343","cve":"CVE-2023-34343","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (SPX REST API): Shell command injection through the BMC's REST API. An administrative BMC user gets a root shell on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (SPX REST API)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Shell command injection through the BMC's REST API. An administrative BMC user gets a root shell on the management controller's Linux - which is a large step up from what the web UI lets them do, because from a BMC shell the attacker can write firmware, install a persistent implant in the BMC's own flash, and pivot to the host. The gap between 'has a BMC admin password' and 'owns the node forever' closes here.","attack_vector":"Network-reachable REST API with an administrative BMC account. Same shared-credential exposure as the rest of the SPX REST API family: assume any leaked BMC admin password is a fleet-wide credential unless you have proven otherwise.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. Config-only compensations that actually reduce blast radius: unique BMC credentials per node, an allowlist ACL restricting who can reach the BMC web/REST port at all, and logging of BMC authentication to your SIEM so a credential-spray across the management VLAN is visible.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34343"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40289","cve":"CVE-2023-40289","aliases":[],"title":"Supermicro BMC (IPMI web interface, command injection): Command injection that turns a BMC administrator account into shell on the BMC's own Linux. That matters more…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, command injection)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Command injection that turns a BMC administrator account into shell on the BMC's own Linux. That matters more than it sounds: BMC admin is a constrained management role, whereas BMC shell means arbitrary firmware modification, access to the host over the internal bridges, and a place to hide that no host-side agent can inspect. Chained after any of the XSS bugs in the same batch, an operator merely visiting a page is enough to reach it.","attack_vector":"An authenticated BMC administrator - or, realistically, an attacker who chained an XSS in the same firmware to ride an admin's session. Requires network reach to the BMC web interface.","remediation":"BMC firmware flash per board, out-of-band. Supermicro fixes ship per-SKU and lag disclosure, so expect a long tail of boards with no image. Interim controls that work today: keep the BMC web UI off any routable network, require a jump host, and stop using shared BMC admin credentials across the fleet so one compromise is not fleet-wide.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40289","https://www.binarly.io/advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-5528","cve":"CVE-2023-5528","aliases":[],"title":"Kubernetes (in-tree storage): Crafted PV/pod on Windows nodes escalates to node admin via in-tree storage plugin","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (in-tree storage)","year":"2023","cvss_score":7.2,"severity":"high","kev":false,"impact":"Crafted PV/pod on Windows nodes escalates to node admin via in-tree storage plugin","attack_vector":"Cluster user able to create pods and PVs","remediation":"Rolling control-plane and kubelet upgrade; Windows node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5528"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0161","cve":"CVE-2024-0161","aliases":["DSA-2024-006"],"title":"Dell PowerEdge Server BIOS (SMM communication buffer): The BIOS fails to properly validate the SMM communication buffer, so a low-privilege local attacker can write…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS (SMM communication buffer)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"The BIOS fails to properly validate the SMM communication buffer, so a low-privilege local attacker can write into SMRAM. System Management Mode is the most privileged execution context on the machine - more privileged than the hypervisor, invisible to it, and reachable regardless of what OS is booted. An attacker who writes to SMRAM owns the node in a way that no reimage, no disk wipe and no hypervisor-level control removes. On a multi-tenant bare-metal GPU fleet this is the canonical persistent-implant primitive.","attack_vector":"A low-privilege local account on the host. Notably this does NOT need root or administrator - so a tenant workload running as an ordinary user is in scope, as is anything that gets modest code execution through an application bug.","remediation":"System BIOS update. Stage it via iDRAC/Lifecycle Controller or OME, but it lands only on the next reboot, so it costs a job drain and a maintenance window per node. Per-platform version floors are in the advisory table. There is no config-only mitigation for an SMM handler bug - you cannot turn SMM off. Prioritise nodes that run untrusted or multi-tenant workloads over internal-only nodes when sequencing the rollout.","references":["https://www.dell.com/support/kbdoc/en-us/000222979/dsa-2024-006-security-update-for-dell-poweredge-server-bios-for-an-improper-smm-communication-buffer-verification-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2024-0161"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-10237","cve":"CVE-2024-10237","aliases":[],"title":"Supermicro BMC firmware validation (MBD-X12DPG-OA6): Root-of-Trust bypass — firmware image authentication design flaw lets a modified image pass BMC inspection…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation (MBD-X12DPG-OA6)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Root-of-Trust bypass — firmware image authentication design flaw lets a modified image pass BMC inspection and signature verification. Persistent below-OS implant","attack_vector":"Network, high-privilege BMC access","remediation":"BMC flash with a fixed Supermicro build; the RoT bypass means prior firmware state cannot be trusted, so treat affected nodes as requiring re-attestation, not just patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10237"],"status":"curated","fleet":{"ubiquity":"very common - Supermicro is a dominant GPU-server ODM for neoclouds","remediation_pain":"firmware-flash + RoT re-provisioning - the flaw is in the update-validation path itself, so a compromised node may need physical-access recovery of the SPI/BMC flash","pain_class":"physical access","why_fleet_wide":"A signed-firmware-validation bypass means the fleet's defense against malicious firmware is itself the bug: an attacker can push a persistent BMC implant that survives host reinstall and is invisible to the OS."}},{"id":"CVE-2024-10238","cve":"CVE-2024-10238","aliases":[],"title":"Supermicro BMC firmware image verification routine on MBD-X12DPG-OA6: A crafted update image smashes the stack of the code that was supposed to be checking that image, giving the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware image verification routine on MBD-X12DPG-OA6","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A crafted update image smashes the stack of the code that was supposed to be checking that image, giving the attacker execution inside the BMC's update path. The result is an operator-invisible firmware implant on a GPU node: control of power sequencing, virtual media, and the console, plus the ability to falsify what the BMC reports about itself. On a board specifically sold for GPU workloads, this is the persistence mechanism an attacker wants after gaining any temporary administrative access. A dual-socket GPU-oriented board. The parser fails to bound-check a length field in the image container before copying, so the verification code itself is the memory-safety bug.","attack_vector":"An authenticated high-privilege BMC session that can submit a firmware image. Reachability is whatever your management network allows - typically the OOB VLAN plus whatever provisioning automation holds BMC admin credentials.","remediation":"Firmware flash with the fixed BMC image from Supermicro's January 2025 BMC/IPMI advisory. There is no config-only fix, because the vulnerable code is the update handler. Practical hardening in the meantime: deny BMC firmware upload from anything except a single hardened provisioning host, and treat any node whose BMC accepted an unexpected image as needing an out-of-band SPI reflash rather than a software update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10238","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/10xxx/CVE-2024-10238.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-10239","cve":"CVE-2024-10239","aliases":[],"title":"Supermicro OpenBMC firmware image verification (MBD-X12DPG-OA6), fat->fsd.max_fld field: The BMC's own firmware-image parser overflows its stack on an unchecked length field in the image header…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro OpenBMC firmware image verification (MBD-X12DPG-OA6), fat->fsd.max_fld field","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"The BMC's own firmware-image parser overflows its stack on an unchecked length field in the image header - meaning the code that is supposed to decide whether an update is trustworthy can be taken over by the untrusted update itself. This is the bug class that defeats a hardware root of trust: it does not matter that the SoC verifies signatures if you get control inside the verifier before verification completes. The prize is a permanent BMC implant that survives host reimage, node redeployment and tenant handover, on a platform explicitly marketed as an OpenBMC server board.","attack_vector":"Requires BMC administrator privileges to submit a firmware image - so it is a second-stage bug. But 'admin on the BMC' is exactly what the credential-disclosure, default-password and authentication-bypass entries elsewhere in this cluster hand an attacker, and it is also what a malicious insider or a compromised firmware-management pipeline already has.","remediation":"Fixed in Supermicro's January 2025 BMC/IPMI firmware drop - a per-node out-of-band BMC flash, with the ordinary brick risk if interrupted. Beyond patching, this argues for treating BMC firmware update authority as a privileged operation in its own right: restrict which systems hold BMC admin credentials, require change control on firmware pushes, and verify BMC flash contents out-of-band after updates rather than trusting the BMC's own report of what it is running.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10239","https://www.supermicro.com/en/support/security_BMC_IPMI_Jan_2025","https://eclypsium.com/wp-content/uploads/OpenBMC-Security-in-Practice.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21820","cve":"CVE-2024-21820","aliases":[],"title":"Intel Xeon memory controller configuration (with SGX): MULTI-TENANT ISOLATION: Incorrect default permissions on Xeon memory controller configuration when SGX is in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon memory controller configuration (with SGX)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Incorrect default permissions on Xeon memory controller configuration when SGX is in use, reachable by a privileged local user for escalation. Same advisory family as the conditions check issue and fixed by the same platform update.","attack_vector":"Privileged local access on the host.","remediation":"OEM platform BIOS update, drain and reboot, then re-attest enclaves.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21820","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01079.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-22095","cve":"CVE-2024-22095","aliases":[],"title":"Intel Server D50DNP UEFI firmware (PlatformVariableInitDxe): Improper input validation in a UEFI DXE driver on Intel Server D50DNP boards gives a privileged local user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server D50DNP UEFI firmware (PlatformVariableInitDxe)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper input validation in a UEFI DXE driver on Intel Server D50DNP boards gives a privileged local user escalation into firmware. D50DNP is a dense datacenter server board, so this is squarely an AI-datacenter platform. Firmware-level escalation means persistence below the OS that survives reimaging.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22095","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01080.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-28248","cve":"CVE-2024-28248","aliases":[],"title":"Cilium: HTTP policies not consistently applied to all traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"HTTP policies not consistently applied to all traffic; L7 policy bypass","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28248"],"status":"curated"},{"id":"CVE-2024-38509","cve":"CVE-2024-38509","aliases":["LEN-156781"],"title":"Lenovo XClarity Controller (XCC) - IPMI command handler: A specially crafted IPMI command gives an authenticated XCC user arbitrary code execution on the controller.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - IPMI command handler","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A specially crafted IPMI command gives an authenticated XCC user arbitrary code execution on the controller. Code execution on the BMC is the terminal outcome for a node: the attacker holds power control, Virtual Media, console and the firmware write path, and can plant an implant that lives below the hypervisor and survives every reimage. It is one of a cluster of XCC command-injection issues Lenovo fixed across 2024 reachable through IPMI, the SSH captive shell and file upload - the IPMI path is the one that matters most because IPMI is so often left enabled for legacy automation.","attack_vector":"An authenticated XCC user with elevated privileges sending IPMI commands - so the exposure is your administrative BMC credentials plus anything on the management VLAN that can reach the IPMI service. Compromise of an automation host that holds XCC admin creds is the realistic path.","remediation":"Flash XCC to the per-model version in LEN-156781 - out-of-band, per-node, no host reboot and no drain. The high-value config-only mitigation is to disable IPMI over LAN on XCC where your tooling has moved to Redfish, which removes this entire command surface rather than fixing one handler in it. Budget for migrating any remaining ipmitool-based automation first.","references":["https://support.lenovo.com/us/en/product_security/LEN-156781","https://nvd.nist.gov/vuln/detail/CVE-2024-38509"],"status":"curated"},{"id":"CVE-2024-41942","cve":"CVE-2024-41942","aliases":[],"title":"JupyterHub: A user granted limited access can escalate","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"JupyterHub","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A user granted limited access can escalate","attack_vector":"Authenticated notebook user on a shared hub","remediation":"Upgrade to 4.1.6/5.1.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41942"],"status":"curated"},{"id":"CVE-2024-42442","cve":"CVE-2024-42442","aliases":["AMI-SA-2024004"],"title":"AMI AptioV UEFI BIOS (SMM): A memory-bounds bug in the BIOS that lets an attacker execute code outside the intended System Management…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SMM)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"A memory-bounds bug in the BIOS that lets an attacker execute code outside the intended System Management Mode. What sets it apart from the rest of the AptioV set is the attack vector: AMI scores it as network-reachable, not local. A firmware bug you can reach over the wire is a different risk class from one that needs host root first - it means the BIOS attack surface is exposed through a management path rather than only through the host OS, and network segmentation of that path becomes load-bearing.","attack_vector":"Network, with high privileges required - an administrative account on whichever management path exposes the BIOS operation. In practice that means a BMC or out-of-band management credential, which is why this bug chains so naturally with the MegaRAC credential and REST API issues in this same cluster: BMC admin access becomes host firmware code execution.","remediation":"BIOS update to BKC_5.37 or later - firmware flash plus a host reboot, per node, vendor-rebase-gated. Because the vector is network with privileges, there is real config-only mitigation available now: isolate the BMC and management plane so that no untrusted party can reach the privileged management interface, and give every node unique management credentials so one leak does not reach the fleet.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024004.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-42442"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-56161","cve":"CVE-2024-56161","aliases":["EntrySign"],"title":"AMD Zen microcode patch loader (CPU ROM signature verification): MULTI-TENANT ISOLATION: The CPU ROM's microcode patch loader verified patch signatures using AES-CMAC with a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zen microcode patch loader (CPU ROM signature verification)","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The CPU ROM's microcode patch loader verified patch signatures using AES-CMAC with a published example key, so an attacker with local administrator privilege can sign and load their own microcode onto Zen 1 through Zen 4 CPUs. Loading arbitrary microcode means redefining what x86 instructions do - the attacker can make RDRAND return a constant, disable checks, or backdoor the CPU beneath every layer of software. For a confidential-computing operator this is the ballgame: SEV-SNP's entire guarantee rests on the CPU behaving as specified, so a host administrator can now read and tamper with any SEV-SNP guest's memory while the attestation report still looks clean. Google's security team demonstrated the full chain.","attack_vector":"Local, requires ring-0 / host administrator privilege. Not remote, and not reachable from a tenant container. But 'host administrator' is exactly the adversary SEV-SNP exists to exclude, which is why a local-admin bug is a confidential-computing catastrophe rather than a routine escalation. Persistence note: microcode does not survive a power cycle, so an attacker must reload it each boot - which also means a cold boot clears an implant.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly. Specifically: AMD shipped fixed microcode plus an updated ASP bootloader in AGESA (December 2024 / released publicly February 2025). Zen 1-4 EPYC and Ryzen are affected; Zen 5 is not. The critical operator step people skip is the attestation side - after the TCB bump, refresh VCEK certs and require tenants to pin the new minimum TCB, otherwise you are patched but still accepting attestations that would have been valid on a backdoored host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56161","https://github.com/google/security-research/security/advisories/GHSA-4xq7-4mgh-gp6w","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-8531","cve":"CVE-2024-8531","aliases":["SEVD-2024-282-01"],"title":"Schneider Electric Data Center Expert - upgrade bundle signature verification: Improper cryptographic signature verification on DCE upgrade bundles: a manipulated bundle can carry…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric Data Center Expert - upgrade bundle signature verification","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper cryptographic signature verification on DCE upgrade bundles: a manipulated bundle can carry arbitrary bash scripts that execute as root. Your patching process becomes the attack. Anyone who can place a bundle in front of the appliance - a compromised mirror, an internal file share, a helpful vendor email - gets root on the system holding the facility's power and cooling credentials.","attack_vector":"Requires getting a crafted upgrade bundle to the appliance. In practice: whoever runs DCE upgrades, or anyone who can tamper with where the bundles are staged.","remediation":"Upgrade DCE per SEVD-2024-282-01 to a build that verifies signatures correctly. Until then, treat upgrade bundles as untrusted code: obtain them only over an authenticated channel from the vendor, verify hashes out of band, and stage them somewhere with restricted write access.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2024-282-01&p_enDocType=Security+and+Safety+Notice&p_File_Name=SEVD-2024-282-01.pdf"],"status":"curated"},{"id":"CVE-2024-9180","cve":"CVE-2024-9180","aliases":[],"title":"HashiCorp Vault: Operator with write on the root namespace identity endpoint escalates self/others to the root policy","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2024","cvss_score":7.2,"severity":"high","kev":false,"impact":"Operator with write on the root namespace identity endpoint escalates self/others to the root policy","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; re-scope operator policies","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9180"],"status":"curated"},{"id":"CVE-2024-9474","cve":"CVE-2024-9474","aliases":[],"title":"Palo Alto PAN-OS: Admin with mgmt-interface access performs firewall actions as root","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Palo Alto PAN-OS","year":"2024","cvss_score":7.2,"severity":"high","kev":true,"impact":"[KEV] Admin with mgmt-interface access performs firewall actions as root; chained with CVE-2024-0012","attack_vector":"Network (remote)","remediation":"Control-plane: same patch window; assume compromise if mgmt was exposed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9474"],"status":"curated"},{"id":"CVE-2025-0032","cve":"CVE-2025-0032","aliases":[],"title":"AMD CPU microcode patch loading - improper cleanup: MULTI-TENANT ISOLATION: Improper cleanup during microcode patch loading gives a local administrator another…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU microcode patch loading - improper cleanup","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper cleanup during microcode patch loading gives a local administrator another route to load malicious CPU microcode, costing integrity of x86 instruction execution itself. This is the same category of failure as EntrySign and lands in the same place: the CPU can be made to lie, and everything built on top of it - SEV-SNP guest isolation included - inherits the lie.","attack_vector":"Local, administrator privilege. Reboot clears a loaded implant, but the attacker who has admin can simply reload it every boot.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0032","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-12006","cve":"CVE-2025-12006","aliases":[],"title":"Supermicro BMC firmware validation logic on the X12STW-F motherboard: An attacker with administrative reach to the BMC installs a firmware image of their own construction. What…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation logic on the X12STW-F motherboard","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"An attacker with administrative reach to the BMC installs a firmware image of their own construction. What they walk away with is a controller that owns node power, console, virtual media and the host's boot path, and that keeps owning it after the operator wipes and reprovisions the machine. Because BMC credentials in most fleets are identical across every node, one admin-level compromise converts directly into a fleet-wide persistent foothold that no host-level EDR or reimage cycle will find. The same class of image-verification weakness as its X13 sibling, showing the flaw spans two board generations rather than one SKU.","attack_vector":"An authenticated BMC session at administrator privilege reachable over the network. That bar is lower than it sounds in real datacenters: shared or default BMC credentials, a leaked provisioning secret, or chaining any of the several authenticated Supermicro BMC command-execution bugs in this same list will get an attacker there.","remediation":"Firmware flash per node using the fixed image from Supermicro's January 2026 BMC/IPMI batch, matched to the exact board SKU. Config-only mitigation is partial but worth doing immediately: rotate BMC administrator credentials so they are unique per node rather than shared fleet-wide, and remove any standing admin accounts used by automation in favour of scoped operator-level accounts. Network isolation of the management VLAN remains the backstop for nodes that cannot be taken down for a flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12006","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/12xxx/CVE-2025-12006.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-20037","cve":"CVE-2025-20037","aliases":[],"title":"Intel CSME firmware (TOCTOU): A time-of-check/time-of-use race in CSME firmware lets a privileged local user escalate into the management…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME firmware (TOCTOU)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"A time-of-check/time-of-use race in CSME firmware lets a privileged local user escalate into the management engine. Recent, and a reminder that the CSME attack surface is still producing findings on current platforms.","attack_vector":"Privileged local access on the host, plus winning a race.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20037","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01280.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-20053","cve":"CVE-2025-20053","aliases":[],"title":"Intel Xeon processor firmware (SGX enabled): MULTI-TENANT ISOLATION: Improper buffer restrictions in Xeon firmware on SGX-enabled parts, giving a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon processor firmware (SGX enabled)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper buffer restrictions in Xeon firmware on SGX-enabled parts, giving a privileged local user an escalation path. Firmware-level, so it lands underneath anything the OS can defend.","attack_vector":"Privileged local access on the host.","remediation":"Platform firmware/BIOS update from the OEM plus SGX TCB recovery. Full drain and reboot per node; OEM release timing dominates the rollout.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20053","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01313.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-26403","cve":"CVE-2025-26403","aliases":[],"title":"Intel Xeon 6 memory subsystem (with SGX or TDX): MULTI-TENANT ISOLATION: An out-of-bounds write in the Xeon 6 memory subsystem reachable when SGX or TDX is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 memory subsystem (with SGX or TDX)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds write in the Xeon 6 memory subsystem reachable when SGX or TDX is enabled, escalating privilege for a privileged local user. An OOB write in the memory subsystem of a confidential-compute platform undercuts both the SGX and the TDX guarantee on the same silicon.","attack_vector":"Privileged local access on a Xeon 6 host with SGX or TDX enabled.","remediation":"OEM platform firmware/BIOS update, plus TDX module and SGX TCB recovery as applicable. Drain and reboot; re-attest every TD and enclave afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-26403","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01367.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-32086","cve":"CVE-2025-32086","aliases":[],"title":"Intel Xeon 6 DDRIO configuration (with SGX or TDX): MULTI-TENANT ISOLATION: An improperly implemented security check in DDRIO configuration on Xeon 6 lets a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 DDRIO configuration (with SGX or TDX)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An improperly implemented security check in DDRIO configuration on Xeon 6 lets a privileged local user escalate on SGX/TDX-enabled platforms. DDRIO sits between the memory controller and DRAM, which is the layer memory-encryption integrity depends on.","attack_vector":"Privileged local access on a Xeon 6 host with SGX or TDX enabled.","remediation":"OEM platform firmware/BIOS update plus TCB recovery for both SGX and TDX. Drain, reboot, re-attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32086","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01367.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-58770","cve":"CVE-2025-58770","aliases":["AMI-SA-2025009"],"title":"AMI AptioV UEFI BIOS: Improper handling of insufficient permissions in the BIOS lets a low-privileged local user escalate their…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Improper handling of insufficient permissions in the BIOS lets a low-privileged local user escalate their authorization, with integrity and availability impact that AMI scores as reaching the subsequent system too. What makes this one worth prioritising over its neighbours is that AMI's own CVSS vector marks exploit maturity as proof-of-concept - meaning working exploit code exists publicly, not just a theoretical write-up. It is also the newest entry in AMI's published series, so ODM rebased images are the least likely to be available.","attack_vector":"Local access with only low privileges required and no user interaction. That is a notably low bar for a firmware bug - it does not need root, so an unprivileged process or a compromised service account on the host is enough to start escalating toward firmware.","remediation":"BIOS update to AptioV_5.041 or later: firmware flash plus a full host reboot, per node, and expect the longest ODM lag of anything in this cluster because the advisory is recent. No config-only fix. Given the low privilege requirement, the interim control is ordinary host hardening - reduce what unprivileged local code exists on GPU nodes at all, and treat any node where untrusted tenant code runs as already exposed until the BIOS is updated.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025009.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-58770"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-6198","cve":"CVE-2025-6198","aliases":[],"title":"Supermicro BMC firmware validation (MBD-X13SEM-F): Second-generation RoT bypass — crafted image passes BMC firmware validation logic (CWE-347, improper…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation (MBD-X13SEM-F)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Second-generation RoT bypass — crafted image passes BMC firmware validation logic (CWE-347, improper signature verification); survives reimaging","attack_vector":"Network, high-privilege","remediation":"Per-node out-of-band BMC flash to the fixed Supermicro build; no host-side mitigation exists","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-6198"],"status":"curated","fleet":{"ubiquity":"very common - affects an additional Supermicro product set beyond CVE-2024-10237","remediation_pain":"firmware-flash - out-of-band per node; once RoT is defeated, prior attestation evidence is untrustworthy and nodes need re-baselining","pain_class":"firmware-flash","why_fleet_wide":"Bypasses the BMC Root of Trust, so the firmware-signing anchor a neocloud relies on for tenant-isolation claims is defeated below the OS."}},{"id":"CVE-2025-62626","cve":"CVE-2025-62626","aliases":[],"title":"AMD CPUs - attacker influence over RDSEED entropy: MULTI-TENANT ISOLATION: A local attacker can influence the values RDSEED returns, causing consumers to draw…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPUs - attacker influence over RDSEED entropy","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A local attacker can influence the values RDSEED returns, causing consumers to draw insufficient entropy. Anything on the node that seeds a key, nonce or token from the hardware RNG - including confidential guests that deliberately chose the hardware source because they do not trust the host - gets attacker-influenced material. The failure is silent: the instruction reports success, and nothing downstream can tell the difference until someone else predicts the key.","attack_vector":"Local. The attacker influences entropy consumed by other software on the same machine, including guests.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Also update guest and host kernels so the OS entropy pool does not lean solely on RDSEED. Any long-lived key generated on an affected host before patching should be rotated - the fix protects future output, not keys already derived.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-62626","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-67818","cve":"CVE-2025-67818","aliases":[],"title":"Weaviate: Crafted entry name with an absolute path","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Weaviate","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Crafted entry name with an absolute path → arbitrary file write","attack_vector":"Tenant with data-insert permission","remediation":"Upgrade past 1.33.4; data insertion is a filesystem-write primitive","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-67818"],"status":"curated"},{"id":"CVE-2025-7937","cve":"CVE-2025-7937","aliases":[],"title":"Supermicro BMC firmware validation (MBD-X12STW): RoT bypass, crafted firmware image accepted","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC firmware validation (MBD-X12STW)","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"RoT bypass, crafted firmware image accepted; demonstrates the 2024 fix was incomplete","attack_vector":"Network, high-privilege","remediation":"Second BMC flash cycle on nodes already patched for CVE-2024-10237 — the operational cost is that a fleet gets flashed twice for the same class of bug","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7937"],"status":"curated","fleet":{"ubiquity":"very common - same Supermicro BMC image family","remediation_pain":"firmware-flash - a second flash cycle on nodes already flashed once; the first remediation round did not hold","pain_class":"firmware-flash","why_fleet_wide":"Shows fleet firmware remediation is not one-and-done: the patch was bypassed, forcing a repeat out-of-band flash campaign across the same node population."}},{"id":"CVE-2025-8076","cve":"CVE-2025-8076","aliases":[],"title":"Supermicro BMC web server request handling on MBD-X13SEDW-F: Any account that can log into the BMC web interface can turn that login into code execution on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC web server request handling on MBD-X13SEDW-F","year":"2025","cvss_score":7.2,"severity":"high","kev":false,"impact":"Any account that can log into the BMC web interface can turn that login into code execution on the controller. For an operator this collapses the distinction between 'someone has a BMC password' and 'someone owns the node out of band' - they get power control, console, virtual-media boot of an attacker image, and a persistence point below the hypervisor that reimaging will not clear. A post-authentication stack buffer overflow triggered by a crafted payload to the management web UI.","attack_vector":"An authenticated high-privilege session against the BMC's HTTP interface, reachable from anywhere routable to the out-of-band management network. In fleets that share one BMC password across every node - still the norm - a single credential leak makes this exploitable everywhere at once.","remediation":"Firmware flash from Supermicro's November 2025 BMC/IPMI advisory. Config-only measures that reduce exposure now: put BMC web access behind a bastion so it is not reachable from the general management subnet, rotate to per-node BMC credentials, and disable the BMC web UI on nodes managed purely via Redfish or IPMI. Flashing is per node and out of band; plan it alongside the other X13 BMC fixes in the same advisory so you only take one flash window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8076","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/8xxx/CVE-2025-8076.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-23759","cve":"CVE-2026-23759","aliases":[],"title":"Perle IOLAN STS/SCS terminal server (firmware before 6.0): A logged-in user of the restricted admin shell (Telnet or SSH) can break out of that shell and run arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Perle IOLAN STS/SCS terminal server (firmware before 6.0)","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"A logged-in user of the restricted admin shell (Telnet or SSH) can break out of that shell and run arbitrary OS commands as root. The 'ps' subcommand doesn't sanitize its arguments before handing them to a shell, so an operator account that was only supposed to have limited diagnostic access ends up with full root on the terminal server.","attack_vector":"Requires an authenticated login to the restricted shell (any account that can reach the 'ps' command), then injects shell metacharacters after the subcommand.","remediation":"Firmware upgrade to 6.0 or later — this is a shell-sanitization bug in the restricted-admin feature, not something a permission change alone fixes. Flash each terminal server one at a time; expect a brief loss of the serial sessions it's terminating during the reboot.","references":["https://www.perle.com/downloads/server_sds_sts_rackmount.shtml","https://www.vulncheck.com/advisories/perle-iolan-sts-scs-authenticated-command-injection-via-shell-ps"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-3820","cve":"CVE-2026-3820","aliases":[],"title":"Supermicro BMC SMTP service configuration handler on AS-2115HS-TNR and related boards: Crafted characters injected into the SMTP configuration are executed by the underlying BMC system when the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC SMTP service configuration handler on AS-2115HS-TNR and related boards","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"Crafted characters injected into the SMTP configuration are executed by the underlying BMC system when the mail process is invoked. The attacker gets arbitrary code execution or a hard denial of service on the controller, and Supermicro's own wording allows for permanent compromise of the controller - meaning an implant that persists in BMC flash. For an operator this is the same nightmare as any BMC RCE: out-of-band power and console control, virtual-media boot of an attacker image, and a foothold below the reimage boundary. The 2026 recurrence of the same notification-service injection class that produced CVE-2023-35861 three years earlier.","attack_vector":"An attacker who has obtained BMC administrator privileges and can reach the BMC's configuration interface over the network. Shared fleet-wide BMC credentials, leaked provisioning secrets, or a chained lower-privilege bug all put an attacker at this level.","remediation":"Firmware flash from Supermicro's June 2026 BMC/IPMI advisory batch, per board SKU. Config-only interim mitigation: disable BMC SMTP alerting and rotate BMC admin credentials to per-node unique values. The fact that this is the second SMTP-handler injection in this codebase in three years is itself operator-relevant - treat the BMC notification subsystem as untrusted attack surface and disable it on nodes that get their alerting from IPMI polling or Redfish instead.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-3820","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2026/3xxx/CVE-2026-3820.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-42306","cve":"CVE-2026-42306","aliases":[],"title":"Docker / moby: Race condition during `docker cp` mount setup allows escape/host access","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"Race condition during `docker cp` mount setup allows escape/host access","attack_vector":"Any tenant workload on a node where docker cp is used","remediation":"Upgrade Docker Engine to 29.5.1+; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-42306"],"status":"curated"},{"id":"CVE-2026-6973","cve":"CVE-2026-6973","aliases":[],"title":"Ivanti Endpoint Manager Mobile: Improper input validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ivanti Endpoint Manager Mobile","year":"2026","cvss_score":7.2,"severity":"high","kev":true,"impact":"[KEV] Improper input validation -> authenticated admin achieves remote code execution","attack_vector":"Network (remote)","remediation":"Control-plane: patch; restrict admin access to the ops network","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-6973"],"status":"curated"},{"id":"CVE-2026-9777","cve":"CVE-2026-9777","aliases":["ZDI-26-381"],"title":"ATEN Unizon fleet management platform: TENANT ISOLATION: Unizon is ATEN's centralized manager for its KVM and PDU fleet. The restoreDB function…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATEN Unizon fleet management platform","year":"2026","cvss_score":7.2,"severity":"high","kev":false,"impact":"TENANT ISOLATION: Unizon is ATEN's centralized manager for its KVM and PDU fleet. The restoreDB function doesn't validate a user-supplied path before writing to it, letting an authenticated attacker write files anywhere on the host and execute code with SYSTEM privileges — full compromise of the platform that has management-plane reach into every KVM and PDU it administers.","attack_vector":"Requires an authenticated account on Unizon (the advisory doesn't specify a high privilege tier is needed), then sends a crafted path to the restoreDB endpoint.","remediation":"Software upgrade to the patched Unizon release. Since Unizon is the single management server for the whole device fleet rather than per-device firmware, this is one upgrade — but treat it as urgent given the blast radius (SYSTEM-level access to the platform that manages every connected KVM/PDU).","references":["https://www.aten.com/global/en/supportcenter/info/security-advisory/30/","https://www.zerodayinitiative.com/advisories/ZDI-26-381/"],"status":"curated"},{"id":"CVE-2019-0090","cve":"CVE-2019-0090","aliases":["Intel x86 Root of Trust: loss of trust"],"title":"Intel CSME / Converged Security and Management Engine (mask ROM): A flaw in the CSME boot ROM window before memory protections engage allows code execution in the engine that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME / Converged Security and Management Engine (mask ROM)","year":"2019","cvss_score":7.1,"severity":"high","kev":false,"impact":"A flaw in the CSME boot ROM window before memory protections engage allows code execution in the engine that is the hardware root of trust for the whole platform - it is what verifies BIOS under Boot Guard, backs Intel PTT/fTPM, and holds the chipset key from which platform-unique keys derive. Researchers demonstrated extraction of that chipset key, which forges the identity of the platform itself: EPID-based attestation, DRM, and any measurement chain rooted in CSME become untrustworthy, and the compromise is not visible from the OS at all.","attack_vector":"Local or physical access to the machine during the early boot window. On bare-metal GPU rental, a tenant with root plus a reboot is the realistic actor; on a colo floor, so is anyone with hands.","remediation":"Cannot be fully fixed. The vulnerable code is in mask ROM, so no firmware update replaces it - Intel's CSME updates only narrow the exploitation window. The durable answer is hardware generations that are not affected, and until then treating CSME-rooted attestation as advisory rather than authoritative. If you sell attestation guarantees, do not root them here; root them in a discrete device you control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0090","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00213.html"],"status":"curated"},{"id":"CVE-2019-5687","cve":"CVE-2019-5687","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A kernel object created by the escape handler gets default permissions that expose it to accounts that should…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":7.1,"severity":"high","kev":false,"impact":"A kernel object created by the escape handler gets default permissions that expose it to accounts that should not reach it. Local privilege boundary erosion inside the driver.","attack_vector":"Any local user on the host who can open the over-permissive object.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://nvd.nist.gov/vuln/detail/CVE-2019-5687"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5697","cve":"CVE-2019-5697","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the vGPU Manager grants a guest access to memory the guest does not own. That is the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2019","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the vGPU Manager grants a guest access to memory the guest does not own. That is the core vGPU isolation promise failing - a tenant VM reads memory belonging to the host or to another tenant's vGPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5697"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5366","cve":"CVE-2020-5366","aliases":["DSA-2020-128"],"title":"Dell iDRAC9 (web interface, local file inclusion): A path-traversal / local-file-inclusion flaw lets a low-privilege iDRAC user read files outside the intended…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 (web interface, local file inclusion)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"A path-traversal / local-file-inclusion flaw lets a low-privilege iDRAC user read files outside the intended directory on the BMC. The value to an attacker is escalation material - configuration, credentials and session data that upgrade a read-only monitoring account into real control of the service processor. This is the classic privilege-ladder rung: the account you handed to a monitoring agent becomes the account that owns the node's out-of-band plane.","attack_vector":"An authenticated low-privilege iDRAC operator or read-only account reaching the iDRAC web interface over the management VLAN. No host access needed.","remediation":"Flash iDRAC9 to 4.20.20.20 or later - out-of-band, per-node, no host reboot, no drain. There is no clean config-only mitigation for this one beyond tightening who holds iDRAC accounts at all, so treat it as a firmware campaign. Worth pairing with an audit of low-privilege iDRAC service accounts, which are usually shared and rarely rotated.","references":["https://www.dell.com/support/article/en-us/sln322125/dsa-2020-128-idrac-local-file-inclusion-vulnerability?lang=en","https://nvd.nist.gov/vuln/detail/CVE-2020-5366"],"status":"curated"},{"id":"CVE-2020-5970","cve":"CVE-2020-5970","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: guest-supplied data size is not validated, letting a tenant tamper with host-side…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: guest-supplied data size is not validated, letting a tenant tamper with host-side state or take the GPU down for co-tenants. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5970"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5972","cve":"CVE-2020-5972","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: uninitialised local pointers later freed in the vGPU plugin - a guest-triggered free…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: uninitialised local pointers later freed in the vGPU plugin - a guest-triggered free of an arbitrary pointer, which is a host memory-corruption primitive. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5972"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5983","cve":"CVE-2020-5983","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin + host kernel module): MULTI-TENANT ISOLATION: the host can be made to write outside the frame-buffer region allocated to a guest.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin + host kernel module)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the host can be made to write outside the frame-buffer region allocated to a guest. That is a direct breach of the vGPU frame-buffer partition - one tenant's writes landing in memory belonging to the host or another tenant's vGPU. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5983"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5985","cve":"CVE-2020-5985","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: guest-supplied length not validated, allowing host-side data tampering or a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: guest-supplied length not validated, allowing host-side data tampering or a shared-GPU outage. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5985"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5988","cve":"CVE-2020-5988","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: double free in the vGPU plugin, guest-triggered. Host heap corruption or disclosure.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: double free in the vGPU plugin, guest-triggered. Host heap corruption or disclosure. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5988"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1056","cve":"CVE-2021-1056","aliases":[],"title":"NVIDIA Linux GPU Display Driver (nvidia.ko): MULTI-TENANT ISOLATION: nvidia.ko does not fully honour filesystem permissions when providing GPU…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Linux GPU Display Driver (nvidia.ko)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: nvidia.ko does not fully honour filesystem permissions when providing GPU device-level isolation. This is the one that matters for containerised GPU fleets - it means the device-node permission model your container runtime relies on to give each container only its assigned GPUs is not actually enforced by the driver, so a container can reach GPUs it was not allocated. On a shared Kubernetes GPU node running multiple tenants' pods, that is a straight isolation break between tenants. Debian and Gentoo both shipped it as a security update, so distro-packaged fleets are in scope.","attack_vector":"Any container or local user on a Linux GPU host with access to some subset of the NVIDIA device nodes - which is every GPU workload.","remediation":"Install the fixed Linux GPU Display Driver branch. nvidia.ko / nvidia-uvm.ko cannot be replaced while any process holds a GPU, so plan a node drain: cordon the node, stop every CUDA job and GPU container, unload the modules or reboot, install, reload. Container runtimes that bind-mount the driver libraries (nvidia-container-toolkit) need restarting so running pods pick up the new userspace. No firmware flash. On multi-tenant nodes, do not treat device-node permissions or the container toolkit's device isolation as a security boundary until the driver is patched; one tenant per node or MIG-backed partitioning is the only reliable interim control.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1056"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1058","cve":"CVE-2021-1058","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): MULTI-TENANT ISOLATION: unvalidated input size across the guest kernel-mode driver and vGPU plugin boundary…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated input size across the guest kernel-mode driver and vGPU plugin boundary - a tenant tampers with host-side data or crashes the shared GPU. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1058"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1060","cve":"CVE-2021-1060","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): MULTI-TENANT ISOLATION: unvalidated index across the guest driver and vGPU plugin boundary, giving a tenant…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated index across the guest driver and vGPU plugin boundary, giving a tenant host-side data tampering or a shared-GPU outage. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1060"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1062","cve":"CVE-2021-1062","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated guest-supplied length in the vGPU plugin","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated guest-supplied length in the vGPU plugin; tenant-triggered host data tampering or denial of service on the shared GPU. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1062"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1064","cve":"CVE-2021-1064","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the plugin takes a value from the guest, casts it to a pointer and dereferences it. A…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the plugin takes a value from the guest, casts it to a pointer and dereferences it. A tenant chooses a host address to read - arbitrary host-memory disclosure or a crash. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1064"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1065","cve":"CVE-2021-1065","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated guest input in the vGPU plugin leading to host data tampering or a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated guest input in the vGPU plugin leading to host data tampering or a shared-GPU outage. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1065"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1086","cve":"CVE-2021-1086","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the vGPU Manager lets guests control resources they are not entitled to, which NVIDIA…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: the vGPU Manager lets guests control resources they are not entitled to, which NVIDIA describes as integrity and confidentiality loss. A tenant reaching resources outside its partition is the vGPU security model failing. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1086"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1090","cve":"CVE-2021-1090","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Out-of-bounds read or write in the kernel-mode layer's control-call handler on both Windows and Linux.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Out-of-bounds read or write in the kernel-mode layer's control-call handler on both Windows and Linux. Unprivileged local code corrupts kernel memory or crashes the node; Gentoo shipped it as a security update.","attack_vector":"Any local user or GPU container with access to the NVIDIA device nodes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1090"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1091","cve":"CVE-2021-1091","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Hard-link attack: an unprivileged user makes the driver overwrite a file that needs elevated privilege to…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Hard-link attack: an unprivileged user makes the driver overwrite a file that needs elevated privilege to modify. Arbitrary privileged file overwrite, which is a well-trodden route to SYSTEM.","attack_vector":"Any local unprivileged user on a Windows GPU host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1091"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1119","cve":"CVE-2021-1119","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: double free in the vGPU Manager that NVIDIA explicitly describes as a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: double free in the vGPU Manager that NVIDIA explicitly describes as a write-what-where condition allowing arbitrary code execution. A tenant VM writing chosen values to chosen host addresses is a full hypervisor-host compromise from inside a guest.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1119"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-21539","cve":"CVE-2021-21539","aliases":[],"title":"Dell iDRAC9: TOCTOU race during simultaneous web-interface access — state corruption on the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"TOCTOU race during simultaneous web-interface access — state corruption on the BMC","attack_vector":"Network, authenticated","remediation":"iDRAC firmware update","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21539"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26332","cve":"CVE-2021-26332","aliases":[],"title":"AMD SEV-ES firmware - TMR placement in MMIO space: MULTI-TENANT ISOLATION: SEV-ES firmware does not verify that the Trusted Memory Region is outside MMIO space.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES firmware - TMR placement in MMIO space","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: SEV-ES firmware does not verify that the Trusted Memory Region is outside MMIO space. Point the TMR at MMIO and the secure firmware's private working memory is suddenly backed by device registers the host controls - a route to observing or influencing what the SEV firmware does, costing integrity or availability of confidential guests.","attack_vector":"Hypervisor-privileged attacker who controls where the TMR is placed.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26332","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26402","cve":"CVE-2021-26402","aliases":[],"title":"AMD Secure Processor firmware - BIOS mailbox command bounds checking: MULTI-TENANT ISOLATION: Insufficient bounds checking while the ASP firmware handles BIOS mailbox commands…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor firmware - BIOS mailbox command bounds checking","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient bounds checking while the ASP firmware handles BIOS mailbox commands lets an attacker write partially-controlled data out of bounds into SMM or SEV-protected memory. Both destinations are places the OS is explicitly not allowed to reach: SMM is the most privileged execution mode on x86, and SEV memory belongs to confidential guests.","attack_vector":"Local, via the BIOS mailbox interface - requires host privilege.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26402","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-27364","cve":"CVE-2021-27364","aliases":[],"title":"Linux iSCSI: Unprivileged user can craft Netlink messages to scsi_transport_iscsi","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iSCSI","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unprivileged user can craft Netlink messages to scsi_transport_iscsi","attack_vector":"Local","remediation":"Data-plane: kernel patch + reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27364"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-28507","cve":"CVE-2021-28507","aliases":[],"title":"Arista EOS (service ACLs): Service ACL bypass for OpenConfig gNOI and RESTCONF — the compensating control for the above does not hold","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (service ACLs)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"Service ACL bypass for OpenConfig gNOI and RESTCONF — the compensating control for the above does not hold","attack_vector":"Network","remediation":"EOS upgrade; important because it invalidates \"we ACL'd the management API\" as a mitigation","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28507"],"status":"curated"},{"id":"CVE-2021-28692","cve":"CVE-2021-28692","aliases":[],"title":"Xen - x86 IOMMU command timeout detection and handling: MULTI-TENANT ISOLATION: Xen's IOMMU command timeout handling is inappropriate, so IOMMU operations that do…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen - x86 IOMMU command timeout detection and handling","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Xen's IOMMU command timeout handling is inappropriate, so IOMMU operations that do not complete in time are mishandled. On a host doing PCI passthrough - which is every GPU cloud - the IOMMU is the component enforcing that an assigned device can only DMA into its owner's memory. Mishandled timeouts mean that enforcement can be left in an indeterminate state while devices keep running.","attack_vector":"Requires a guest able to generate IOMMU load, i.e. a guest with an assigned device doing heavy DMA - normal behaviour for a passed-through GPU or NIC.","remediation":"Fixed in Xen (XSA-373). Hypervisor update plus host reboot. Especially relevant on GPU nodes, where passed-through accelerators and RDMA NICs push the IOMMU hard enough to hit timeout paths that idle hosts never reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28692","https://xenbits.xen.org/xsa/advisory-373.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-3571","cve":"CVE-2021-3571","aliases":[],"title":"linuxptp / ptp4l (transparent clock on little-endian): A crafted PTP packet against ptp4l running as a transparent clock on a little-endian machine — i.e. every x86…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"linuxptp / ptp4l (transparent clock on little-endian)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"A crafted PTP packet against ptp4l running as a transparent clock on a little-endian machine — i.e. every x86 and ARM64 server in your cluster — produces a fault. Transparent-clock mode is exactly the configuration used when PTP is carried across switches inside the cluster, so the affected deployment is the mainstream one, not an edge case.","attack_vector":"Remote, unauthenticated — a crafted PTP message from anything that can reach the node's PTP port.","remediation":"linuxptp package upgrade plus ptp4l restart. Same segmentation advice as the companion issue: PTP traffic should not be sourceable by tenant workloads.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3571"],"status":"curated"},{"id":"CVE-2021-36309","cve":"CVE-2021-36309","aliases":[],"title":"Dell Enterprise SONiC OS (information disclosure): An authenticated user can extract sensitive information from Enterprise SONiC 3.3.0 and earlier. On a switch…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC OS (information disclosure)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"An authenticated user can extract sensitive information from Enterprise SONiC 3.3.0 and earlier. On a switch, 'sensitive information' generally means credentials for the things the switch talks to — TACACS/RADIUS secrets, SNMP communities, image-server logins — so the impact propagates outward from the device.","attack_vector":"Authenticated user with system access on the switch.","remediation":"NOS image upgrade plus reboot, then rotate every shared secret configured on the switch. The rotation is the expensive part on a fleet that shares TACACS keys across all devices.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-36309"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-4460","cve":"CVE-2021-4460","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2021","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdkfd: Fix UBSAN shift-out-of-bounds warning","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-4460","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-28184","cve":"CVE-2022-28184","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): MULTI-TENANT ISOLATION: The DxgkDdiEscape handler lets an unprivileged user reach administrator-privileged…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The DxgkDdiEscape handler lets an unprivileged user reach administrator-privileged registers, giving direct read/write to GPU control state that should be off-limits to tenant code. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28184","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-284"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-2989","cve":"CVE-2022-2989","aliases":[],"title":"Podman: Incorrect supplementary group handling","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Incorrect supplementary group handling; information disclosure or data modification","attack_vector":"Any tenant workload","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2989"],"status":"curated"},{"id":"CVE-2022-2995","cve":"CVE-2022-2995","aliases":[],"title":"CRI-O: Incorrect supplementary group handling leads to information disclosure between workloads","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Incorrect supplementary group handling leads to information disclosure between workloads","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-2995"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31612","cve":"CVE-2022-31612","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read through DxgkDdiEscape leaks internal kernel information or crashes the node. Only…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds read through DxgkDdiEscape leaks internal kernel information or crashes the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31612","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31613","cve":"CVE-2022-31613","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Any local user can null-pointer-dereference the kernel mode layer and panic the node - no privileges required…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Any local user can null-pointer-dereference the kernel mode layer and panic the node - no privileges required at all. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31613","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-32521","cve":"CVE-2022-32521","aliases":["SEVD-2023-010-06"],"title":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - Java deserialization: Unsafe deserialization of data posted to the web server yields remote code execution on the DCIM appliance.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric Data Center Expert (versions prior to v7.9.0) - Java deserialization","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unsafe deserialization of data posted to the web server yields remote code execution on the DCIM appliance. Standard deserialization bug, non-standard consequence: the host it lands on controls the power and cooling telemetry and credentials for the building.","attack_vector":"Remote, by posting crafted serialized data to the DCE web server.","remediation":"Upgrade to DCE v7.9.0 or later. Rotate stored credentials. Restrict who can reach the DCE web interface to a management jump host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32521"],"status":"curated"},{"id":"CVE-2022-34676","cve":"CVE-2022-34676","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An out-of-bounds read in the kernel mode handler leaks kernel memory contents or crashes the node. Everything…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds read in the kernel mode handler leaks kernel memory contents or crashes the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34676","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-197"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-35929","cve":"CVE-2022-35929","aliases":[],"title":"cosign / sigstore: `cosign verify-attestation --type` returns a false positive if any attestation exists","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"`cosign verify-attestation --type` returns a false positive if any attestation exists; unsigned images pass policy","attack_vector":"Malicious image","remediation":"Upgrade cosign; re-verify every image admitted while the vulnerable version was in the gate","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-35929"],"status":"curated"},{"id":"CVE-2022-41737","cve":"CVE-2022-41737","aliases":[],"title":"IBM Storage Scale Container Native Storage Access (namespace boundary): TENANT ISOLATION: a local attacker can initiate connections from a container outside its current namespace.…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"IBM Storage Scale Container Native Storage Access (namespace boundary)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"TENANT ISOLATION: a local attacker can initiate connections from a container outside its current namespace. Network-namespace escape from a storage-access container is a direct route from one tenant's pod onto networks the pod was never meant to touch — including, on most cluster designs, the storage back-end network where authentication is weak because it is assumed to be private.","attack_vector":"A local attacker inside a container using Storage Scale container-native access, versions 5.1.2.1 through 5.1.7.0.","remediation":"Upgrade Container Native Storage Access past 5.1.7.0 — rolling operator/DaemonSet upgrade. Companion issue CVE-2022-41738 allows connections *into* containers from external networks; both are closed by the same upgrade path. Also treat the storage back-end network as authenticated rather than trusted, which is an architectural change and the durable answer.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41737","https://nvd.nist.gov/vuln/detail/CVE-2022-41738"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-42262","cve":"CVE-2022-42262","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A second unvalidated index path in the vGPU plugin with the same guest-driven host…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A second unvalidated index path in the vGPU plugin with the same guest-driven host buffer overrun. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5415. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42262","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42263","cve":"CVE-2022-42263","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in the kernel mode handler leaks kernel information or crashes the node. Everything with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An integer overflow in the kernel mode handler leaks kernel information or crashes the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42263","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42264","cve":"CVE-2022-42264","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user drives the kernel to use an out-of-range pointer offset, producing data corruption, data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"An unprivileged user drives the kernel to use an out-of-range pointer offset, producing data corruption, data loss, kernel information disclosure or a crash. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42264","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-823"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42280","cve":"CVE-2022-42280","aliases":[],"title":"DGX-2 BMC: Path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Path traversal -> file read/write on BMC","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42280","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42327","cve":"CVE-2022-42327","aliases":["XSA-412"],"title":"Xen (x86): Unintended memory sharing between guests - cross-tenant data exposure","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Unintended memory sharing between guests - cross-tenant data exposure","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + evacuation/reboot. Direct multi-tenant isolation break","references":["https://xenbits.xen.org/xsa/advisory-412.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-43755","cve":"CVE-2022-43755","aliases":[],"title":"Rancher: Insufficient entropy means a leaked cattle-token stays usable after rotation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Rancher","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Insufficient entropy means a leaked cattle-token stays usable after rotation","attack_vector":"An attacker who once observed the token","remediation":"Upgrade Rancher; force full token regeneration","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43755"],"status":"curated"},{"id":"CVE-2022-48797","cve":"CVE-2022-48797","aliases":[],"title":"Linux memory management (NUMA balancing) as used by Gaudi accelerators: Automatic NUMA balancing migrated copy-on-write pages that the Gaudi accelerator still had pinned, silently…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux memory management (NUMA balancing) as used by Gaudi accelerators","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"Automatic NUMA balancing migrated copy-on-write pages that the Gaudi accelerator still had pinned, silently corrupting data in flight to the device. This is a correctness and integrity bug rather than an access-control one, but on a training cluster silent tensor corruption is worse than a crash because it poisons checkpoints before anyone notices.","attack_vector":"No attacker needed - it fires on normal accelerator workloads whenever automatic NUMA balancing is enabled on the node.","remediation":"Take the kernel fix and reboot. As an interim mitigation, disable automatic NUMA balancing (numa_balancing=disable) on accelerator nodes, which is a boot-parameter change and therefore still needs a drain and reboot. No firmware component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48797","https://git.kernel.org/stable/c/254090925e16abd914c87b4ad1b489440d89c4c3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50026","cve":"CVE-2022-50026","aliases":[],"title":"habanalabs kernel driver (Gaudi NIC queue validation): A shift-out-of-bounds in Gaudi NIC queue validation: the driver computes a queue offset for queue types that…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (Gaudi NIC queue validation)","year":"2022","cvss_score":7.1,"severity":"high","kev":false,"impact":"A shift-out-of-bounds in Gaudi NIC queue validation: the driver computes a queue offset for queue types that are not NIC queues, producing undefined behaviour in the kernel. Practical effect is a kernel oops taking the node out of service; on a hardened kernel it is a panic.","attack_vector":"Local user with the habanalabs device node mapped in - i.e. any Gaudi tenant container.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50026","https://git.kernel.org/stable/c/01622098aeb05a5efbb727199bbc2a4653393255"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0180","cve":"CVE-2023-0180","aliases":[],"title":"GPU Display Driver (GeForce/RTX/Quadro/Tesla/vGPU): Info disclosure + DoS (OOB read in kernel driver)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (GeForce/RTX/Quadro/Tesla/vGPU)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Info disclosure + DoS (OOB read in kernel driver)","attack_vector":"Any tenant with GPU device access (local, unprivileged)","remediation":"Upgrade driver to the Jan-2023 branch; drain + reboot node (kernel module reload evicts all GPU workloads)","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0180","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0181","cve":"CVE-2023-0181","aliases":[],"title":"GPU Display Driver: Local privesc / data tampering (missing access control)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc / data tampering (missing access control)","attack_vector":"Any tenant with a container holding /dev/nvidia*","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0181","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-280"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0183","cve":"CVE-2023-0183","aliases":[],"title":"GPU Display Driver: Local privesc (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc (kernel buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0183","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0191","cve":"CVE-2023-0191","aliases":[],"title":"GPU Display Driver: Local privesc / data tampering (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc / data tampering (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0191","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-25516","cve":"CVE-2023-25516","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user triggers an integer overflow in the Linux kernel mode layer, yielding information…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"An unprivileged user triggers an integer overflow in the Linux kernel mode layer, yielding information disclosure and denial of service. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5468. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25516","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-25517","cve":"CVE-2023-25517","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: The vGPU plugin lets a guest OS control resources it is not authorised for, reaching…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The vGPU plugin lets a guest OS control resources it is not authorised for, reaching information disclosure and data tampering on the host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Same authorisation-failure class as CVE-2022-31609 - the guest gets legitimate-looking access to something belonging to the host or another tenant.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5468. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25517","https://github.com/NVIDIA/product-security/tree/main/2023/5468"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-285"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-25518","cve":"CVE-2023-25518","aliases":[],"title":"Jetson AGX Xavier / Xavier NX: Arbitrary memory R/W (PCIe controller without IOMMU)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson AGX Xavier / Xavier NX","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"Arbitrary memory R/W (PCIe controller without IOMMU)","attack_vector":"Local attacker with physical access","remediation":"Flash JetPack 32.7.4+; edge fleet only, not core DC","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5466/5466.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-923"]},{"id":"CVE-2023-31316","cve":"CVE-2023-31316","aliases":[],"title":"AMD Secure Processor - hardware config integrity across power save/restore: MULTI-TENANT ISOLATION: Hardware configuration state is not properly preserved across a power save/restore…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - hardware config integrity across power save/restore","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Hardware configuration state is not properly preserved across a power save/restore cycle in the ASP, so an attacker who can write outside the Trusted Memory Range can change security-relevant configuration that comes back wrong after resume. Suspend/resume is a soft spot in every confidential-computing design; here it is the ASP's own configuration that fails to survive intact.","attack_vector":"Local, requires the ability to write outside the TMR - host-privileged code - and a power-state transition to land on.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Servers rarely suspend, so on a datacenter fleet the practical exposure is lower than the score suggests; still worth catching in the next BIOS wave.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31316","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34338","cve":"CVE-2023-34338","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC TLS certificate / cryptographic keys): A hard-coded certificate and its private key ship inside the firmware, so the same key is on every BMC built…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC TLS certificate / cryptographic keys)","year":"2023","cvss_score":7.1,"severity":"high","kev":false,"impact":"A hard-coded certificate and its private key ship inside the firmware, so the same key is on every BMC built from that image across every customer of that ODM. Anyone who extracts it once - and firmware images are downloadable from vendor support sites - can impersonate any BMC's HTTPS endpoint or decrypt intercepted management traffic. The operator consequence is that TLS on your management plane is decorative: admin passwords, Redfish tokens and KVM sessions are recoverable by an attacker who can interpose.","attack_vector":"Adjacent network with the ability to interpose on BMC traffic and some operator interaction (an admin logging into the BMC). No credentials needed - the attacker supplies the trust. A compromised management jump host, a rogue device on the management VLAN, or an ARP/DHCP position on that segment is enough.","remediation":"Firmware flash to SPx_12.3 / SPx_13.0 or later, but the flash is only half the fix - a fixed image does not retroactively replace a certificate already in place. After flashing you must generate and install a unique per-node BMC certificate signed by your own internal CA, which is a config operation over Redfish and can be automated, no reboot required. Do the certificate rotation even on nodes you cannot yet flash; it is the part that actually removes the shared key.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34338"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-0074","cve":"CVE-2024-0074","aliases":[],"title":"GPU Display Driver: Local privesc / data tampering (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc / data tampering (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0074","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-788"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0128","cve":"CVE-2024-0128","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: The host vGPU plugin lets a guest reach global GPU resources, producing information…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The host vGPU plugin lets a guest reach global GPU resources, producing information disclosure, data tampering and privilege escalation across the partition. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. 'Global resources' on a shared physical GPU means state belonging to other tenants - this is a cross-tenant read/write primitive, not just a host bug.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5586. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0128","https://github.com/NVIDIA/product-security/tree/main/2024/5586"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-732"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0150","cve":"CVE-2024-0150","aliases":[],"title":"GPU Display Driver: Local privesc (kernel driver buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc (kernel driver buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0150","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-21801","cve":"CVE-2024-21801","aliases":[],"title":"Intel TDX module: Insufficient control-flow management in the TDX module lets a privileged host user deny service - i.e. the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Insufficient control-flow management in the TDX module lets a privileged host user deny service - i.e. the host can wedge confidential VMs. Availability rather than confidentiality, but on a confidential-compute product the host being able to kill TDs at will is still a boundary the design claims to hold.","attack_vector":"Privileged host user.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21801","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01070.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-25743","cve":"CVE-2024-25743","aliases":["Heckler-class","interrupt injection"],"title":"SEV-ES / SEV-SNP guest kernel - injection of virtual interrupts 0 and 14: MULTI-TENANT ISOLATION: An untrusted hypervisor can inject virtual interrupts 0 (divide error) and 14 (page…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SEV-ES / SEV-SNP guest kernel - injection of virtual interrupts 0 and 14","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An untrusted hypervisor can inject virtual interrupts 0 (divide error) and 14 (page fault) into an SEV-SNP or SEV-ES guest at arbitrary points, reaching userspace signal handlers - in particular SIGFPE - inside the confidential VM. The Heckler research showed this class turning into authentication bypass inside the guest: inject an interrupt at the right instruction and a login check returns the wrong answer. The host never touches guest memory, so no memory-integrity mechanism catches it.","attack_vector":"Malicious hypervisor against its own guest. Requires timing precision but no guest vulnerability.","remediation":"Fixed in the **guest** kernel, not the host - the hardening lives in the SEV-ES/SNP guest's #VC handler and interrupt entry code. That inverts the usual rollout: you can patch every hypervisor you own and still be exposed, because the protection has to be in the tenant's own VM image. As an operator your job is to ship updated confidential-guest images (or tell tenants which minimum kernel to run) and, where you can, enforce it as an admission requirement. Each guest picks the fix up on its next boot; no host reboot, no firmware update. Fixed by guest-kernel hardening (Linux 6.9+ restricts which interrupts a guest accepts from the hypervisor). Enable Restricted Injection where the platform and guest support it. Again this is a guest-image problem, so operators should publish a minimum kernel and gate confidential workloads on it rather than assuming host patching covers them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25743","https://ahoi-attacks.github.io/heckler/","https://www.amd.com/en/resources/product-security.html"],"status":"curated"},{"id":"CVE-2024-26672","cve":"CVE-2024-26672","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix variable 'mca_funcs' dereferenced before NULL check in 'amdgpu_mca_smu_get_mca_entry()'","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26672","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27029","cve":"CVE-2024-27029","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix mmhub client id out-of-bounds access","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27029","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-32878","cve":"CVE-2024-32878","aliases":[],"title":"llama.cpp (`gguf_init_from_file`): Use of uninitialized heap variable","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (`gguf_init_from_file`)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Use of uninitialized heap variable → double free","attack_vector":"Customer-supplied GGUF file","remediation":"Rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32878"],"status":"curated"},{"id":"CVE-2024-46722","cve":"CVE-2024-46722","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix mc_data out-of-bounds read warning","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46722","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46723","cve":"CVE-2024-46723","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix ucode out-of-bounds read warning","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46723","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46724","cve":"CVE-2024-46724","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu kernel driver core - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix out-of-bounds read of df_v1_7_channel_number","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46724","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46731","cve":"CVE-2024-46731","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/pm: fix the Out-of-bounds read warning","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46731","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46815","cve":"CVE-2024-46815","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check num_valid_sets before accessing reader_wm_sets[]","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46815","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49894","cve":"CVE-2024-49894","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix index out of bounds in degamma hardware format translation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49894","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53150","cve":"CVE-2024-53150","aliases":[],"title":"Linux kernel (ALSA usb-audio): Out-of-bounds read finding clock sources","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ALSA usb-audio)","year":"2024","cvss_score":7.1,"severity":"high","kev":true,"impact":"Out-of-bounds read finding clock sources [KEV]","attack_vector":"Local user with USB device access","remediation":"Livepatchable; otherwise drain + reboot. Blacklist snd-usb-audio on servers","references":["https://access.redhat.com/security/cve/CVE-2024-53150"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-22226","cve":"CVE-2025-22226","aliases":[],"title":"VMware ESXi / Workstation / Fusion: Out-of-bounds read in HGFS leaks vmx process memory to the guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Fusion","year":"2025","cvss_score":7.1,"severity":"high","kev":true,"impact":"Out-of-bounds read in HGFS leaks vmx process memory to the guest [KEV]","attack_vector":"Tenant VM guest (local admin inside the VM)","remediation":"ESXi patch + host reboot with evacuation","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22226"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23270","cve":"CVE-2025-23270","aliases":[],"title":"IGX Orin / Jetson (power mgmt): DoS / hardware damage via power-management abuse","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"IGX Orin / Jetson (power mgmt)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"DoS / hardware damage via power-management abuse","attack_vector":"Local attacker on the device","remediation":"Flash firmware; edge fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23270","https://github.com/NVIDIA/product-security/tree/main/2025/5662"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:H","cwe":["CWE-392"]},{"id":"CVE-2025-23278","cve":"CVE-2025-23278","aliases":[],"title":"GPU Display Driver: Local privesc (OOB write)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"Local privesc (OOB write)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23278","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-129"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23360","cve":"CVE-2025-23360","aliases":[],"title":"NVIDIA NeMo Framework: A relative path traversal gives arbitrary file write, reaching code execution. In an AI datacenter this is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo Framework","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"A relative path traversal gives arbitrary file write, reaching code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5623 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23360","https://github.com/NVIDIA/product-security/tree/main/2025/5623"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-23"]},{"id":"CVE-2025-41239","cve":"CVE-2025-41239","aliases":[],"title":"VMware ESXi / Workstation / Tools: Uninitialised memory in vSockets discloses host memory to the guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi / Workstation / Tools","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"Uninitialised memory in vSockets discloses host memory to the guest","attack_vector":"Tenant VM guest","remediation":"ESXi patch + host reboot; also VMware Tools update in guests","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-41239"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-6202","cve":"CVE-2025-6202","aliases":["Phoenix"],"title":"SK Hynix DDR5 DIMMs (manufactured 2021-01 through 2024-12): Rowhammer bit flips on DDR5, which had been assumed out of reach because of on-die ECC and improved TRR.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SK Hynix DDR5 DIMMs (manufactured 2021-01 through 2024-12)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"Rowhammer bit flips on DDR5, which had been assumed out of reach because of on-die ECC and improved TRR. Affects SK Hynix DDR5 DIMMs produced between January 2021 and December 2024 - a very large share of DDR5 installed in AI host nodes bought in that window. Impact is integrity of host memory, with the usual escalation to privilege via page-table corruption.","attack_vector":"Local attacker on the node. High attack complexity, low privileges - a tenant workload with sustained memory access is the model.","remediation":"Take an inventory of DIMM vendor and date code across the fleet (dmidecode -t memory) before anything else, because the exposure is specific. Mitigation guidance is to raise the DRAM refresh rate - tripling it substantially raises the bar at a measurable memory-bandwidth cost - which is a BIOS-level change requiring drain and reboot per node. There is no microcode or OS patch. Monitor correctable ECC error rates as the detection signal.","references":["https://comsec.ethz.ch/phoenix","https://security.googleblog.com/2025/09/supporting-rowhammer-research-to.html","https://nvd.nist.gov/vuln/detail/CVE-2025-6202"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-6242","cve":"CVE-2025-6242","aliases":[],"title":"vLLM (`MediaConnector` SSRF): SSRF via `load_from_url` in multimodal input handling","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`MediaConnector` SSRF)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"SSRF via `load_from_url` in multimodal input handling","attack_vector":"Unauthenticated request supplying an image URL — reaches cloud metadata endpoints and internal control planes","remediation":"Upgrade and block link-local metadata (169.254.169.254) at the pod network layer. Provider-owned: IMDS reachability from tenant pods is an infrastructure decision","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-6242"],"status":"curated"},{"id":"CVE-2025-66448","cve":"CVE-2025-66448","aliases":[],"title":"vLLM (`Nemotron_Nano_VL_Config`): RCE via a config class evaluated at model load","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`Nemotron_Nano_VL_Config`)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"RCE via a config class evaluated at model load","attack_vector":"Customer-supplied model config","remediation":"Upgrade to 0.11.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-66448"],"status":"curated"},{"id":"CVE-2025-68313","cve":"CVE-2025-68313","aliases":[],"title":"AMD Zen 5 RDSEED (16-bit and 32-bit variants): MULTI-TENANT ISOLATION: On Zen 5, the 16-bit and 32-bit forms of RDSEED return zero far more often than…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zen 5 RDSEED (16-bit and 32-bit variants)","year":"2025","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: On Zen 5, the 16-bit and 32-bit forms of RDSEED return zero far more often than randomness allows, while still setting the carry flag to signal success. Any software that trusts RDSEED's success indicator - kernel entropy pools, TLS libraries, key generation inside confidential guests - silently consumes zeros as if they were seed material. This is a hardware entropy failure that announces nothing; you find it by auditing, not by observing symptoms.","attack_vector":"Not an attack in the usual sense - it is a silicon defect that any workload on affected Zen 5 parts hits passively. The exposure is that an attacker who knows a target derived keys from RDSEED on affected hardware can search a drastically reduced keyspace.","remediation":"Fixed in the Linux kernel (x86/CPU/AMD) by masking off the broken RDSEED variants on affected Zen 5 parts, so the OS stops trusting them and falls back to working entropy sources. Take the distro kernel update and reboot the node; no firmware, microcode or BIOS step. Guest kernels need the same fix, so confidential-VM images must be updated too. Rotate any long-lived key material generated on affected Zen 5 hosts before the fix - the patch stops the bleeding but does not un-weaken existing keys.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68313"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24185","cve":"CVE-2026-24185","aliases":[],"title":"NVIDIA NVOS (network switches): With PKA-only SSH mode enabled, an administrator can inadvertently leave an alternative authentication path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVOS (network switches)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"With PKA-only SSH mode enabled, an administrator can inadvertently leave an alternative authentication path open. If the default password was never changed, that path grants unauthorised switch access - so the switch looks key-only while still accepting a known password.","attack_vector":"Network, adjacent, low privileges. The exposure only exists where the default password survived deployment, which is exactly the switch nobody revisited after racking.","remediation":"Update NVOS per bulletin 5817 and, more importantly, verify that no switch still holds a default password - the configuration audit matters more than the patch here. Cost: switch reboot for the upgrade; the password audit costs nothing and should happen today.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24185","https://github.com/NVIDIA/product-security/tree/main/2026/5817"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-288"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24195","cve":"CVE-2026-24195","aliases":[],"title":"vGPU Manager: Guest-to-host impact via invalid guest-driver input","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Guest-to-host impact via invalid guest-driver input","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate guest VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24195","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-20"]},{"id":"CVE-2026-24196","cve":"CVE-2026-24196","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB read of graphics memory)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Info disclosure (OOB read of graphics memory)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24196","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24779","cve":"CVE-2026-24779","aliases":[],"title":"vLLM (`MediaConnector`): SSRF, recurrence of CVE-2025-6242","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`MediaConnector`)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"SSRF, recurrence of CVE-2025-6242","attack_vector":"Unauthenticated request with an attacker-supplied media URL","remediation":"Upgrade to 0.14.1+; enforce IMDSv2 / metadata deny at the network layer","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24779"],"status":"curated"},{"id":"CVE-2026-25960","cve":"CVE-2026-25960","aliases":[],"title":"vLLM (`load_from_url_async`): Bypass of the CVE-2026-24779 SSRF fix","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`load_from_url_async`)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Bypass of the CVE-2026-24779 SSRF fix","attack_vector":"Unauthenticated request","remediation":"Upgrade past 0.15.1. Third SSRF in the same connector — network-layer egress control is the only durable fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-25960"],"status":"curated"},{"id":"CVE-2026-31395","cve":"CVE-2026-31395","aliases":[],"title":"Linux bnxt_en driver (DBG_BUF_PRODUCER async event handler): The async-event handler indexes a fixed array with a `type` field supplied by the NIC firmware, without…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (DBG_BUF_PRODUCER async event handler)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"The async-event handler indexes a fixed array with a `type` field supplied by the NIC firmware, without bounds checking — so firmware controls a kernel array index. This is the concrete version of a threat operators often wave at abstractly: if the adapter's firmware is compromised or buggy, it has a direct path into kernel memory corruption on the host. Every argument for verifying NIC firmware provenance at intake rests on bugs of exactly this shape.","attack_vector":"The NIC firmware itself, or anything that can influence what the firmware reports — which includes a firmware image installed at build time or by a previous tenant on bare metal.","remediation":"Kernel/driver upgrade plus host reboot. The durable control is separate: verify and reflash NIC firmware from a known-good image at rack intake and at tenant handoff, so the host is not trusting whatever firmware happens to be on the card.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31395"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-31766","cve":"CVE-2026-31766","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu user-mode queues (doorbell submission path) - a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu user-mode queues (doorbell submission path) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: validate doorbell_offset in user queue creation","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31766","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-35155","cve":"CVE-2026-35155","aliases":["DSA-2026-187"],"title":"Dell iDRAC10 (credential handling, race condition): A race in iDRAC10's credential handling leaves secrets insufficiently protected, letting an authenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC10 (credential handling, race condition)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A race in iDRAC10's credential handling leaves secrets insufficiently protected, letting an authenticated low-privilege user win the race and come back with elevated access to the BMC. Elevated on the BMC means the whole out-of-band toolkit: power, Virtual Media, KVM, firmware. This one matters disproportionately because iDRAC10 ships on 17G PowerEdge - the newest GPU platforms - so the affected fleet is the freshly-racked capacity, not the legacy tier, and it is likely still inside its burn-in window where firmware is whatever shipped from the factory.","attack_vector":"An authenticated low-privilege iDRAC10 account. The exposure is anyone you have handed a non-admin BMC login: remote hands, an integrator, a monitoring service account, or a tenant-facing self-service console that proxies BMC actions.","remediation":"Flash iDRAC10 to 1.30.10.50 or later. Affected builds are 1.20.70.50 and 1.30.05.10 specifically. Out-of-band, per-node, no host reboot and no job drain. Because this hits new deployments, fold the check into rack-acceptance: verify iDRAC10 build before a node ever takes tenant traffic, rather than discovering it in a later sweep. No config-only mitigation - reduce exposure meanwhile by pruning low-privilege iDRAC accounts.","references":["https://www.dell.com/support/kbdoc/en-us/000452298/dsa-2026-187-security-update-for-dell-idrac10-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2026-35155"],"status":"curated"},{"id":"CVE-2026-45856","cve":"CVE-2026-45856","aliases":[],"title":"Linux kernel InfiniBand core (ib_uverbs post_send): ib_uverbs_post_send() takes the work-queue-entry size straight from userspace with no validation, allocates…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel InfiniBand core (ib_uverbs post_send)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"ib_uverbs_post_send() takes the work-queue-entry size straight from userspace with no validation, allocates that size, then reads fields past the allocation - an out-of-bounds read of the kernel heap that leaks kernel memory to an unprivileged process. The receive path validated this; the send path did not. Any tenant process holding an RDMA verbs handle - which on a GPU cluster is every job using NCCL, UCX or MPI - can read host kernel memory.","attack_vector":"A local unprivileged user with an open RDMA verbs device handle. On a shared GPU node that is any tenant running a distributed training job.","remediation":"Upgrade the host kernel to 7.0 or a stable backport (5.10.252, 5.15.202, 6.1.165, 6.6.128, 6.12.75, 6.18.14, 6.19.4). Rolling reboot of every node that exposes /dev/infiniband/uverbs* to workloads - which is all of them on an RDMA cluster.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45856","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-45856.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46199","cve":"CVE-2026-46199","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vcn4: Prevent OOB reads when parsing dec msg","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46199","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46204","cve":"CVE-2026-46204","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn4)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vcn4: Prevent OOB reads when parsing IB","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46204","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46218","cve":"CVE-2026-46218","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Add bounds checking to ib_{get,set}_value","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46218","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46230","cve":"CVE-2026-46230","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn3): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/vcn3)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/vcn3: Prevent OOB reads when parsing dec msg","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46230","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53138","cve":"CVE-2026-53138","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Bound VBIOS record-chain walk loops","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53138","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53330","cve":"CVE-2026-53330","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix out-of-bounds read in dp_get_eq_aux_rd_interval()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53330","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53875","cve":"CVE-2026-53875","aliases":[],"title":"picklescan: `scan_pytorch` bypass via forged magic numbers","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"`scan_pytorch` bypass via forged magic numbers","attack_vector":"Customer-supplied model file","remediation":"Upgrade to 1.0.3+. There is no complete fix for pickle scanning — this is the fifth bypass in the same tool","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53875"],"status":"curated"},{"id":"CVE-2026-64172","cve":"CVE-2026-64172","aliases":["Erratum 1235"],"title":"Linux KVM/SVM - AVIC IPI virtualization on Hygon Family 18h: MULTI-TENANT ISOLATION: AVIC inter-processor-interrupt virtualization is unsafe on Hygon Family 18h parts…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM/SVM - AVIC IPI virtualization on Hygon Family 18h","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: AVIC inter-processor-interrupt virtualization is unsafe on Hygon Family 18h parts, which are derived from AMD Family 17h. With AVIC active, a guest's IPIs can be delivered incorrectly - interrupt delivery reaching the wrong target is a cross-VM correctness failure on the interrupt path, which is the same seam the Heckler-class attacks exploit.","attack_vector":"From inside a guest VM on affected Hygon silicon with AVIC enabled.","remediation":"Fixed in the Linux kernel by disabling AVIC IPI virtualization on affected parts. Distro kernel update plus host reboot. Interim mitigation: disable AVIC in the kvm_amd module parameters (avic=0), which costs some interrupt-heavy performance but needs only a module reload with guests drained. Relevant only if you have Hygon parts in the fleet - worth checking, since they show up in some regional supply chains.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64172"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-65918","cve":"CVE-2026-65918","aliases":[],"title":"torchvision (GIF decoder): Out-of-bounds heap read in `read_from_tensor` GIF decode","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"torchvision (GIF decoder)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"Out-of-bounds heap read in `read_from_tensor` GIF decode","attack_vector":"Customer-supplied image data reaching a vision preprocessing pipeline","remediation":"Rebuild images with torchvision > 0.28.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65918"],"status":"curated"},{"id":"CVE-2026-68103","cve":"CVE-2026-68103","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): A correctness defect in the amdgpu user-mode queues (doorbell submission path) reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu user-mode queues (doorbell submission path) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: reject mapping a reserved doorbell to a new queue","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68103","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68258","cve":"CVE-2026-68258","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Check bounds on CRIU restore queue type and mqd size","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68258","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68293","cve":"CVE-2026-68293","aliases":["net/mlx5 MCIA register buffer overflow on 32 dword reads"],"title":"Linux kernel mlx5_core port / transceiver module EEPROM (MCIA register): The MCIA register can return 32 dwords when the device advertises the capability, but the kernel's structure…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_core port / transceiver module EEPROM (MCIA register)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"The MCIA register can return 32 dwords when the device advertises the capability, but the kernel's structure defines only 12, so reading module EEPROM copies past the end of the buffer. In practice this panics the host when an operator or monitoring agent runs ethtool against an optical module - so routine transceiver telemetry becomes a way to take a node down, and any agent that polls optics fleet-wide becomes a fleet-wide outage trigger.","attack_vector":"Local, low-privileged in the sense that the read is triggered from ethtool module-EEPROM queries; the overflowing length comes from what the adapter firmware advertises.","remediation":"Upgrade the host kernel to 7.2 or a stable backport (6.12.101, 6.18.42, 7.1.6). Rolling reboot. Interim: stop automated optics/EEPROM polling on affected kernels - a monitoring config change, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68293","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-68293.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68447","cve":"CVE-2026-68447","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":7.1,"severity":"high","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: clamp v9 CRIU control stack checkpoint copy to BO size","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68447","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-11983","cve":"CVE-2019-11983","aliases":["HPESBHF03917"],"title":"HPE iLO 4 / iLO 5 (remote buffer overflow): Remotely triggerable buffer overflow in the iLO firmware on both the Gen9 (iLO 4) and Gen10 (iLO 5)…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 4 / iLO 5 (remote buffer overflow)","year":"2019","cvss_score":7,"severity":"high","kev":false,"impact":"Remotely triggerable buffer overflow in the iLO firmware on both the Gen9 (iLO 4) and Gen10 (iLO 5) generations. Memory corruption on a service processor is the highest-value bug class in a fleet, because success means code on the BMC and therefore power control, Virtual Media, console, and an implant that persists across host reinstalls. Worth noting alongside the older iLO 4 authentication-bypass work that made this platform a known research target - the Gen9 tier tends to be the part of a fleet that stopped receiving attention.","attack_vector":"Reachable over the network to the iLO address on the out-of-band management VLAN.","remediation":"Flash iLO 4 to v2.61b or later and iLO 5 to v1.39 or later. Out-of-band, per-node, no host reboot and no job drain. On Gen9 hardware the practical obstacle is inventory and change control rather than the flash itself - these are usually the nodes with the least recent firmware campaign. Interim control: hard-ACL the iLO management network.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03917en_us","https://nvd.nist.gov/vuln/detail/CVE-2019-11983"],"status":"curated"},{"id":"CVE-2019-19921","cve":"CVE-2019-19921","aliases":[],"title":"runc: Volume-mount race gives incorrect access control and privilege escalation to host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2019","cvss_score":7,"severity":"high","kev":false,"impact":"Volume-mount race gives incorrect access control and privilege escalation to host","attack_vector":"Any tenant workload able to spawn two containers with crafted mounts","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19921"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-1099","cve":"CVE-2021-1099","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: stack buffer overflow in the vGPU Manager with enough control for a guest to place a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: stack buffer overflow in the vGPU Manager with enough control for a guest to place a ROP chain on the host stack. This is the most explicitly weaponisable of the 2021 vGPU set - guest-to-host code execution on the hypervisor. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU on the affected host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1099"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1120","cve":"CVE-2021-1120","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: a guest-supplied string may not be null-terminated, and the host plugin reads past it…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: a guest-supplied string may not be null-terminated, and the host plugin reads past it - disclosure, tampering or unauthorised code execution on the host, though NVIDIA notes the guest cannot choose the content it pushes.","attack_vector":"Any user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1120"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-20188","cve":"CVE-2021-20188","aliases":[],"title":"Podman: File permissions not checked for non-root users in a privileged container","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"File permissions not checked for non-root users in a privileged container; cross-user file access","attack_vector":"Any tenant workload in a privileged container","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20188"],"status":"curated"},{"id":"CVE-2021-22543","cve":"CVE-2021-22543","aliases":[],"title":"KVM: Improper handling of VM_IO/VM_PFNMAP vmas in KVM lets a guest bypass read-only checks - host privilege…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"Improper handling of VM_IO/VM_PFNMAP vmas in KVM lets a guest bypass read-only checks - host privilege escalation","attack_vector":"Tenant VM guest / local user with /dev/kvm","remediation":"Kernel patch. Livepatchable on some vendors; KVM module changes often are not - budget drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-22543"],"status":"curated"},{"id":"CVE-2021-22600","cve":"CVE-2021-22600","aliases":[],"title":"Linux kernel (af_packet): Double free in packet_set_ring(), local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (af_packet)","year":"2021","cvss_score":7,"severity":"high","kev":true,"impact":"Double free in packet_set_ring(), local privilege escalation [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_RAW","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2021-22600"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-4204","cve":"CVE-2021-4204","aliases":[],"title":"Linux kernel (eBPF): eBPF improper input validation leading to local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF)","year":"2021","cvss_score":7,"severity":"high","kev":false,"impact":"eBPF improper input validation leading to local privilege escalation","attack_vector":"Any tenant process in a container with BPF access","remediation":"Livepatchable; otherwise drain + reboot. Set `kernel.unprivileged_bpf_disabled=1`","references":["https://access.redhat.com/security/cve/CVE-2021-4204"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-0492","cve":"CVE-2022-0492","aliases":[],"title":"Linux kernel (cgroups v1): cgroups v1 release_agent lets a container with CAP_SYS_ADMIN (or an unconfined userns) run arbitrary…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (cgroups v1)","year":"2022","cvss_score":7,"severity":"high","kev":true,"impact":"cgroups v1 release_agent lets a container with CAP_SYS_ADMIN (or an unconfined userns) run arbitrary host-root commands [KEV]","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot. Strong compensating controls: cgroup v2 only, seccomp/AppArmor default profiles, drop CAP_SYS_ADMIN","references":["https://access.redhat.com/security/cve/CVE-2022-0492"],"status":"curated","fleet":{"ubiquity":"Universal - the `release_agent` path exists on every pre-5.17 kernel; cgroups v1 was still the default on most GPU host images","remediation_pain":"`node-reboot` - kernel upgrade; the only mitigation is relying on AppArmor/SELinux/seccomp being correctly applied, which is exactly what privileged AI workloads often disable","pain_class":"node-reboot","why_fleet_wide":"A root-in-container process abuses user namespaces to mount cgroups v1, sets `release_agent` and runs arbitrary commands as host root; GPU containers routinely run privileged or with `--cap-add`, which removes the default hardening that would have blocked it"}},{"id":"CVE-2022-2602","cve":"CVE-2022-2602","aliases":[],"title":"Linux kernel (io_uring): Use-after-free between io_uring and the unix GC - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (io_uring)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Use-after-free between io_uring and the unix GC - local root","attack_vector":"Any tenant process in a container with io_uring enabled","remediation":"Livepatchable; otherwise drain + reboot. Durable control: block io_uring via seccomp in the default container profile","references":["https://access.redhat.com/security/cve/CVE-2022-2602"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-31614","cve":"CVE-2022-31614","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: The vGPU plugin double-frees host resources","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: The vGPU plugin double-frees host resources; chained with another bug it reaches code execution on the hypervisor host. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5383. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31614","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-32469","cve":"CVE-2022-32469","aliases":["INSYDE-SA-2023001"],"title":"Insyde InsydeH2O (PnpSmm shared SMM/non-SMM buffer, DMA TOCTOU): A buffer shared between SMM and non-SMM code in the plug-and-play driver can be rewritten by DMA between…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm shared SMM/non-SMM buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"A buffer shared between SMM and non-SMM code in the plug-and-play driver can be rewritten by DMA between validation and use, corrupting SMRAM and escalating privilege. PnpSmm owns SMBIOS/platform description, so alongside ring -2 the attacker can poison the hardware inventory your fleet tooling trusts.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32469","https://www.insyde.com/security-pledge/SA-2023001"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32470","cve":"CVE-2022-32470","aliases":["INSYDE-SA-2023002"],"title":"Insyde InsydeH2O (FwBlockServiceSmm shared buffer, DMA TOCTOU): The firmware block service's shared buffer is racy, so DMA lands an attacker on the SPI flash write path with…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FwBlockServiceSmm shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The firmware block service's shared buffer is racy, so DMA lands an attacker on the SPI flash write path with SMRAM corruption alongside it. Insyde's own suggested fix - copy the firmware block services data into SMRAM before checking it - tells you the shape of the bug: validation happening on memory the attacker still owns.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Confirm SPI flash write protection is enforced in the interim. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32470","https://www.insyde.com/security-pledge/SA-2023002"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32471","cve":"CVE-2022-32471","aliases":["INSYDE-SA-2023003"],"title":"Insyde InsydeH2O (IhisiSmm / IhisiDxe command buffer): One representative of a family of roughly a dozen Insyde advisories covering the same defect across different…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IhisiSmm / IhisiDxe command buffer)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"One representative of a family of roughly a dozen Insyde advisories covering the same defect across different drivers: the SMI handler validates its parameters in a buffer that sits outside SMRAM, then uses them - and a DMA-capable device can rewrite that buffer in between. The result is SMRAM corruption and privilege escalation to ring -2 driven from a peripheral rather than from the CPU. In a GPU chassis the DMA-capable devices are GPUs, NICs and NVMe drives, several of which run tenant-flashable firmware, so this is not a theoretical adversary.","attack_vector":"An attacker with control of a DMA-capable device on the node - a compromised NIC or GPU firmware, a malicious PCIe device, or a tenant who can drive DMA from a passed-through device - racing the SMI handler. Does not require host root, which is what makes the family interesting.","remediation":"OEM BIOS update on the fixed Insyde kernel. Firmware flash, reboot per node. The real compensating control here is the IOMMU: enable VT-d/AMD-Vi and DMA protection (including pre-boot DMA protection where the platform supports it) so untrusted devices cannot reach arbitrary host memory. That is a config change and should be standard on any multi-tenant GPU node regardless of this CVE. Note the sibling advisories SA-2023001 through SA-2023015 cover the same bug in PnpSmm, FwBlockServiceSmm, HddPassword, AhciBusDxe, IdeBusDxe, NvmExpressDxe, SdHostDriver, SdMmcDevice, StorageSecurityCommandDxe, VariableRuntimeDxe and FvbServicesRuntimeDxe - patching one does not patch the rest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32471","https://www.insyde.com/security-pledge/SA-2023003"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32473","cve":"CVE-2022-32473","aliases":["INSYDE-SA-2023005"],"title":"Insyde InsydeH2O (HddPassword shared buffer, DMA TOCTOU): Racy shared buffer in the ATA security driver. What this one reaches is drive locking and unlock credentials…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (HddPassword shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Racy shared buffer in the ATA security driver. What this one reaches is drive locking and unlock credentials, so a successful race yields ring -2 plus a foothold in the code holding drive secrets - which matters for any operator using ATA drive locks as part of the between-tenant wipe.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Do not rely on ATA HDD passwords as the at-rest control on unpatched nodes. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32473","https://www.insyde.com/security-pledge/SA-2023005"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32474","cve":"CVE-2022-32474","aliases":["INSYDE-SA-2023006"],"title":"Insyde InsydeH2O (StorageSecurityCommandDxe shared buffer, DMA TOCTOU): The TCG/Opal security-command driver shares a buffer between SMM and non-SMM code without protecting it from…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (StorageSecurityCommandDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The TCG/Opal security-command driver shares a buffer between SMM and non-SMM code without protecting it from DMA. The reachable target is self-encrypting-drive authentication - the mechanism a GPU cloud points at when asked how tenant data is protected at rest between leases.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32474","https://www.insyde.com/security-pledge/SA-2023006"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32475","cve":"CVE-2022-32475","aliases":["INSYDE-SA-2023007"],"title":"Insyde InsydeH2O (VariableRuntimeDxe shared buffer, DMA TOCTOU): Racy shared buffer in the UEFI variable driver - the store holding the Secure Boot key databases…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (VariableRuntimeDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Racy shared buffer in the UEFI variable driver - the store holding the Secure Boot key databases (PK/KEK/db/dbx). Corrupting SMRAM through this driver attacks the arbiter of what firmware and bootloaders are permitted to run, so the loss is the boot-integrity guarantee itself rather than any one workload. Insyde notes the fix also hardened chipset and OEM chipset code, which means OEM-specific builds needed their own rebase on top of the kernel fix.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32475","https://www.insyde.com/security-pledge/SA-2023007"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32476","cve":"CVE-2022-32476","aliases":["INSYDE-SA-2023008"],"title":"Insyde InsydeH2O (AhciBusDxe shared buffer, DMA TOCTOU): DMA race on the SATA/AHCI driver's shared buffer produces SMRAM corruption and privilege escalation. Reaches…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AhciBusDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SATA/AHCI driver's shared buffer produces SMRAM corruption and privilege escalation. Reaches the SATA storage path, typically the node's boot device - firmware persistence plus a position on the disk the node starts from.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32476","https://www.insyde.com/security-pledge/SA-2023008"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32477","cve":"CVE-2022-32477","aliases":["INSYDE-SA-2023009"],"title":"Insyde InsydeH2O (FvbServicesRuntimeDxe shared buffer, DMA TOCTOU): Firmware Volume Block services again, this time via the shared SMM/non-SMM buffer. Same consequence as its…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FvbServicesRuntimeDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Firmware Volume Block services again, this time via the shared SMM/non-SMM buffer. Same consequence as its 2022 twin: an attacker influencing SPI flash writes, which is the difference between a node you can clean by reimaging and a node you have to physically reflash before it goes back in the pool.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Verify SPI write protection is set. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32477","https://www.insyde.com/security-pledge/SA-2023009"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32478","cve":"CVE-2022-32478","aliases":["INSYDE-SA-2023010"],"title":"Insyde InsydeH2O (IdeBusDxe shared buffer, DMA TOCTOU): Racy shared buffer in the legacy IDE/ATA driver leading to SMRAM corruption. As with its 2022 counterpart…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IdeBusDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Racy shared buffer in the legacy IDE/ATA driver leading to SMRAM corruption. As with its 2022 counterpart, this driver is generally only active where CSM/legacy storage compatibility is enabled, so a UEFI-only server profile may not expose it at all.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory. Disable CSM / legacy storage on UEFI-only nodes to remove the driver rather than patch it. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32478","https://www.insyde.com/security-pledge/SA-2023010"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32953","cve":"CVE-2022-32953","aliases":["INSYDE-SA-2023013"],"title":"Insyde InsydeH2O (SdHostDriver shared buffer, DMA TOCTOU): DMA race on the SD host controller's shared buffer. Insyde's remediation note is more specific for this…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdHostDriver shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SD host controller's shared buffer. Insyde's remediation note is more specific for this sub-group - copy the link data into SMRAM before checking it AND verify every pointer falls inside the buffer - which is worth knowing because it indicates the earlier fixes in this driver were incomplete rather than wrong.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32953","https://www.insyde.com/security-pledge/SA-2023013"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32954","cve":"CVE-2022-32954","aliases":["INSYDE-SA-2023014"],"title":"Insyde InsydeH2O (SdMmcDevice shared buffer, DMA TOCTOU): Shared-buffer DMA race in the SD/MMC device layer, kernel 5.1 through 5.5. Pairs with the SdHostDriver entry…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdMmcDevice shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"Shared-buffer DMA race in the SD/MMC device layer, kernel 5.1 through 5.5. Pairs with the SdHostDriver entry - Insyde consistently files the controller and device layers separately, so both need to be present in whatever BIOS release you accept as the fix.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32954","https://www.insyde.com/security-pledge/SA-2023014"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32955","cve":"CVE-2022-32955","aliases":["INSYDE-SA-2023015"],"title":"Insyde InsydeH2O (NvmExpressDxe shared buffer, DMA TOCTOU): The NVMe driver's SMM/non-SMM shared buffer is racy, giving SMRAM corruption and ring -2 escalation. On a GPU…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressDxe shared buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The NVMe driver's SMM/non-SMM shared buffer is racy, giving SMRAM corruption and ring -2 escalation. On a GPU node this is the driver sitting on the datasets, checkpoints and weights, and it is the second separately-filed NVMe DMA defect after SA-2022055 - strong evidence that this driver deserves standing attention in a firmware patch policy rather than case-by-case triage.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Insyde lists kernel 5.0 through 5.5 affected; take the per-kernel fixed version from the advisory.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. This is Insyde's second pass at the same defect class in a different set of buffers - a fleet that took the 2022 BIOS release is NOT covered for this batch, and OEM release notes rarely make that distinction clear. Verify by kernel version, not by 'we patched the Insyde DMA bugs'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32955","https://www.insyde.com/security-pledge/SA-2023015"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33905","cve":"CVE-2022-33905","aliases":["INSYDE-SA-2022047"],"title":"Insyde InsydeH2O (AhciBusDxe SMI input buffer, DMA TOCTOU): DMA race on the SATA/AHCI controller driver's SMI input buffer yields SMRAM corruption and escalation to ring…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (AhciBusDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SATA/AHCI controller driver's SMI input buffer yields SMRAM corruption and escalation to ring -2. What this driver reaches is the SATA storage path - on a GPU node that is typically the boot drive or a bulk data volume, so an attacker lands both firmware persistence and a position astride the disk the node boots from.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33905","https://www.insyde.com/security-pledge/SA-2022047"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33908","cve":"CVE-2022-33908","aliases":["INSYDE-SA-2022050"],"title":"Insyde InsydeH2O (SdHostDriver SMI input buffer, DMA TOCTOU): DMA race on the SD host controller driver gives SMRAM corruption and ring -2 escalation. On server hardware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdHostDriver SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the SD host controller driver gives SMRAM corruption and ring -2 escalation. On server hardware the SD/eMMC controller is often wired to platform or BMC-adjacent boot media rather than to anything a tenant uses, which makes this an easy driver to forget about - and an attractive one for an attacker, because nobody is watching it.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25. Where the platform exposes it, disabling the unused SD/eMMC controller in BIOS removes the surface. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33908","https://www.insyde.com/security-pledge/SA-2022050"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33909","cve":"CVE-2022-33909","aliases":["INSYDE-SA-2022051"],"title":"Insyde InsydeH2O (HddPassword SMI input buffer, DMA TOCTOU): The HddPassword driver handles ATA security - drive locking and unlock credentials. A DMA race here corrupts…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (HddPassword SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"The HddPassword driver handles ATA security - drive locking and unlock credentials. A DMA race here corrupts SMRAM and puts the attacker inside the code that holds drive-unlock secrets in memory, so the reachable prize is not just ring -2 but the credentials protecting the drive. Relevant to any fleet that leans on ATA drive locking as part of its between-tenant wipe story.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23. Do not treat ATA HDD passwords as the confidentiality control on unpatched nodes - prefer SED keys or software FDE with keys held off the node. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33909","https://www.insyde.com/security-pledge/SA-2022051"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33983","cve":"CVE-2022-33983","aliases":["INSYDE-SA-2022053"],"title":"Insyde InsydeH2O (NvmExpressLegacy SMI input buffer, DMA TOCTOU): DMA race on the legacy NVMe SMI handler. NVMe is the data path on a GPU node - datasets, checkpoints, model…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressLegacy SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the legacy NVMe SMI handler. NVMe is the data path on a GPU node - datasets, checkpoints, model weights - so a driver that reaches NVMe reaches tenant data as well as SMRAM. This is the legacy-path twin of the NvmExpressDxe issue filed as SA-2022055; both ship, and patching one does not patch the other.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33983","https://www.insyde.com/security-pledge/SA-2022053"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33984","cve":"CVE-2022-33984","aliases":["INSYDE-SA-2022054"],"title":"Insyde InsydeH2O (SdMmcDevice SMI input buffer, DMA TOCTOU): SMRAM corruption through a DMA race on the SD/MMC device driver. Pairs with the SdHostDriver issue…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SdMmcDevice SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"SMRAM corruption through a DMA race on the SD/MMC device driver. Pairs with the SdHostDriver issue (SA-2022050) - Insyde filed the controller and the device layer separately, so a fleet that patched one BIOS release for 'the SD bug' may still be carrying the other.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33984","https://www.insyde.com/security-pledge/SA-2022054"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33985","cve":"CVE-2022-33985","aliases":["INSYDE-SA-2022055"],"title":"Insyde InsydeH2O (NvmExpressDxe SMI input buffer, DMA TOCTOU): DMA race on the primary NVMe driver's SMI input buffer gives SMRAM corruption and ring -2 escalation. This is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (NvmExpressDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"DMA race on the primary NVMe driver's SMI input buffer gives SMRAM corruption and ring -2 escalation. This is the one that matters most on a modern GPU server: NVMe is where the training data, checkpoints and weights live, and the same driver that touches them is the one exposing a racy SMI handler.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33985","https://www.insyde.com/security-pledge/SA-2022055"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-50079","cve":"CVE-2022-50079","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Check correct bounds for stream encoder instances for DCN303","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50079","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0386","cve":"CVE-2023-0386","aliases":[],"title":"Linux kernel (OverlayFS/FUSE): OverlayFS copies setuid files from a nosuid FUSE mount - unprivileged local user to root, a live…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (OverlayFS/FUSE)","year":"2023","cvss_score":7,"severity":"high","kev":true,"impact":"OverlayFS copies setuid files from a nosuid FUSE mount - unprivileged local user to root, a live container-escape chain [KEV]","attack_vector":"Any tenant process in a container with a user namespace","remediation":"Livepatchable; otherwise drain + reboot. Compensating control: disallow unprivileged FUSE mounts","references":["https://access.redhat.com/security/cve/CVE-2023-0386"],"status":"curated","fleet":{"ubiquity":"Universal - kernels 5.11-6.1.8, and OverlayFS *is* the container storage driver on every containerized GPU host","remediation_pain":"`node-reboot` - kernel upgrade; no meaningful runtime mitigation since disabling OverlayFS breaks container storage","pain_class":"node-reboot","why_fleet_wide":"Copying a setuid binary across a `nosuid` OverlayFS mount preserves capabilities, giving any regular user root; inside a container it gives container-root, which then chains into the cgroup/procfs escapes above on a shared GPU host"}},{"id":"CVE-2023-27561","cve":"CVE-2023-27561","aliases":[],"title":"runc: Regression of CVE-2019-19921: incorrect access control leading to privilege escalation via volume mounts","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Regression of CVE-2019-19921: incorrect access control leading to privilege escalation via volume mounts","attack_vector":"Any tenant workload able to spawn two containers with custom mounts","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27561"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-41914","cve":"CVE-2023-41914","aliases":[],"title":"Slurm: Filesystem race conditions allow gaining ownership of, overwriting, or deleting files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Filesystem race conditions allow gaining ownership of, overwriting, or deleting files","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-41914"],"status":"curated","fleet":{"ubiquity":"Very common - Slurm 22.05.x / 23.02.x branches","remediation_pain":"`daemon-restart` - upgrade slurmd/slurmctld on every node; running jobs survive a slurmd restart in many configs, so this is cheaper than a runc bug but still fleet-wide","pain_class":"daemon-restart","why_fleet_wide":"TOCTOU race in Slurm file handling lets a low-privilege job take ownership of, overwrite or delete arbitrary files, escalating on any compute node the scheduler touches"}},{"id":"CVE-2023-46813","cve":"CVE-2023-46813","aliases":[],"title":"Linux kernel SEV-ES #VC handler - MMIO access checking: MULTI-TENANT ISOLATION: Incorrect access checking in the SEV-ES #VC handler and instruction emulation lets a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel SEV-ES #VC handler - MMIO access checking","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Incorrect access checking in the SEV-ES #VC handler and instruction emulation lets a local user with userspace access to MMIO registers escalate. Inside a confidential VM, a merely local user reaches privileged guest state through the exception handler that SEV-ES uses to virtualise MMIO - so the confidential VM's own internal privilege boundary breaks, not just the host/guest one.","attack_vector":"Local, from userspace inside an SEV-ES guest that has userspace-accessible MMIO. Affects Linux before 6.5.9.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. The fix belongs in the **guest** kernel, so update your confidential-VM images (or publish a minimum guest kernel to tenants) rather than assuming host patching covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46813"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-51767","cve":"CVE-2023-51767","aliases":["OpenSSH single-bit auth bypass","Rowhammer-assisted authentication bypass"],"title":"OpenSSH through 10.0 - mm_answer_authpassword uses an integer 'authenticated' flag that does not resist a single bit flip; the same class was fixed in sudo as CVE-2023-42465: Shows the second half…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenSSH through 10.0 - mm_answer_authpassword uses an integer 'authenticated' flag that does not resist a single bit…","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Shows the second half of the Rowhammer chain that operators usually skip: the flip has to land somewhere useful, and a lot of privileged software stores 'is this user authenticated' as a plain integer that a one-bit change turns from false to true. A co-resident attacker who can hammer the sshd privilege-separation monitor's memory authenticates without credentials. Sudo had the identical pattern. For a bare-metal or shared-host fleet, this converts 'Rowhammer is a research curiosity' into 'a co-tenant logs in as root on the host'.","attack_vector":"Requires attacker-victim co-location - unprivileged local code on the same machine as the sshd being attacked, with hammerable DRAM. It is a threat model that only exists on shared hosts, which is exactly the neocloud and multi-tenant cluster model.","remediation":"Patch OpenSSH and sudo to versions with the hardened comparisons (sudo 1.9.15 and later; OpenSSH per your distribution's backport) - that is a normal package update with an sshd restart, no reboot. But understand what it buys you: it removes two known landing spots, not the underlying ability to flip bits. Treat this as a prompt to audit other privileged daemons in your control plane for single-integer authorisation flags, and to fix the DRAM-sharing policy underneath.","references":["https://arxiv.org/abs/2309.02545","https://nvd.nist.gov/vuln/detail/CVE-2023-51767","https://nvd.nist.gov/vuln/detail/CVE-2023-42465"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2023-6931","cve":"CVE-2023-6931","aliases":[],"title":"Linux kernel (perf): Out-of-bounds write in perf_read_group() via read_size overflow - local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (perf)","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Out-of-bounds write in perf_read_group() via read_size overflow - local root","attack_vector":"Any tenant process in a container with perf access","remediation":"Livepatchable; otherwise drain + reboot. Set `kernel.perf_event_paranoid=3` - but note profiling access is a feature many AI tenants expect","references":["https://access.redhat.com/security/cve/CVE-2023-6931"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-6932","cve":"CVE-2023-6932","aliases":[],"title":"Linux kernel (IGMP): Use-after-free in IPv4 IGMP - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (IGMP)","year":"2023","cvss_score":7,"severity":"high","kev":false,"impact":"Use-after-free in IPv4 IGMP - local privilege escalation","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2023-6932"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0646","cve":"CVE-2024-0646","aliases":[],"title":"Linux kernel (kTLS): splice() into a kTLS socket overwrites read-only kernel pages - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (kTLS)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"splice() into a kTLS socket overwrites read-only kernel pages - local privilege escalation","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-0646"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26660","cve":"CVE-2024-26660","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Implement bounds check for stream encoder creation in DCN301","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26660","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27045","cve":"CVE-2024-27045","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix a potential buffer overflow in 'dp_dsc_clock_en_read()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27045","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27134","cve":"CVE-2024-27134","aliases":[],"title":"MLflow (`spark_udf` dir perms): Excessive directory permissions","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"MLflow (`spark_udf` dir perms)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Excessive directory permissions → local privilege escalation","attack_vector":"Co-tenant local user on the same node","remediation":"Upgrade; matters on shared bare-metal nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27134"],"status":"curated"},{"id":"CVE-2024-36334","cve":"CVE-2024-36334","aliases":[],"title":"AMD Radeon RGB tool - signature verification on files in the installation directory: The Radeon RGB tool does not verify signatures on files placed in its installation directory, so a planted…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Radeon RGB tool - signature verification on files in the installation directory","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"The Radeon RGB tool does not verify signatures on files placed in its installation directory, so a planted file runs with elevated privileges. A cosmetic utility that escalates to code execution - the reason it appears here is that vendor GPU tooling gets installed wholesale on GPU hosts without anyone asking what the LED control daemon is doing running as root.","attack_vector":"Local, requires write access to the tool's installation directory.","remediation":"Update or, better, uninstall - RGB lighting control has no business on a datacenter GPU node. Removing unnecessary vendor tooling from the golden image is the durable fix and costs nothing at runtime. No reboot needed to uninstall.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36334","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-39766","cve":"CVE-2024-39766","aliases":[],"title":"Intel Neural Compressor (SQL injection, second instance): A second SQL-injection path in Neural Compressor reachable by an authenticated user. Same consequence as the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Neural Compressor (SQL injection, second instance)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"A second SQL-injection path in Neural Compressor reachable by an authenticated user. Same consequence as the first: control of the service's backing store.","attack_vector":"Any authenticated user of the service.","remediation":"Upgrade to v3.0 or later. Userspace, restart only.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39766","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated"},{"id":"CVE-2024-41022","cve":"CVE-2024-41022","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Fix signedness bug in sdma_v4_0_process_trap_irq()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41022","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42123","cve":"CVE-2024-42123","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A double free in the amdgpu firmware, ACPI and IP-block initialisation. The same allocation is released…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"A double free in the amdgpu firmware, ACPI and IP-block initialisation. The same allocation is released twice, corrupting the slab allocator's freelist. This is a classic heap-corruption primitive: with slab grooming it becomes arbitrary kernel memory write and therefore host compromise from an unprivileged GPU workload. The cheap outcome is a node panic. Upstream fix: drm/amdgpu: fix double free err_addr pointer warnings","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42123","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42228","cve":"CVE-2024-42228","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu kernel driver core. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: Using uninitialized value *size when calling amdgpu_vce_cs_reloc","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42228","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47975","cve":"CVE-2024-47975","aliases":["Solidigm SA-000563"],"title":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620 Opal) - improper access control validation: Improper access-control validation…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620…","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Improper access-control validation in the Opal-enabled firmware lets an attacker with physical access gain unauthorized access to the drive, or a local attacker knock it offline. The highest-scored entry in Solidigm's 2024 advisory set, and it hits exactly the SKUs an operator chose specifically BECAUSE they are Opal drives - the ones bought so that locking ranges would separate tenants and make decommissioning safe. BREAKS TENANT HANDOFF: the locking that your reclaim process depends on can be walked around, so a drive you believed was cryptographically locked between customers is readable. Also carries an availability tail - a local attacker can take the drive down, and on a shared bare-metal node that is a tenant-triggered outage.","attack_vector":"An attacker with physical access to the drive - the RMA return path, a decommissioned node in the resale channel, or a colo/rack tech - for the unauthorized-access half; a tenant with local access on the host for the denial-of-service half.","remediation":"Firmware flash per SKU with drive offline and node drained, using Solidigm Storage Tool: VEV10294/VDV10194/VEV10394 for the P4510/P4511/P4610 Opal variants, 3DV10132 for D5-P4320 Opal, 8DV10564 for D5-P4326 Opal, ACV10310 for D5-P5316 Opal, JCV10300 for D7-P5510 Opal, 9CV10410 for D7-P5520/P5620 Opal. Stop treating an Opal locking range as your tenant boundary on its own - it is a vendor-attested control you cannot audit. Layer LUKS/dm-crypt above it so a locking-range bypass yields ciphertext. Chain-of-custody matters as much as the flash here: because the attack is physical, tighten RMA and decommission handling (destroy rather than return where contract allows, or degauss/shred media that held tenant data) - a firmware update on drives still in the rack does nothing for the ones already on a pallet heading back to the vendor.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47975","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-6409","cve":"CVE-2024-6409","aliases":[],"title":"OpenSSH (sshd, RHEL9): Signal-handling race in the privsep child - possible RCE, RHEL 9 specific","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (sshd, RHEL9)","year":"2024","cvss_score":7,"severity":"high","kev":false,"impact":"Signal-handling race in the privsep child - possible RCE, RHEL 9 specific","attack_vector":"Unauthenticated network","remediation":"Package update + sshd restart","references":["https://access.redhat.com/security/cve/CVE-2024-6409"],"status":"curated"},{"id":"CVE-2025-23279","cve":"CVE-2025-23279","aliases":[],"title":"GPU Display Driver: Local privesc (race condition)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privesc (race condition)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23279","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-367"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23280","cve":"CVE-2025-23280","aliases":[],"title":"GPU Display Driver: Local privesc (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privesc (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23280","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23281","cve":"CVE-2025-23281","aliases":[],"title":"GPU Display Driver: Local privesc (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privesc (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23281","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23282","cve":"CVE-2025-23282","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): MULTI-TENANT ISOLATION: A race condition in the Linux display driver is winnable by a local attacker and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"MULTI-TENANT ISOLATION: A race condition in the Linux display driver is winnable by a local attacker and escalates to code execution in kernel context - the strongest container-escape primitive in this batch of driver bugs. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23282","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-415"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-32023","cve":"CVE-2025-32023","aliases":[],"title":"Redis: Authenticated user triggers a stack/heap out-of-bounds write in hyperloglog ops","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Authenticated user triggers a stack/heap out-of-bounds write in hyperloglog ops -> potential RCE","attack_vector":"Local","remediation":"Control-plane: upgrade to 8.0.3/7.4.5/7.2.10/6.2.19; restrict the command surface","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32023"],"status":"curated"},{"id":"CVE-2025-32462","cve":"CVE-2025-32462","aliases":[],"title":"sudo: Local privilege escalation via the `--host` option against host-specific sudoers rules","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"sudo","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Local privilege escalation via the `--host` option against host-specific sudoers rules","attack_vector":"Local user with any sudoers entry","remediation":"Package update; no reboot","references":["https://access.redhat.com/security/cve/CVE-2025-32462"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-3770","cve":"CVE-2025-3770","aliases":["GHSA-vx5v-4gg6-6qxr"],"title":"EDK II (SMM environment, Machine Check Exception handling): Machine Check Exceptions are enabled before SMM installs a handler for them, so an MCE fired in that window…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II (SMM environment, Machine Check Exception handling)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Machine Check Exceptions are enabled before SMM installs a handler for them, so an MCE fired in that window is delivered to whatever the IDT happens to point at - which an attacker can arrange. That is arbitrary code execution at ring -2. SMM sits below the hypervisor and below the OS: an implant there survives OS reinstall and node reimaging, can forge or suppress the measurements that feed remote attestation, and is invisible to every agent a tenant or operator runs. On multi-tenant GPU hardware it is the difference between wiping a node between customers and believing you wiped it.","attack_vector":"Local attacker with sufficient privilege on the host OS to trigger a machine check at the right moment - realistically root/admin on the node, or a tenant with kernel-level access on bare metal.","remediation":"Firmware flash from the server OEM. This is a core edk2 fix published August 2025, so the IBV rebase into AMI/Insyde/Phoenix trees and then into OEM BIOS payloads is the long pole - budget one to two OEM BIOS release cycles and track it per platform generation. One reboot per node, drain first. No configuration workaround: SMM is always present and cannot be disabled. Compensating control is to not hand kernel-level access on shared bare metal to untrusted tenants without a full firmware re-flash between leases.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-3770","https://github.com/tianocore/edk2/security/advisories/GHSA-vx5v-4gg6-6qxr"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-4802","cve":"CVE-2025-4802","aliases":[],"title":"glibc: Static setuid binaries incorrectly search LD_LIBRARY_PATH during dlopen - local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"glibc","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"Static setuid binaries incorrectly search LD_LIBRARY_PATH during dlopen - local privilege escalation","attack_vector":"Local user","remediation":"Package update; restart or reboot for full coverage of long-running processes","references":["https://access.redhat.com/security/cve/CVE-2025-4802"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-54518","cve":"CVE-2025-54518","aliases":["XSA-490"],"title":"Xen (x86): CPU opcode cache corruption - host instability triggerable from a guest","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86)","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"CPU opcode cache corruption - host instability triggerable from a guest","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + host reboot","references":["https://xenbits.xen.org/xsa/advisory-490.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-6019","cve":"CVE-2025-6019","aliases":[],"title":"libblockdev / udisks: allow_active to root via libblockdev through udisks - second half of the chain, works on nearly every…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"libblockdev / udisks","year":"2025","cvss_score":7,"severity":"high","kev":false,"impact":"allow_active to root via libblockdev through udisks - second half of the chain, works on nearly every mainstream distro","attack_vector":"Local user","remediation":"Package update + restart udisksd; no reboot","references":["https://access.redhat.com/security/cve/CVE-2025-6019"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2026-20885","cve":"CVE-2026-20885","aliases":["TDX module improper authentication"],"title":"Intel TDX module, Ring 0 / Trust Domain context, multiple Intel platforms - INTEL-SA-01436: Improper authentication in the TDX module allows both information disclosure and privilege escalation…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module, Ring 0 / Trust Domain context, multiple Intel platforms - INTEL-SA-01436","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"Improper authentication in the TDX module allows both information disclosure and privilege escalation from a system-software adversary. Authentication failures in the TDX module are the worst shape of TDX bug because the module's job is to decide who is allowed to ask it for what - a flaw there means the untrusted hypervisor can present itself as authorized for operations reserved to the trust domain. Result is tenant data disclosure plus escalation into the protected domain. This is one of the most recent items in the set, from Intel's August 2026 advisory batch, so many fleets have not yet deployed the fix.","attack_vector":"A system-software adversary with privileged host access - the hypervisor or host root - against the TDX module. Intel rates attack complexity as high, but the required position is one every infrastructure operator already occupies.","remediation":"Update the TDX module to the fixed version per INTEL-SA-01436, plus the associated platform firmware. Host reboot and full drain of trust domains. Because this landed in the August 2026 batch alongside several other TDX and transient-execution CVEs, deploy it as one consolidated firmware and TDX-module campaign rather than a series of single-CVE windows - each window costs a full fleet drain. Then raise the accepted TDX module SVN in your attestation policy; until you do, patched and unpatched hosts are indistinguishable to relying parties.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-20885","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01436.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24200","cve":"CVE-2026-24200","aliases":[],"title":"vGPU Manager: Guest-to-host escape (use-after-free in GPU context mgmt)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"Guest-to-host escape (use-after-free in GPU context mgmt)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate all guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24200","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53329","cve":"CVE-2026-53329","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Use krealloc_array() in dal_vector_reserve()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53329","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64219","cve":"CVE-2026-64219","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":7,"severity":"high","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Validate payload length and link_index in dc_process_dmub_aux_transfer_async","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64219","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-25647","cve":"CVE-2020-25647","aliases":[],"title":"GRUB2 (USB device initialization): Out-of-bounds write in grub_usb_device_initialize from a malicious USB descriptor. In a datacenter this is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (USB device initialization)","year":"2020","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Out-of-bounds write in grub_usb_device_initialize from a malicious USB descriptor. In a datacenter this is not a 'someone walks up with a USB stick' story - the BMC presents virtual media as a USB device, so anyone with BMC credentials can trigger it entirely remotely.","attack_vector":"Physical USB, or - the one that matters - BMC virtual media, which turns this into a remote attack for anyone on the management VLAN with iDRAC/iLO/XCC credentials.","remediation":"grub2 package update + reboot. Meaningful compensating control: disable virtual media on the BMC where you do not use it for provisioning, and keep the management network off any tenant-reachable path.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-25647","https://access.redhat.com/security/cve/CVE-2020-25647"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-41333","cve":"CVE-2023-41333","aliases":[],"title":"Cilium: A user who can create CiliumNetworkPolicy in one namespace affects traffic cluster-wide","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":6.9,"severity":"medium","kev":false,"impact":"A user who can create CiliumNetworkPolicy in one namespace affects traffic cluster-wide; cross-tenant policy tampering","attack_vector":"Cluster user with namespace access","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-41333"],"status":"curated"},{"id":"CVE-2024-24557","cve":"CVE-2024-24557","aliases":[],"title":"Docker / moby: Classic builder cache poisoning for images built FROM scratch","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Classic builder cache poisoning for images built FROM scratch","attack_vector":"Anyone sharing a build cache, e.g. a multi-tenant CI builder","remediation":"Upgrade moby; give each tenant an isolated build cache","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24557"],"status":"curated"},{"id":"CVE-2025-0045","cve":"CVE-2025-0045","aliases":[],"title":"AMD Secure Processor PCI driver - input validation: Improper input validation in the ASP PCI driver lets a local attacker trigger a buffer overflow and crash the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor PCI driver - input validation","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Improper input validation in the ASP PCI driver lets a local attacker trigger a buffer overflow and crash the node. Practically this is availability: a tenant-adjacent process that can reach the driver can take the host down, and on a GPU node that means every co-resident training job dies with it.","attack_vector":"Local, through the ASP PCI driver interface. Not tenant-container reachable on a normally configured node - the ccp/PSP device is not handed to workload containers.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. There is also a kernel-side component (the ccp driver); take the distro kernel update as well as the BIOS, since the OEM firmware will lag.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0045","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-29939","cve":"CVE-2025-29939","aliases":[],"title":"AMD SEV firmware - RMP write during SNP initialization: MULTI-TENANT ISOLATION: A privileged attacker can write to the reverse map page during secure nested paging…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - RMP write during SNP initialization","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A privileged attacker can write to the reverse map page during secure nested paging initialization, corrupting the ownership map before any guest launches. Because the RMP is initialised once and then trusted, poisoning it at init time means every guest that subsequently launches on that host inherits a compromised isolation boundary.","attack_vector":"Local, privileged, during SNP init - so an attacker who controls the host's boot or SNP initialisation path.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands. Pair with host measured boot: if you cannot attest the boot sequence, you cannot rule out that SNP was initialised under an attacker's influence.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29939","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-48516","cve":"CVE-2025-48516","aliases":[],"title":"AMD AGESA bootloader - DDR5 PMIC default configuration: The AGESA bootloader leaves DDR5 memory modules in an insecure default state with the on-DIMM power…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD AGESA bootloader - DDR5 PMIC default configuration","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"The AGESA bootloader leaves DDR5 memory modules in an insecure default state with the on-DIMM power management IC interface unprotected. A local user can then reprogram the PMIC and destroy the module - a **permanent**, physical denial of service. This is one of the rare software bugs whose remediation is an RMA: an attacker who runs this across a fleet does not take your nodes offline for a reboot, they take them offline for a parts order.","attack_vector":"Local user privilege on the host. No physical access needed - the PMIC is reachable over the platform's memory-module management interface.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. There is no software undo once a DIMM is bricked. Until the OEM BIOS lands, the mitigation is access control: this needs local execution on the host, so it is a strong argument for not giving semi-trusted workloads a shell on bare metal. Budget for spare DIMMs on any fleet where you cannot patch quickly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48516","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-48518","cve":"CVE-2025-48518","aliases":[],"title":"AMD Graphics Driver - out-of-bounds write: MULTI-TENANT ISOLATION: Improper input validation lets a local attacker write out of bounds through the AMD…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Graphics Driver - out-of-bounds write","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper input validation lets a local attacker write out of bounds through the AMD graphics driver, costing integrity or availability. Out-of-bounds kernel writes reachable from a GPU device handle are the standard escape route out of a GPU container.","attack_vector":"Local, via the graphics driver interface - reachable from a tenant container with a GPU device node.","remediation":"Update the AMD graphics driver package and reload the driver or reboot the node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48518","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-48521","cve":"CVE-2025-48521","aliases":[],"title":"AMD Secure Processor PCI driver - use-after-free: A use-after-free reachable through the ASP PCI driver. Beyond the crash, freed-then-reused kernel memory is a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor PCI driver - use-after-free","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"A use-after-free reachable through the ASP PCI driver. Beyond the crash, freed-then-reused kernel memory is a corruption primitive, so the honest read is loss of platform integrity rather than simple denial of service.","attack_vector":"Local, via the ASP PCI driver interface; requires access to the crypto/PSP device, i.e. host-level not tenant-level.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Pair the BIOS update with the corresponding kernel ccp driver fix.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48521","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-58354","cve":"CVE-2025-58354","aliases":[],"title":"Kata Containers: A malicious host can circumvent guest protections","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"A malicious host can circumvent guest protections","attack_vector":"A compromised or hostile host operator","remediation":"Upgrade Kata; relevant when Kata is used to isolate tenants from each other, not from you","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-58354"],"status":"curated"},{"id":"CVE-2025-64329","cve":"CVE-2025-64329","aliases":[],"title":"containerd: CRI Attach implementation bug lets a user attach to a container they should not reach","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"CRI Attach implementation bug lets a user attach to a container they should not reach","attack_vector":"Cluster user with namespace access","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64329"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-64436","cve":"CVE-2025-64436","aliases":[],"title":"KubeVirt: virt-handler service-account permissions (update VMI, patch nodes) can be abused to force VMI migration","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":6.9,"severity":"medium","kev":false,"impact":"virt-handler service-account permissions (update VMI, patch nodes) can be abused to force VMI migration","attack_vector":"An attacker with virt-handler credentials","remediation":"Upgrade KubeVirt; scope down the service account","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64436"],"status":"curated"},{"id":"CVE-2026-15789","cve":"CVE-2026-15789","aliases":[],"title":"BuildKit: Crafted upload request lets files escape the BuildKit state directory onto the host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Crafted upload request lets files escape the BuildKit state directory onto the host","attack_vector":"Anyone with build control API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15789"],"status":"curated"},{"id":"CVE-2026-31838","cve":"CVE-2026-31838","aliases":[],"title":"Istio: Envoy RBAC header matching flaw bypasses header-based authorization policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":6.9,"severity":"medium","kev":false,"impact":"Envoy RBAC header matching flaw bypasses header-based authorization policy","attack_vector":"Unauthenticated network","remediation":"Rolling istiod and proxy upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31838"],"status":"curated"},{"id":"CVE-2026-8810","cve":"CVE-2026-8810","aliases":["INSYDE-SA-2026005"],"title":"Insyde InsydeH2O on ARM platforms (HDD password storage in UEFI variables): HDD passwords are recoverable from UEFI variables on affected ARM platforms - insufficiently protected…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O on ARM platforms (HDD password storage in UEFI variables)","year":"2026","cvss_score":6.9,"severity":"medium","kev":false,"impact":"HDD passwords are recoverable from UEFI variables on affected ARM platforms - insufficiently protected credentials, in CWE terms. The operator concern is drive-level secrets surviving in firmware storage where the next person to hold the box can read them, which matters for any fleet where drive locking is part of the between-tenant wipe story, and for ARM-based nodes appearing in AI inference and edge fleets.","attack_vector":"Requires physical access plus local privilege and user interaction per Insyde's own scoring - so this is a returned-hardware, decommissioning, or colo-access risk rather than a remote one.","remediation":"OEM firmware update on Insyde kernel 5.6 / 05.63.21 or 5.7 / 05.72.21. Firmware flash, reboot per node. Advisory dated 2026-08-18, so OEM images are not yet widely available. Operational mitigation that does not wait for the flash: do not rely on ATA HDD passwords as the confidentiality control on affected ARM platforms - use self-encrypting-drive keys or software full-disk encryption with keys held off the node, and cryptographically erase rather than password-lock drives at decommission.","references":["https://www.insyde.com/security-pledge/SA-2026005/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2017-8371","cve":"CVE-2017-8371","aliases":[],"title":"Schneider Electric StruxureWare Data Center Expert before 7.4.0: Passwords held in cleartext in RAM on the DCIM appliance, recoverable remotely. Included here because it is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric StruxureWare Data Center Expert before 7.4.0","year":"2017","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Passwords held in cleartext in RAM on the DCIM appliance, recoverable remotely. Included here because it is the earliest entry in a seven-year pattern: DCE has repeatedly failed to protect the device credentials it must hold, and any operator running an old DCE build should assume the facility credential set is compromised rather than assume otherwise.","attack_vector":"Remote, per the advisory; unspecified vectors, but the practical read is that a foothold on or near the appliance yields the credentials.","remediation":"Upgrade to 7.4.0 or later - though anyone still on a pre-7.4 build has far larger problems from the 2021-2024 RCEs above. Rotate all device credentials.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-8371"],"status":"curated"},{"id":"CVE-2018-15776","cve":"CVE-2018-15776","aliases":[],"title":"Dell iDRAC (u-boot): Improper error handling grants access to the u-boot shell — pre-BMC-OS control, i.e. below even the BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC (u-boot)","year":"2018","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper error handling grants access to the u-boot shell — pre-BMC-OS control, i.e. below even the BMC firmware image","attack_vector":"Local / serial-adjacent","remediation":"Firmware update; a u-boot foothold survives BMC firmware reflash, so affected nodes need physical verification rather than a remote fix","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-15776"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-18095","cve":"CVE-2018-18095","aliases":["INTEL-SA-00267","LEN-28116"],"title":"Intel SSD DC S4500 and SSD DC S4600 series firmware before SCV10150 - improper authentication: Improper authentication in the drive firmware lets an UNPRIVILEGED user with physical access escalate…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SSD DC S4500 and SSD DC S4600 series firmware before SCV10150 - improper authentication","year":"2018","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper authentication in the drive firmware lets an UNPRIVILEGED user with physical access escalate privilege on the drive, with full confidentiality, integrity and availability impact. These are mainstream datacenter SATA SSDs that shipped in enormous volume as boot and scratch media in ProLiant/PowerEdge/ThinkSystem-class nodes, which is exactly the tier of hardware that gets recycled into GPU builds and rented out bare-metal. BREAKS TENANT HANDOFF: authentication on the drive is the mechanism that is supposed to keep a departing tenant from reaching a reprovisioned drive's contents, and it is bypassable by someone with no privilege at all. The integrity impact matters as much as the read - an attacker who owns the drive controller can persist there through your reimage.","attack_vector":"An unprivileged attacker with physical access to the drive. No credentials of any kind required. Realistically: anyone in the hardware's physical path - decommission, RMA, colo, or a rack tech - and any second-hand S4500/S4600 you bought to fill out a build.","remediation":"Flash to firmware SCV10150 or later; drive offline, node drained, vendor tooling (Intel MAS, or the OEM's own bundle - Lenovo shipped this as LEN-28116 and F5 issued its own bulletin, so if these drives came inside an OEM chassis check the OEM's firmware bundle rather than Intel's, because the OEM-qualified image is often the only one their controller will accept). This is a 2018 fix, so the operational question in 2026 is not whether a patch exists but whether anyone ever applied it to the second-hand S4500/S4600 inventory now sitting in your fleet - assume not, and audit. Any drive you cannot confirm was flashed and cannot account for the custody of should be treated as potentially firmware-implanted and destroyed rather than redeployed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-18095","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00267.html","https://support.lenovo.com/us/en/product_security/LEN-28116"],"status":"curated"},{"id":"CVE-2019-18424","cve":"CVE-2019-18424","aliases":["XSA-302","Xen passthrough DMA host privilege escalation"],"title":"Xen through 4.12.x - passed-through PCI devices left able to DMA into host memory after being handed to an untrusted domain: A guest that has been assigned a physical device gains host privileges…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen through 4.12.x - passed-through PCI devices left able to DMA into host memory after being handed to an untrusted…","year":"2019","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A guest that has been assigned a physical device gains host privileges through DMA, because the device retains reach into memory it should have lost. For a GPU-passthrough cloud this is the direct form of the failure everyone worries about: the tenant you gave a GPU to uses that GPU to read and write the hypervisor. It also breaks tenant handoff, since the state that makes the device dangerous is set up around assignment and de-assignment - the exact transition that happens between customers.","attack_vector":"A tenant in a guest domain with a physical device assigned to it. That is the normal configuration of a GPU-passthrough product, not an unusual one.","remediation":"Patch Xen (XSA-302) and reboot the hypervisor - a rolling drain across the fleet. Structurally, use the assignable-add workflow so devices are explicitly quarantined before and after assignment rather than being handed straight from host to guest. Beyond this specific CVE, the general lesson holds for every passthrough platform including KVM/VFIO: reset the device, flush its DMA mappings and re-verify its firmware between tenants, and treat 'the device was assigned to someone else five minutes ago' as untrusted state.","references":["https://xenbits.xen.org/xsa/advisory-302.html","https://nvd.nist.gov/vuln/detail/CVE-2019-18424"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-12355","cve":"CVE-2020-12355","aliases":["INTEL-SA-00391"],"title":"RPMB protocol message authentication subsystem in Intel TXE before 4.0.30 (replay-protected memory block): Capture-replay authentication bypass in RPMB - the mechanism that is supposed to make…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RPMB protocol message authentication subsystem in Intel TXE before 4.0.30 (replay-protected memory block)","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Capture-replay authentication bypass in RPMB - the mechanism that is supposed to make firmware anti-rollback and monotonic-counter state tamper-evident. Break RPMB and you break rollback protection: an attacker can replay old, signed state to roll firmware back to a known-vulnerable version, or to reset counters that firmware relies on to detect tampering. The operator consequence is that 'we patched that' stops being verifiable from the platform itself, and a node that you believe is on current firmware can be silently downgraded and left that way through a tenant handoff.","attack_vector":"Physical access to the platform, capturing and replaying RPMB traffic. Relevant for hardware that passes through untrusted hands - shared cages, remote-hands, RMA and resale channels, and any secondhand GPU capacity you have taken on.","remediation":"TXE/CSME firmware update from the OEM bundle, host reboot and drain. More importantly, stop treating the platform's own report of its firmware version as authoritative: read firmware versions out of band via the BMC and compare against an externally held inventory, and re-flash the full firmware stack on any node that has been out of your physical custody before it re-enters a tenant pool.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12355","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00391.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-13799","cve":"CVE-2020-13799","aliases":["VU#231329","WDC-20008","RPMB replay"],"title":"Replay Protected Memory Block (RPMB) protocol as specified for eMMC, UFS and ALL versions of NVMe - multi-vendor storage device flaw: RPMB is the small authenticated region storage devices provide…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Replay Protected Memory Block (RPMB) protocol as specified for eMMC, UFS and ALL versions of NVMe - multi-vendor…","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"RPMB is the small authenticated region storage devices provide so a host can keep trusted firmware and anti-rollback state where software cannot forge it. The protocol as SPECIFIED - not one vendor's bug, but the standard itself, across eMMC, UFS and every version of NVMe - permits replay attacks that let an attacker roll the protected region back to an earlier state. That undermines the anti-rollback guarantee that firmware-integrity schemes are built on, so an attacker can reinstate previously-revoked firmware or state. This is a foundational-trust issue rather than a data-read issue: the mechanism your platform uses to prove firmware has not been downgraded can itself be replayed. Because it is written into the standard, it is present across vendors and generations simultaneously - the widest blast radius of anything in this category.","attack_vector":"An attacker with physical access to the device, or with the ability to interpose on the host-to-device command path, capturing and replaying RPMB message sequences. Multi-vendor by construction, since the flaw is in the specification that every implementer followed.","remediation":"No single patch exists - remediation is per-implementation and depends on the host software and the device firmware cooperating, which is why CERT/CC coordinated it across vendors rather than issuing one fix. Check each storage vendor's advisory for your specific SKUs (Western Digital published WDC-20008; CERT/CC VU#231329 tracks the multi-vendor response) and apply device firmware plus any host-side platform firmware updates they name. Practically, most operators will not be able to close this on existing fleet hardware, so treat RPMB-backed anti-rollback as a control you cannot fully rely on: do not let it be the only thing preventing a firmware downgrade. Keep an independent record of expected firmware versions per drive serial, and alert on any drive whose reported firmware version goes BACKWARDS between inventory scans - a downgrade you detect is far more useful than an anti-rollback guarantee you cannot verify.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13799","https://www.kb.cert.org/vuls/id/231329"],"status":"curated"},{"id":"CVE-2020-16844","cve":"CVE-2020-16844","aliases":[],"title":"Istio: DENY AuthorizationPolicy with wildcard-suffix principals silently fails to deny","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DENY AuthorizationPolicy with wildcard-suffix principals silently fails to deny","attack_vector":"Any pod on the mesh","remediation":"Rolling istiod upgrade; re-verify deny policies","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-16844"],"status":"curated"},{"id":"CVE-2020-8705","cve":"CVE-2020-8705","aliases":[],"title":"Intel Boot Guard in Intel CSME / TXE / SPS: Insecure default initialisation in Boot Guard means the S3 resume path does not re-verify the boot chain, so…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Boot Guard in Intel CSME / TXE / SPS","year":"2020","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Insecure default initialisation in Boot Guard means the S3 resume path does not re-verify the boot chain, so an attacker who can modify firmware while the machine is suspended defeats verified boot. Boot Guard is the hardware root of trust the rest of your platform attestation chains to - if it does not hold on resume, it does not hold.","attack_vector":"An attacker able to modify platform firmware, typically with physical access or an existing firmware-write primitive, against a machine that suspends.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines. Servers that never suspend are largely out of scope, which is most of a datacenter fleet - but verify rather than assume, because management controllers do use low-power states.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8705","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00391"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-21284","cve":"CVE-2021-21284","aliases":[],"title":"Docker / moby: With --userns-remap, remapped root can escalate to real host root","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"With --userns-remap, remapped root can escalate to real host root","attack_vector":"Any tenant workload on a userns-remapped daemon","remediation":"Upgrade Docker Engine; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21284"],"status":"curated"},{"id":"CVE-2021-28694","cve":"CVE-2021-28694","aliases":[],"title":"Xen on AMD-Vi (AMD IOMMU) - ACPI IVMD unity-map page permissions: MULTI-TENANT ISOLATION: Xen honours ACPI-described IOMMU unity mappings but applied page permissions…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi (AMD IOMMU) - ACPI IVMD unity-map page permissions","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Xen honours ACPI-described IOMMU unity mappings but applied page permissions inconsistently, so a passed-through device can end up with wider DMA access than the hypervisor intended. On a GPU cloud that does device passthrough, this is the boundary that stops one tenant's assigned accelerator from DMA-ing into another guest's memory - and it was not holding.","attack_vector":"Requires a guest with a passed-through PCI device, which on a GPU cloud is the standard configuration. The malicious guest drives DMA from its own assigned device.","remediation":"Fixed in Xen via XSA-378. Update the hypervisor and reboot the host; no firmware step. If you run GPU passthrough on Xen, this is a top-priority class - the whole security model of passthrough rests on the IOMMU restricting device DMA correctly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28694","https://xenbits.xen.org/xsa/advisory-378.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-28695","cve":"CVE-2021-28695","aliases":[],"title":"Xen on AMD-Vi - IOMMU page mapping permissions: MULTI-TENANT ISOLATION: Second of the XSA-378 IOMMU page-mapping issues on AMD-Vi. Incorrect mapping…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - IOMMU page mapping permissions","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Second of the XSA-378 IOMMU page-mapping issues on AMD-Vi. Incorrect mapping permissions give a passed-through device DMA reach beyond its guest, which is guest-to-host and guest-to-guest memory access via a device rather than via the CPU.","attack_vector":"Guest with a passed-through PCI device - the normal GPU-passthrough configuration.","remediation":"Fixed in Xen (XSA-378). Hypervisor update plus host reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28695","https://xenbits.xen.org/xsa/advisory-378.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-28696","cve":"CVE-2021-28696","aliases":[],"title":"Xen on AMD-Vi - IOMMU page mapping permissions: MULTI-TENANT ISOLATION: Third of the XSA-378 AMD-Vi mapping issues. Same practical consequence: a device…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD-Vi - IOMMU page mapping permissions","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Third of the XSA-378 AMD-Vi mapping issues. Same practical consequence: a device assigned to one guest can reach memory it should not, defeating passthrough isolation.","attack_vector":"Guest with an assigned PCI device.","remediation":"Fixed in Xen (XSA-378). Update and reboot; patch all three of the XSA-378 CVEs together.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28696","https://xenbits.xen.org/xsa/advisory-378.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-32690","cve":"CVE-2021-32690","aliases":[],"title":"Helm: Helm repository credentials leaked to a redirected third-party host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Helm repository credentials leaked to a redirected third-party host","attack_vector":"Malicious or compromised chart repository","remediation":"Upgrade Helm; rotate repo credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32690"],"status":"curated"},{"id":"CVE-2021-33077","cve":"CVE-2021-33077","aliases":["CVE-2021-33080","INTEL-SA-00563"],"title":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware - control-flow flaw and uncleared debug data reachable over physical access: Two firmware defects in the same advisory, both scored 6.8 and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware - control-flow flaw and uncleared debug data reachable over…","year":"2021","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Two firmware defects in the same advisory, both scored 6.8 and both giving full confidentiality-plus-integrity impact to someone holding the drive. The control-flow flaw escalates privilege inside the drive controller; the second leaves debug information in the firmware that was never cleared for production, exposing internal state and enabling further escalation. Together they mean the drive controller itself is takeable - and a controller you control is a controller that can lie about sanitize, lie about encryption, and read every block regardless of Opal locking. BREAKS TENANT HANDOFF at the root: every erase-and-report guarantee in your reclaim pipeline is only as trustworthy as the firmware making the report, and here that firmware is compromisable. Also a leftover-debug-in-shipping-firmware finding, which tells you what the vendor's production hardening was actually worth on these SKUs.","attack_vector":"An unauthenticated attacker with physical access to the drive - no password, no host credential, no prior privilege. In an operator's world that means the RMA return path, a decommissioned node, a colo cage with shared access, or a supply-chain touchpoint before the drive was ever racked.","remediation":"Firmware flash per SKU using Intel MAS / Solidigm Storage Tool, drive offline and node drained. Match your inventory against INTEL-SA-00563 - it covers a wide spread of Intel SSD, SSD DC, Optane SSD and Optane SSD DC families with different fixed-firmware versions each, and some end-of-support SKUs get no fix at all. Beyond patching, this is the entry that should drive a chain-of-custody policy rather than a technical one: because exploitation needs only physical possession, the highest-leverage change is to stop letting drives that held tenant data leave your control intact. Destroy media on decommission instead of reselling, negotiate destroy-in-place terms with the vendor for RMA, and treat any drive with an unexplained gap in custody as compromised at the firmware level rather than re-racking it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33077","https://nvd.nist.gov/vuln/detail/CVE-2021-33080","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00563.html","https://www.solidigm.com/support-page/support-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-0004","cve":"CVE-2022-0004","aliases":[],"title":"Intel Boot Guard and Intel TXT (hardware debug / INIT): Hardware debug modes and processor INIT handling can override the locks that Boot Guard and TXT rely on…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Boot Guard and Intel TXT (hardware debug / INIT)","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Hardware debug modes and processor INIT handling can override the locks that Boot Guard and TXT rely on, letting an unauthenticated attacker escalate privilege and defeat measured/verified boot. This reaches both of Intel's boot-integrity technologies at once - the two things a remote attestation claim about a bare-metal node ultimately rests on.","attack_vector":"An attacker with the ability to trigger the relevant debug or INIT conditions - realistically physical or deep platform access.","remediation":"Platform BIOS/firmware update from the OEM. Drain and reboot per node, and wait on OEM packaging. Until then, treat Boot Guard and TXT measurements from affected platforms as advisory rather than authoritative in any attestation policy you enforce.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0004","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00613.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-21657","cve":"CVE-2022-21657","aliases":[],"title":"Envoy: Envoy accepts any peer certificate rather than restricting to configured CAs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Envoy accepts any peer certificate rather than restricting to configured CAs; mTLS trust bypass","attack_vector":"Unauthenticated network with any valid-looking cert","remediation":"Upgrade Envoy; sidecar restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21657"],"status":"curated"},{"id":"CVE-2022-28185","cve":"CVE-2022-28185","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An out-of-bounds write in the driver's ECC layer, reachable by an unprivileged user, corrupts state and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the driver's ECC layer, reachable by an unprivileged user, corrupts state and crashes the node. Relevant because ECC handling is the layer operators lean on for memory integrity on Tesla parts. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5353. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28185","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-31611","cve":"CVE-2022-31611","aliases":[],"title":"GeForce Experience installer: Local privesc during install","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce Experience installer","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Local privesc during install","attack_vector":"Local user","remediation":"Consumer-only; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31611","https://github.com/NVIDIA/product-security/tree/main/2023/5384"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:L/I:H/A:H","cwe":["CWE-427"]},{"id":"CVE-2022-34674","cve":"CVE-2022-34674","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): MULTI-TENANT ISOLATION: A kernel helper maps more physical pages than were actually requested, so a caller…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":6.8,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A kernel helper maps more physical pages than were actually requested, so a caller can read physical memory that was never meant to be exposed to it - a direct residual-memory read from an unprivileged GPU workload. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34674","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:L/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-0778","cve":"CVE-2023-0778","aliases":[],"title":"Podman: TOCTOU during volume export lets a symlink swap expose arbitrary host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"TOCTOU during volume export lets a symlink swap expose arbitrary host files","attack_vector":"Any tenant workload with volume access","remediation":"Upgrade Podman","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0778"],"status":"curated"},{"id":"CVE-2023-20589","cve":"CVE-2023-20589","aliases":["faulTPM"],"title":"AMD Secure Processor secure boot - voltage fault injection (AMD-SB-4005): MULTI-TENANT ISOLATION: Voltage fault injection against the ASP defeats its secure boot, yielding arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor secure boot - voltage fault injection (AMD-SB-4005)","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Voltage fault injection against the ASP defeats its secure boot, yielding arbitrary code execution on the secure processor and extraction of fTPM-sealed secrets - the published work pulls BitLocker keys out. The affected list is Zen 1/2/3 **client** parts and EPYC is not on it, but the technique is the reason to read this entry: a datacenter operator with hardware in colocation, in transit, or coming back from RMA has to assume physical access to some nodes by someone.","attack_vector":"Physical access plus specialised fault-injection hardware. Not remote, not local-software.","remediation":"Fixed in AGESA/PI firmware for the affected client parts - OEM BIOS package, drain and power cycle. For a server fleet the practical answer is not patching but physical control: tamper-evident handling, chain of custody for RMAs and redeployments, and not trusting fTPM-sealed secrets on any node that has left your custody. AMD's position on the server analogue (AMD-SB-3028, voltage fault injection against SEV VMs on EPYC 7272) is **WONTFIX** - physical attacks are declared outside the SEV-SNP threat model, so there is no patch coming for the server case at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20589","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-4005.html","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3028.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-28005","cve":"CVE-2023-28005","aliases":[],"title":"Trend Micro Endpoint Encryption Full Disk Encryption (UEFI pre-boot): A signed pre-boot component that allows Secure Boot bypass on affected machines. Relevant to any fleet where…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Trend Micro Endpoint Encryption Full Disk Encryption (UEFI pre-boot)","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A signed pre-boot component that allows Secure Boot bypass on affected machines. Relevant to any fleet where an endpoint-security vendor's UEFI component is part of the boot chain - the security product itself becomes the way in.","attack_vector":"Local access to a machine running the affected pre-boot component.","remediation":"Vendor product update plus, where the binary is revoked, a dbx update. Worth a general lesson for fleet operators: every third-party signed EFI component you allow into the boot chain is a permanent addition to your Secure Boot attack surface, and you cannot remove it later without a revocation rollout.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28005"],"status":"curated"},{"id":"CVE-2023-28841","cve":"CVE-2023-28841","aliases":[],"title":"Docker / moby: Encrypted overlay network traffic can be unencrypted due to missing rules","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Encrypted overlay network traffic can be unencrypted due to missing rules","attack_vector":"Unauthenticated network adjacent to the underlay","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28841"],"status":"curated"},{"id":"CVE-2023-28842","cve":"CVE-2023-28842","aliases":[],"title":"Docker / moby: Unauthenticated injection of traffic into an encrypted overlay network","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Unauthenticated injection of traffic into an encrypted overlay network","attack_vector":"Unauthenticated network adjacent to the underlay","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28842"],"status":"curated"},{"id":"CVE-2023-31010","cve":"CVE-2023-31010","aliases":[],"title":"DGX H100 BMC (IPMI): DoS of the BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS of the BMC","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-20"]},{"id":"CVE-2023-31033","cve":"CVE-2023-31033","aliases":[],"title":"DGX A100 BMC: Missing authentication on BMC service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Missing authentication on BMC service","attack_vector":"Network-adjacent unauthenticated","remediation":"Flash BMC 00.22.05+ out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31033","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"]},{"id":"CVE-2024-0111","cve":"CVE-2024-0111","aliases":[],"title":"CUDA Toolkit: Memory-safety issue","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Memory-safety issue -> code exec","attack_vector":"Malicious model/binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0111","https://github.com/NVIDIA/product-security/tree/main/2024/5564"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:L/A:L","cwe":["CWE-1284"]},{"id":"CVE-2024-0140","cve":"CVE-2024-0140","aliases":[],"title":"RAPIDS cuDF / cuML: RCE via unsafe deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"RAPIDS cuDF / cuML","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"RCE via unsafe deserialization","attack_vector":"Malicious model/dataset artifact loaded by a tenant job","remediation":"Bump RAPIDS packages; rebuild data-science base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0140","https://github.com/NVIDIA/product-security/tree/main/2025/5597"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:L/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2024-0141","cve":"CVE-2024-0141","aliases":[],"title":"Hopper HGX 8-GPU: DoS (inadequate loop termination in firmware)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Hopper HGX 8-GPU","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS (inadequate loop termination in firmware)","attack_vector":"Local privileged host access","remediation":"Flash HGX baseboard firmware; node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0141","https://github.com/NVIDIA/product-security/tree/main/2025/5561"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-782"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-0142","cve":"CVE-2024-0142","aliases":[],"title":"nvJPEG2000: Code exec via OOB write on malicious image","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via OOB write on malicious image","attack_vector":"Malicious dataset / user-uploaded image","remediation":"Bump nvJPEG2000 in preprocessing images; rebuild","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0142","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-787"]},{"id":"CVE-2024-0143","cve":"CVE-2024-0143","aliases":[],"title":"nvJPEG2000: Code exec via OOB write","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via OOB write","attack_vector":"Malicious dataset","remediation":"Bump nvJPEG2000; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0143","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-787"]},{"id":"CVE-2024-0144","cve":"CVE-2024-0144","aliases":[],"title":"nvJPEG2000: Code exec via buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via buffer overflow","attack_vector":"Malicious dataset","remediation":"Bump nvJPEG2000; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0144","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-120"]},{"id":"CVE-2024-0145","cve":"CVE-2024-0145","aliases":[],"title":"nvJPEG2000: Code exec via heap overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvJPEG2000","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Code exec via heap overflow","attack_vector":"Malicious dataset","remediation":"Bump nvJPEG2000; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0145","https://github.com/NVIDIA/product-security/tree/main/2025/5596"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:R/S:U/C:N/I:H/A:H","cwe":["CWE-122"]},{"id":"CVE-2024-2315","cve":"CVE-2024-2315","aliases":["AMI-SA-2024004"],"title":"AMI AptioV UEFI BIOS (SPI flash access control): Improper access control in the BIOS that lets a local attacker make unexpected SPI flash modifications and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SPI flash access control)","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper access control in the BIOS that lets a local attacker make unexpected SPI flash modifications and launch a BIOS bootkit. This is the direct route to firmware persistence: write the flash, own every subsequent boot, and become invisible to the OS and to every agent running in it. AMI also calls out an availability impact - a bad write bricks the board, which on a GPU node means an RMA and weeks of lost capacity rather than a reboot.","attack_vector":"Local access with low privileges, no user interaction. Code on the host OS is enough; it does not require root by AMI's scoring. That makes it one of the cheaper firmware-persistence paths in this cluster for an attacker who has landed anywhere on the node.","remediation":"BIOS update to BKC_5.37 or later - firmware flash plus a host reboot, per node, gated on your server vendor's rebase. Alongside the update, verify that the platform's flash write protections are actually enabled in your BIOS configuration (flash descriptor lock, BIOS write-protect, boot guard where the platform supports it) - operators routinely find these left open by an OEM's default profile, and that config check costs one reboot rather than a flash campaign.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024004.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-2315"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-37085","cve":"CVE-2024-37085","aliases":[],"title":"VMware ESXi: AD-integrated ESXi grants full host admin to any member of a re-created \"ESX Admins\" group - used by Akira…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware ESXi","year":"2024","cvss_score":6.8,"severity":"medium","kev":true,"impact":"AD-integrated ESXi grants full host admin to any member of a re-created \"ESX Admins\" group - used by Akira and Black Basta to mass-encrypt VMs [KEV]","attack_vector":"Attacker with Active Directory write access (post-initial-access, not tenant-facing)","remediation":"Patch ESXi and stop using AD for ESXi user management. Configuration change, not just a binary update - the real fix is removing the AD trust from the hypervisor plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37085"],"status":"curated"},{"id":"CVE-2024-42488","cve":"CVE-2024-42488","aliases":[],"title":"Cilium: Agent race condition drops pod labels, so the wrong (often more permissive) policy applies","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Agent race condition drops pod labels, so the wrong (often more permissive) policy applies","attack_vector":"Any tenant workload, timing-dependent","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42488"],"status":"curated"},{"id":"CVE-2024-45101","cve":"CVE-2024-45101","aliases":["LEN-154748"],"title":"Lenovo XClarity Administrator (LXCA) - single sign-on to XCC: Where LXCA acts as the single sign-on provider for XCC, an attacker who gets an authenticated LXCA user to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Administrator (LXCA) - single sign-on to XCC","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Where LXCA acts as the single sign-on provider for XCC, an attacker who gets an authenticated LXCA user to click a crafted URL can intercept that user's XCC session. The stolen session is a live, authenticated connection to a node's service processor, carrying whatever privileges the victim held - typically enough for power control, Virtual Media and console. The interesting property is that SSO, which operators deploy to reduce credential sprawl across BMCs, becomes the mechanism that spreads one phished click into out-of-band access. Affects LXCA before 4.1.","attack_vector":"Requires an authenticated LXCA user to click an attacker-supplied link, and requires that SSO between LXCA and XCC is enabled. The attacker needs no reachability to the management VLAN themselves - the operator's browser session is the bridge.","remediation":"Upgrade LXCA to 4.1 or later - a single appliance upgrade, no per-node work, no reboots and no job drain. Config-only mitigation if you cannot upgrade immediately: disable LXCA-to-XCC single sign-on and fall back to direct XCC authentication, accepting the credential-management cost. Independently, administer LXCA and XCC from a dedicated browser profile or privileged access workstation so a crafted link cannot reach a live management session.","references":["https://support.lenovo.com/us/en/product_security/LEN-154748","https://nvd.nist.gov/vuln/detail/CVE-2024-45101"],"status":"curated"},{"id":"CVE-2024-7726","cve":"CVE-2024-7726","aliases":["GHSA-3hh8-94j4-62rh","Kioxia JTAG"],"title":"Kioxia CM6 (GPK5 and earlier), PM6 (BD0D and earlier), PM7 (C40A and earlier) enterprise NVMe/SAS SSDs - unauthenticated JTAG debug port: An open, unauthenticated JTAG debug port on the drive's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Kioxia CM6 (GPK5 and earlier), PM6 (BD0D and earlier), PM7 (C40A and earlier) enterprise NVMe/SAS SSDs…","year":"2024","cvss_score":6.8,"severity":"medium","kev":false,"impact":"An open, unauthenticated JTAG debug port on the drive's PCB exposes both main SoC CPU cores - and the enclosure cutout is wide enough that you do not even need to open the drive to reach it. With a cheap ARM JTAG probe an attacker executes arbitrary code on the controller, reads firmware and memory, and bypasses the RSA firmware signature check at boot. These are Kioxia's flagship DATACENTER drives; CM6/PM6/PM7 are standard fitment in the exact GPU and AI-server platforms this database is about. BREAKS TENANT HANDOFF and does it in the worst direction: an attacker who owns the controller can read everything regardless of Opal state, can make sanitize report success while preserving data, and can attempt to leave an implant that survives every reimage you perform. Note the researchers' own caveat - fully PERSISTENT firmware modification additionally requires a shared secret used to compute the firmware MAC, so persistence is not demonstrated as trivially achievable, but transient full control of the controller is.","attack_vector":"Anyone with brief physical access to the drive and a low-cost JTAG probe. The enclosure does not have to be opened, so this is minutes of unsupervised contact, not a lab teardown - a rack tech, a colo neighbour with cage access, a courier in the RMA path, a decommission handler, or anyone in the supply chain before the drive reached you.","remediation":"UNPATCHABLE on the affected SKUs - the Google advisory records no fixed version, and an exposed JTAG pad is a board-design property that no firmware update removes. This is therefore a physical-security and procurement control, not a patch: enforce tamper-evident seals and chain of custody on CM6/PM6/PM7 media, never return or resell a drive that held tenant data (destroy it), and treat any of these drives with an unexplained custody gap as compromised at controller level rather than reimaging and re-renting it. Because you cannot trust the controller's own attestations on these drives, run LUKS/dm-crypt with an operator-held key so the drive only ever sees ciphertext and reclaim means destroying your key. Raise it with Kioxia as a procurement question for future SKUs - the fix has to come in hardware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-7726","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2025-0012","cve":"CVE-2025-0012","aliases":[],"title":"AMD - overlap between segmented reverse map table (RMP) and SMM memory: MULTI-TENANT ISOLATION: Improper handling of overlap between the segmented RMP and System Management Mode…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD - overlap between segmented reverse map table (RMP) and SMM memory","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper handling of overlap between the segmented RMP and System Management Mode memory lets a privileged attacker corrupt or partially infer SMM memory. SMM is the most privileged execution context on x86 - above the hypervisor - so reaching it from the RMP path is both a route to total platform control and, in the inference direction, a leak out of the one context nothing else can inspect.","attack_vector":"Local, privileged attacker on a platform using segmented RMP (large-memory SEV-SNP configurations).","remediation":"Fixed in AMD firmware/AGESA, delivered as an OEM SBIOS package with **one to six months of OEM lag** and a drained-node power cycle. Because it touches the SEV-SNP trust boundary, refresh VCEK certificates and update tenant attestation policy after the TCB version moves. Segmented RMP is used on very large memory configurations - exactly the shape of an AI training host - so do not assume this is an edge case on a GPU fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0012","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-14304","cve":"CVE-2025-14304","aliases":[],"title":"Motherboards from ASRock and its subsidiaries ASRockRack and ASRockInd built on Intel 500-series chipsets: A DMA-capable PCIe device reads and writes system memory without restriction. On a GPU…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Motherboards from ASRock and its subsidiaries ASRockRack and ASRockInd built on Intel 500-series chipsets","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"A DMA-capable PCIe device reads and writes system memory without restriction. On a GPU platform this is a direct hit on the assumption the whole tenant-isolation model rests on: IOMMU enforcement is what stops a device - or a tenant-controlled device - from reading arbitrary host memory, and it is what makes GPU passthrough to a VM safe. With it off, an attacker with a malicious or reprogrammed peripheral, or with control of a passed-through device, reads host kernel memory, extracts keys, and writes to memory to escalate. For any operator doing GPU passthrough or accepting tenant-supplied hardware, this invalidates the isolation guarantee. The IOMMU is not properly enabled, so the protection that is supposed to confine what a PCIe device can reach in system memory is simply not active.","attack_vector":"Physical access sufficient to attach a DMA-capable PCIe device - which includes Thunderbolt/USB4 ports, open PCIe slots, and any peripheral in a colocation or edge environment where the chassis is not under your exclusive control. Also relevant wherever a device is passed through to an untrusted guest, since the confinement that passthrough relies on is absent.","remediation":"BIOS/UEFI update from ASRock, then verify - do not assume. After flashing, confirm the IOMMU is actually active at runtime: check for DMAR/IVRS tables and that the kernel reports IOMMU groups, rather than trusting a BIOS setting. ASRock and ASRockRack publish a security page, but it sits behind an Imperva challenge and returns unreadable content to automated clients, so advisory tracking has to be manual or via NVD and TWCERT. Config-only hardening in the meantime: enable IOMMU explicitly in BIOS and on the kernel command line, and disable unused Thunderbolt/PCIe hot-plug paths.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-14304","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/14xxx/CVE-2025-14304.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23216","cve":"CVE-2025-23216","aliases":[],"title":"Argo CD: Secret values exposed in error messages and the diff view when an invalid Secret is synced","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Argo CD","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Secret values exposed in error messages and the diff view when an invalid Secret is synced","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + rotate any secret rendered in the UI","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23216"],"status":"curated"},{"id":"CVE-2025-26465","cve":"CVE-2025-26465","aliases":[],"title":"OpenSSH (client): Machine-in-the-middle against the client when VerifyHostKeyDNS is enabled","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (client)","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Machine-in-the-middle against the client when VerifyHostKeyDNS is enabled","attack_vector":"Unauthenticated network (MITM)","remediation":"Package update; no reboot. Audit client configs for VerifyHostKeyDNS","references":["https://access.redhat.com/security/cve/CVE-2025-26465"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-2713","cve":"CVE-2025-2713","aliases":[],"title":"gVisor: runsc mishandles file access permissions, letting unprivileged users read restricted files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"runsc mishandles file access permissions, letting unprivileged users read restricted files; local privilege escalation","attack_vector":"Any tenant workload inside the sandbox","remediation":"Upgrade runsc; restart sandboxed pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2713"],"status":"curated"},{"id":"CVE-2025-33215","cve":"CVE-2025-33215","aliases":[],"title":"SNAP-4 container (BlueField storage): DoS of the storage dataplane","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"SNAP-4 container (BlueField storage)","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS of the storage dataplane","attack_vector":"Authenticated network attacker","remediation":"Bump SNAP container image; restart DPU storage service","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33215","https://github.com/NVIDIA/product-security/tree/main/2026/5744"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-823"]},{"id":"CVE-2025-33216","cve":"CVE-2025-33216","aliases":[],"title":"SNAP-4 container: DoS (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"SNAP-4 container","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"DoS (integer overflow)","attack_vector":"Network-accessible client","remediation":"Bump SNAP container image","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33216","https://github.com/NVIDIA/product-security/tree/main/2026/5744"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-131"]},{"id":"CVE-2025-35979","cve":"CVE-2025-35979","aliases":["guest-mode predictor state leak","VMX non-root transient execution disclosure"],"title":"Intel processors, exploitable from within VMX non-root (guest) operation - INTEL-SA-01420: Shared microarchitectural predictor state influences transient execution inside guest (VMX non-root)…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors, exploitable from within VMX non-root (guest) operation - INTEL-SA-01420","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Shared microarchitectural predictor state influences transient execution inside guest (VMX non-root) operation, letting unprivileged software in a VM observe data it should not. This is the shape of bug that matters most to anyone renting VMs on shared hosts: the leak is reachable from inside a guest by ordinary unprivileged code, targeting state shared with whatever else the host is running. For a neocloud running multiple tenant VMs per physical machine, it is a tenant-boundary issue by construction; for a bare-metal-per-tenant product it is contained to that tenant.","attack_vector":"Unprivileged software inside a guest VM. Intel rates the attack complexity as high and notes attack requirements must be present, so this is a capable-adversary scenario rather than a commodity exploit - but the position required is just 'a customer with a VM'.","remediation":"Microcode/BIOS update via OEM firmware - firmware flash, host reboot, job drain - plus hypervisor updates where the VMM must invoke the new controls. This is part of Intel's 2026 quarterly advisory batch, so bundle it with the other CVEs in that IPU rather than scheduling separately; the marginal cost of adding it to an existing firmware window is zero and the cost of its own window is a full fleet drain. No SMT decision attached. Verify by microcode revision and by the hypervisor's own mitigation reporting.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-35979","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01420.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-9708","cve":"CVE-2025-9708","aliases":[],"title":"Kubernetes C# client: Improper certificate validation in custom-CA mode enables MITM on the Kubernetes API connection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes C# client","year":"2025","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Improper certificate validation in custom-CA mode enables MITM on the Kubernetes API connection","attack_vector":"Unauthenticated network in a MITM position","remediation":"Upgrade any internal tooling built on the C# client","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-24234","cve":"CVE-2026-24234","aliases":[],"title":"TensorRT-LLM: SSRF via plugin loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"SSRF via plugin loading","attack_vector":"Tenant-supplied plugin reference","remediation":"Bump TensorRT-LLM; add egress policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24234","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:L/A:L","cwe":["CWE-918"]},{"id":"CVE-2026-27765","cve":"CVE-2026-27765","aliases":[],"title":"Intel Gaudi / vLLM hardware plugin: Malformed input to the Gaudi vLLM plugin crashes or wedges the serving process. On an inference fleet this is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi / vLLM hardware plugin","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Malformed input to the Gaudi vLLM plugin crashes or wedges the serving process. On an inference fleet this is a request-triggered outage of the model server rather than a data-confidentiality problem, but a single tenant can repeatedly take down a shared endpoint that is pinned to expensive accelerators.","attack_vector":"Anyone who can reach the vLLM endpoint with an authenticated request - so any tenant of a shared inference service, or any workload inside the cluster if the endpoint is not network-segmented.","remediation":"Upgrade the vLLM hardware plugin for Gaudi to 0.16.0 or later. Pure Python/userspace package update - restart the serving process, no node reboot, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27765","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01487.html"],"status":"curated"},{"id":"NCVD-2026-004-openbmc-bmcweb-mtls-client-certi","cve":null,"aliases":["AUTH-F6"],"title":"OpenBMC bmcweb mTLS client-certificate UPN validation: Where mTLS is configured, bmcweb matches the certificate's UPN by walking dot-separated labels with no bound.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenBMC bmcweb mTLS client-certificate UPN validation","year":"2026","cvss_score":6.8,"severity":"medium","kev":false,"impact":"Where mTLS is configured, bmcweb matches the certificate's UPN by walking dot-separated labels with no bound. A certificate issued for user@com authenticates against any host under that TLD, and a parent-domain certificate authenticates against every child deployment. For an operator running mTLS across a fleet - which is the hardened configuration, chosen by the most security-conscious teams - this means one certificate from anywhere in the hierarchy authenticates to every BMC in it, silently, with no failed-auth log to notice. It is the rare bug that punishes you specifically for having done the harder thing.","attack_vector":"Requires mTLS to be enabled on the BMC (not the default) and possession of any certificate the BMC's trust store chains to, including one issued for a parent domain or a different deployment. Network access to the BMC's HTTPS port.","remediation":"Unpatched at disclosure. An April 2026 commit fixed only case-insensitivity in the comparison and left the suffix-walking logic intact, so do not assume a recent bmcweb clears it. If you run mTLS on BMCs, audit which CAs are in each BMC's trust store and narrow them to a per-fleet issuing CA that signs nothing else - that is a config-only change and is the effective mitigation today. Do not put a broadly-scoped corporate CA in a BMC trust store. A real fix will require a BMC firmware flash once upstream lands one.","references":["https://seclists.org/fulldisclosure/2026/May/24","https://binreaper.pages.dev/posts/2026-05-27-bmcweb-disclosure/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2017-12334","cve":"CVE-2017-12334","aliases":[],"title":"Cisco NX-OS CLI: CLI command injection giving root-level execution on the switch OS for an authenticated admin — the same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS CLI","year":"2017","cvss_score":6.7,"severity":"medium","kev":false,"impact":"CLI command injection giving root-level execution on the switch OS for an authenticated admin — the same pattern later re-appeared as the exploited CVE-2024-20399","attack_vector":"Network, authenticated admin","remediation":"NX-OS upgrade with fabric failover; recurrence of the class means CLI-injection hardening on the switch OS is a standing item, not a one-off patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-12334"],"status":"curated"},{"id":"CVE-2018-12204","cve":"CVE-2018-12204","aliases":[],"title":"Intel Server Board / Server System / Compute Module platform firmware: Improper memory initialisation in platform sample/silicon reference firmware on Intel server boards, allowing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server Board / Server System / Compute Module platform firmware","year":"2018","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Improper memory initialisation in platform sample/silicon reference firmware on Intel server boards, allowing privilege escalation. Reference firmware defects propagate into whatever the OEM built on top of it, so the affected population is wider than Intel-branded boards.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12204","https://www.intel.com/content/www/us/en/security-center/advisory/INTEL-SA-00191.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-13787","cve":"CVE-2018-13787","aliases":[],"title":"SPI flash descriptor region configuration on a wide range of Supermicro boards: Any software running with sufficient privilege on the host operating system can rewrite the platform…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SPI flash descriptor region configuration on a wide range of Supermicro boards","year":"2018","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Any software running with sufficient privilege on the host operating system can rewrite the platform firmware. That is a UEFI/BIOS implant written from userland-reachable code - it survives disk wipes, OS reinstalls, and node reprovisioning, and it is invisible to every host-based security tool the operator runs. On bare-metal GPU rental this is the canonical tenant-persistence attack: a tenant with root on their leased node writes firmware, releases the node, and retains a foothold on hardware that is subsequently handed to someone else. X11S, X10, X9, X8SI, K1SP, C9X299, C7, B1, A2 and A1 families. The descriptor is what tells the chipset which flash regions the host CPU may write; Supermicro shipped it misconfigured so the OS could write firmware.","attack_vector":"Host-side privileged code execution - root on Linux or an equivalent. No network position on the management VLAN is required at all, which makes it the mirror image of the BMC bugs in this list: the threat comes from inside the node, from whoever you rented it to.","remediation":"BIOS/firmware flash with a Supermicro image that ships a locked descriptor region, per board family. Verification matters more than the flash here: after updating, confirm the descriptor is actually locked - Intel's chipsec `common.bios_wp` and descriptor checks will tell you - because the fix is a configuration inside the firmware image and it is easy to assume it landed when it did not. For a bare-metal rental fleet, add a descriptor/flash-lock check to the node turnup and node return pipeline rather than treating it as a one-time patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-13787","https://eclypsium.com/blog/firmware-vulnerabilities-in-supermicro-systems/","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2018/13xxx/CVE-2018-13787.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-0119","cve":"CVE-2019-0119","aliases":[],"title":"Intel Xeon D / Xeon Scalable system firmware, Server Board and Server System: A buffer overflow in system firmware across Xeon D and Xeon Scalable server boards, letting a privileged user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon D / Xeon Scalable system firmware, Server Board and Server System","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A buffer overflow in system firmware across Xeon D and Xeon Scalable server boards, letting a privileged user escalate into firmware. Affects the exact generations that built out the first wave of GPU datacenter capacity, and many of those nodes are still running.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0119","https://www.intel.com/content/www/us/en/security-center/advisory/INTEL-SA-00223.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-11157","cve":"CVE-2019-11157","aliases":["Plundervolt","V0ltPwn"],"title":"Intel SGX / dynamic voltage and frequency scaling interface: MULTI-TENANT ISOLATION: Undervolting the CPU through the privileged voltage-scaling MSR induces faults inside…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX / dynamic voltage and frequency scaling interface","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Undervolting the CPU through the privileged voltage-scaling MSR induces faults inside SGX enclaves, which is enough to extract AES keys and to corrupt RSA-CRT signatures into key-recovering faulty signatures. The attacker is the host operator - which is precisely the party SGX is supposed to be defended against - so this collapses confidential compute on the affected generations.","attack_vector":"Privileged local access on the host (ring 0). That is the SGX threat model: the platform owner attacking a tenant's enclave, or a compromised hypervisor attacking a confidential workload.","remediation":"BIOS/firmware update that locks the undervolting MSR, and a TCB recovery: enclaves must re-attest against the new SVN and any secrets sealed under the old TCB should be considered compromised. Because the lock is applied by platform firmware, this needs an OEM BIOS release and a full drain-and-reboot per node - OEM availability historically lagged Intel's advisory by months on server boards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00289.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-5676","cve":"CVE-2019-5676","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The driver loads Windows system DLLs without validating path or signature. A local user who can write to a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The driver loads Windows system DLLs without validating path or signature. A local user who can write to a searched directory gets code execution in a privileged NVIDIA process.","attack_vector":"Any local user with write access to a directory in the driver process's DLL search path.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-27815","https://nvd.nist.gov/vuln/detail/CVE-2019-5676"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5688","cve":"CVE-2019-5688","aliases":[],"title":"NVIDIA NVFlash / NVUFlash / GPUModeSwitch kernel driver (nvflash.sys, nvflsh32/64.sys): MULTI-TENANT ISOLATION: NVIDIA's signed flashing driver hands an administrator raw access to device memory…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NVFlash / NVUFlash / GPUModeSwitch kernel driver (nvflash.sys, nvflsh32/64.sys)","year":"2019","cvss_score":6.7,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: NVIDIA's signed flashing driver hands an administrator raw access to device memory and registers of arbitrary PCIe devices - including hardware NVIDIA does not own. That is a general-purpose physical-memory and device-register primitive wrapped in a legitimately signed driver, so it defeats driver signature enforcement and lets an admin (or malware that reached admin) reach across to the BMC, NICs, NVMe and other tenants' passthrough devices on the same host. Even on hosts that never flashed a GPU, its presence is a lasting bring-your-own-vulnerable-driver weapon.","attack_vector":"An authenticated administrator on the host, or any code that reached admin - which is exactly the boundary a signed kernel driver is supposed to hold.","remediation":"Update NVFlash/NVUFlash to 5.588.0 or later and GPUModeSwitch to the 2019-11 build, and remove the old nvflash.sys / nvflsh32.sys / nvflsh64.sys drivers from every host they were ever installed on. These are flashing tools, not runtime components: the durable fix is to keep the signed vulnerable driver off production hosts entirely and add its hashes to your driver blocklist, since it is a bring-your-own-vulnerable-driver primitive independent of NVIDIA. Removing the driver needs a reboot; no GPU firmware flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5688"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-11488","cve":"CVE-2020-11488","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): The BMC does not validate the RSA-1024 public key used to verify firmware signatures. That breaks the root of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The BMC does not validate the RSA-1024 public key used to verify firmware signatures. That breaks the root of trust for BMC firmware updates: an attacker who can push an update can install their own firmware image and persist below the host OS indefinitely. DGX-1 before 3.38.30, DGX-2 before 1.06.06.","attack_vector":"An attacker with access to the BMC's firmware update path - network reach to the management interface, or an administrator session obtained through the other bugs in this bulletin.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11488"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-27779","cve":"CVE-2020-27779","aliases":[],"title":"GRUB2 (cutmem command): The cutmem command was not gated by Secure Boot lockdown, so a privileged user could carve memory regions out…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (cutmem command)","year":"2020","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The cutmem command was not gated by Secure Boot lockdown, so a privileged user could carve memory regions out of the map GRUB hands the kernel. Used to remove the regions that hold verification state, which downgrades a verified boot to an unverified one without tripping anything.","attack_vector":"Local privileged user at the GRUB shell.","remediation":"grub2 package update + reboot. This is the lockdown-coverage class of bug: the fix is that the command is now refused when Secure Boot is on, so there is no config workaround short of a GRUB password.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-27779","https://access.redhat.com/security/cve/CVE-2020-27779"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-8692","cve":"CVE-2020-8692","aliases":[],"title":"Intel Ethernet 700 Series Controller firmware (access control): Insufficient access control inside 700-series NIC firmware lets a privileged host user escalate further or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet 700 Series Controller firmware (access control)","year":"2020","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insufficient access control inside 700-series NIC firmware lets a privileged host user escalate further or deny service. On bare-metal GPU rental the 'privileged host user' is the tenant, and the thing they are escalating into is firmware that the next tenant will inherit. This is the concrete mechanism behind the tenant-handoff problem: patching the host OS between customers does nothing about the NIC.","attack_vector":"A privileged local user on the host — in a bare-metal rental model, the customer with root.","remediation":"Flash 700-series firmware to 7.3 or later; cold power cycle. Operationally the stronger control is to reflash NIC firmware from a known-good image at every tenant handoff and verify the resulting version, rather than trusting whatever the previous tenant left behind. Related issues fixed in the same family: CVE-2020-8691, CVE-2020-8693, CVE-2019-0139, CVE-2019-0144.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8692","https://nvd.nist.gov/vuln/detail/CVE-2020-8693","https://nvd.nist.gov/vuln/detail/CVE-2019-0139"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-0157","cve":"CVE-2021-0157","aliases":[],"title":"Intel BIOS firmware: Insufficient control-flow management in Intel BIOS firmware lets a privileged user escalate. Broad advisory…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel BIOS firmware","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insufficient control-flow management in Intel BIOS firmware lets a privileged user escalate. Broad advisory covering many platform SKUs - check your specific board rather than assuming you are out of scope.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00562.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-0186","cve":"CVE-2021-0186","aliases":["SmashEx"],"title":"Intel SGX SDK (asynchronous exit / exception handling): MULTI-TENANT ISOLATION: SmashEx: an asynchronous exception delivered at the right moment during an enclave…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (asynchronous exit / exception handling)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: SmashEx: an asynchronous exception delivered at the right moment during an enclave entry/exit leaves the enclave's internal state inconsistent, and the SDK's own exception handling can then be steered into an in-enclave control-flow hijack. Because the host controls interrupt delivery, the attacker in this model is the platform - so it is a direct break of the confidential-compute promise, and it recovers enclave secrets in the published attack.","attack_vector":"Privileged host code that can inject exceptions/interrupts into a running enclave - i.e. the hypervisor or host OS on a confidential-compute node.","remediation":"Rebuild enclaves against a fixed SGX SDK and re-attest. This is a software fix in the SDK's AEX handling, so it needs a new enclave binary from the enclave author; no microcode, BIOS or reboot on the operator side, but also nothing the operator can do alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0186","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00548.html"],"status":"curated"},{"id":"CVE-2021-20225","cve":"CVE-2021-20225","aliases":[],"title":"GRUB2 (short-form option parser): Heap out-of-bounds write in the short-form option parser. Another arbitrary-write primitive inside the signed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (short-form option parser)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Heap out-of-bounds write in the short-form option parser. Another arbitrary-write primitive inside the signed bootloader, usable to load unsigned code with Secure Boot enabled.","attack_vector":"Local, via GRUB command line or grub.cfg.","remediation":"grub2 package update + reboot per node, followed by the dbx pass if you are actually revoking old binaries.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20225","https://access.redhat.com/security/cve/CVE-2021-20225"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-20233","cve":"CVE-2021-20233","aliases":[],"title":"GRUB2 (option quoting): Miscalculated buffer size when quoting options produces a heap out-of-bounds write. Same outcome: pre-boot…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (option quoting)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Miscalculated buffer size when quoting options produces a heap out-of-bounds write. Same outcome: pre-boot code execution and a Secure Boot bypass on a node the attacker already touched once.","attack_vector":"Local, via GRUB command line or a modified grub.cfg.","remediation":"grub2 package update + reboot per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20233","https://access.redhat.com/security/cve/CVE-2021-20233"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-29213","cve":"CVE-2021-29213","aliases":["HPESBHF04197"],"title":"HPE ProLiant Gen10 System ROM (security restriction bypass): A local bypass of security restrictions in the System ROM (BIOS) of ProLiant DL20 Gen10, ML30 Gen10 and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE ProLiant Gen10 System ROM (security restriction bypass)","year":"2021","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A local bypass of security restrictions in the System ROM (BIOS) of ProLiant DL20 Gen10, ML30 Gen10 and MicroServer Gen10 Plus. Firmware-level restriction bypasses matter more than their score suggests, because the restrictions being bypassed are the ones enforcing secure configuration at boot - and anything an attacker changes there persists across OS reinstalls. These are edge and utility SKUs rather than GPU nodes, but they are the boxes that end up as bastion hosts, PXE servers and DHCP/DNS for a GPU cluster, which makes them a useful staging point.","attack_vector":"Local to the host with elevated privileges - an administrator-level account on the operating system, or physical access during maintenance. Not reachable from the management VLAN.","remediation":"Update the System ROM to v2.52 or later. Unlike an iLO flash, a System ROM update only takes effect on the next host reboot, so it needs a maintenance window per node - though for this SKU set that is far cheaper than draining a GPU node. Stage the ROM through iLO or Service Pack for ProLiant and let it apply at the next scheduled restart. No config-only mitigation.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04197en_us","https://nvd.nist.gov/vuln/detail/CVE-2021-29213"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-23222","cve":"CVE-2022-23222","aliases":[],"title":"Linux kernel (eBPF verifier): kernel/bpf/verifier.c mishandles pointer types - unprivileged BPF to local root","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (eBPF verifier)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"kernel/bpf/verifier.c mishandles pointer types - unprivileged BPF to local root","attack_vector":"Any tenant process in a container where unprivileged BPF is enabled","remediation":"Livepatchable; otherwise drain + reboot. `kernel.unprivileged_bpf_disabled=1`","references":["https://access.redhat.com/security/cve/CVE-2022-23222"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-2586","cve":"CVE-2022-2586","aliases":[],"title":"Linux kernel (nf_tables): Cross-table use-after-free in nf_tables leading to local privilege escalation","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (nf_tables)","year":"2022","cvss_score":6.7,"severity":"medium","kev":true,"impact":"Cross-table use-after-free in nf_tables leading to local privilege escalation [KEV]","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-2586"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-25905","cve":"CVE-2022-25905","aliases":[],"title":"Intel oneAPI Data Analytics Library (oneDAL): An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI Data Analytics Library (oneDAL)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-25905","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-26052","cve":"CVE-2022-26052","aliases":[],"title":"Intel MPI Library (oneAPI HPC Toolkit): An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel MPI Library (oneAPI HPC Toolkit)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26052","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-26076","cve":"CVE-2022-26076","aliases":[],"title":"Intel oneAPI Deep Neural Network Library (oneDNN): An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI Deep Neural Network Library (oneDNN)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26076","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-26345","cve":"CVE-2022-26345","aliases":[],"title":"Intel oneAPI OpenMP runtime: An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI OpenMP runtime","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26345","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-26363","cve":"CVE-2022-26363","aliases":["XSA-402"],"title":"Xen (x86 PV): Insufficient care with non-coherent mappings - PV guest to host compromise","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 PV)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insufficient care with non-coherent mappings - PV guest to host compromise","attack_vector":"Tenant VM guest (PV)","remediation":"Hypervisor patch; livepatchable via Xen livepatch, otherwise host reboot + guest evacuation","references":["https://xenbits.xen.org/xsa/advisory-402.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-26421","cve":"CVE-2022-26421","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler runtime: An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI DPC++/C++ compiler runtime","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26421","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-26425","cve":"CVE-2022-26425","aliases":[],"title":"Intel oneAPI Collective Communications Library (oneCCL): An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI Collective Communications Library (oneCCL)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26425","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00674.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-28736","cve":"CVE-2022-28736","aliases":[],"title":"GRUB2 (chainloader): Use-after-free in grub_cmd_chainloader when a chainloaded image fails to start. Gives pre-boot code execution…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (chainloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Use-after-free in grub_cmd_chainloader when a chainloaded image fails to start. Gives pre-boot code execution and, in combination with the other 2022 bugs, a full Secure Boot bypass chain.","attack_vector":"Local, via GRUB command line or grub.cfg on a node the attacker has touched.","remediation":"grub2 package update + reboot per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28736","https://access.redhat.com/security/cve/CVE-2022-28736"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-31601","cve":"CVE-2022-31601","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An out-of-bounds write in the SmbiosPei module gives a highly privileged local attacker firmware-phase code…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the SmbiosPei module gives a highly privileged local attacker firmware-phase code execution. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31601","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-34301","cve":"CVE-2022-34301","aliases":["Three more bootloaders"],"title":"CryptoPro Secure Disk (signed UEFI bootloader): A Microsoft-signed bootloader that can be made to execute arbitrary pre-boot code. Same portable-bypass shape…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CryptoPro Secure Disk (signed UEFI bootloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A Microsoft-signed bootloader that can be made to execute arbitrary pre-boot code. Same portable-bypass shape as the Howyar case: the attacker brings the signed binary with them, so a fleet that never deployed this product is still exploitable simply because the firmware trusts the signature.","attack_vector":"Write access to the EFI System Partition on the target node.","remediation":"dbx revocation update pushed to every node via firmware update or OS vendor channel - not a package upgrade. Verify the dbx entry is present afterwards; revocation is the only fix because the vulnerable binary is signed and portable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34301","https://kb.cert.org/vuls/id/309662"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-34302","cve":"CVE-2022-34302","aliases":["Three more bootloaders"],"title":"New Horizon Datasys (signed UEFI bootloader): Signed bootloader with a built-in mechanism to bypass Secure Boot enforcement, described by researchers as…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"New Horizon Datasys (signed UEFI bootloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Signed bootloader with a built-in mechanism to bypass Secure Boot enforcement, described by researchers as more subtle than its siblings because it disables verification without obviously tampering with the chain. Yields a bootkit that measured boot will not flag.","attack_vector":"EFI System Partition write access on the target node.","remediation":"dbx revocation, delivered via firmware or OS vendor update. No package fix exists because the vulnerable artifact is a signed third-party binary.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34302","https://kb.cert.org/vuls/id/309662"],"status":"curated"},{"id":"CVE-2022-34303","cve":"CVE-2022-34303","aliases":["Three more bootloaders"],"title":"Eurosoft (UK) Ltd (signed UEFI bootloader): Signed UEFI bootloader containing a shell that executes arbitrary code, bypassing Secure Boot on any machine…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eurosoft (UK) Ltd (signed UEFI bootloader)","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Signed UEFI bootloader containing a shell that executes arbitrary code, bypassing Secure Boot on any machine trusting the Microsoft third-party CA.","attack_vector":"EFI System Partition write access.","remediation":"dbx revocation update. Track revocation-list version per node as a fleet health metric - this class recurs roughly annually and package inventories will never show it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34303","https://kb.cert.org/vuls/id/309662"],"status":"curated"},{"id":"CVE-2022-42281","cve":"CVE-2022-42281","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An out-of-bounds write in the FsRecovery module reaches firmware code execution from a highly privileged…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the FsRecovery module reaches firmware code execution from a highly privileged local account. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42281","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-0185","cve":"CVE-2023-0185","aliases":[],"title":"GPU Display Driver: DoS / data tampering (integer underflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"DoS / data tampering (integer underflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0185","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:C/C:L/I:L/A:H","cwe":["CWE-196"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0201","cve":"CVE-2023-0201","aliases":[],"title":"NVIDIA DGX-2 - SBIOS / SMM firmware: An out-of-bounds write in the Bds phase gives a privileged local user firmware code execution. This is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-2 - SBIOS / SMM firmware","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the Bds phase gives a privileged local user firmware code execution. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5449. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0201","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-118"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20567","cve":"CVE-2023-20567","aliases":[],"title":"AMD Radeon RX Vega M graphics driver installer - signature verification: The driver package launches AMDSoftwareInstaller.exe without validating its signature, so an attacker with…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Radeon RX Vega M graphics driver installer - signature verification","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The driver package launches AMDSoftwareInstaller.exe without validating its signature, so an attacker with admin privileges can substitute the binary and get their code run by a trusted installer flow. It is a signed-update-chain failure rather than a memory-safety bug: the mechanism you use to keep drivers current is the mechanism that runs the attacker's payload.","attack_vector":"Local, requires admin privilege to place the substituted binary. Windows driver packaging.","remediation":"Update the AMD driver package. Relevant only where you deploy AMD's Windows driver installer; Linux ROCm deployments are unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20567","https://www.amd.com/en/resources/product-security.html"],"status":"curated"},{"id":"CVE-2023-24932","cve":"CVE-2023-24932","aliases":["BlackLotus"],"title":"Windows Boot Manager (Secure Boot bypass): The bypass the BlackLotus UEFI bootkit used in the wild. An attacker with admin or physical access installs a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Windows Boot Manager (Secure Boot bypass)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The bypass the BlackLotus UEFI bootkit used in the wild. An attacker with admin or physical access installs a bootkit that survives OS reinstall and disk replacement, disables Secure Boot enforcement from inside the boot chain, and hides from every in-OS security agent. On a mixed Windows/Linux estate this also poisons attestation for anything downstream.","attack_vector":"Local administrator or physical access - which on bare metal means any tenant that rented the node, and on a colo floor means anyone with remote-hands.","remediation":"The most operationally painful entry in this cluster. The security update alone does nothing: Microsoft shipped it behind a manual, multi-stage opt-in requiring boot manager updates, revocation of the old boot manager, and a UEFI CA/dbx update, staged over roughly two years precisely because enabling revocation early bricks machines that still boot old media. Budget for a phased rollout with per-node verification, keep known-good recovery media that is still trusted, and expect to touch firmware settings on some boards. Not a patch-and-forget.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-24932","https://msrc.microsoft.com/update-guide/vulnerability/CVE-2023-24932"],"status":"curated"},{"id":"CVE-2023-25508","cve":"CVE-2023-25508","aliases":[],"title":"NVIDIA DGX BMC (AMI-derived management controller): The DGX-1 BMC's IPMI handler allows an authorised attacker to upload and download arbitrary files, which is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI-derived management controller)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The DGX-1 BMC's IPMI handler allows an authorised attacker to upload and download arbitrary files, which is enough to plant persistence on the controller or exfiltrate its configuration and credentials. A BMC compromise on a DGX gives an attacker power control, virtual media, serial console and a persistent foothold under the host OS on a node holding eight GPUs.","attack_vector":"Network access to the BMC management interface holding credentials at some authorised level. Whether that is 'remote' depends entirely on how genuinely isolated your OOB network is - in practice most fleets have a jump host, a DCIM integration or a monitoring collector that bridges it.","remediation":"Update the DGX BMC firmware bundle per bulletin 5458. Cost: BMC firmware usually updates without draining the GPUs, but the BMC resets and out-of-band access drops for a few minutes. Pair the patch with an actual audit of who can route to the BMC subnet - that control is worth more than the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25508","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-3264","cve":"CVE-2023-3264","aliases":[],"title":"CyberPower PowerPanel Enterprise DCIM: Hard-coded credentials in the DCIM platform","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CyberPower PowerPanel Enterprise DCIM","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Hard-coded credentials in the DCIM platform; chained with the other PowerPanel flaws for full DCIM takeover and, from there, control of every managed power device in the facility","attack_vector":"Network","remediation":"Upgrade PowerPanel Enterprise to 2.6.9; DCIM is often facility-operator-owned rather than tenant-owned, so a colocated neocloud may not control the remediation at all","references":["https://thehackernews.com/2023/08/multiple-flaws-in-cyberpower-and.html"],"status":"curated"},{"id":"CVE-2023-48733","cve":"CVE-2023-48733","aliases":[],"title":"EDK II / OVMF (UEFI Shell left enabled in downstream Ubuntu and LXD firmware builds): Not a memory-safety bug - a packaging default. The interactive UEFI Shell is compiled into the shipped…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II / OVMF (UEFI Shell left enabled in downstream Ubuntu and LXD firmware builds)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Not a memory-safety bug - a packaging default. The interactive UEFI Shell is compiled into the shipped firmware image, and from it an attacker who already has the OS can load unsigned code and walk straight past Secure Boot. For anyone running GPU workloads inside VMs on OVMF, this quietly voids the guest's Secure Boot guarantee and any attestation chain rooted in it, which is precisely the guarantee confidential-compute and tenant-isolation stories are sold on.","attack_vector":"An attacker who already has administrative control of the guest OS (or the VM's boot configuration) and can reach the UEFI Shell on the next boot. Local to the guest; no firmware flash needed by the attacker.","remediation":"Distribution package update, not an OEM BIOS flash - this is a build-flag fix in the edk2/OVMF firmware images shipped by the distro (Ubuntu, LXD). Update the ovmf/edk2 packages on your hypervisor hosts and restart the affected guests to pick up the new firmware blob; no host reboot required. Config workaround: remove the UEFI Shell from the boot order and enforce a firmware password / locked boot order in the guest's variable store, though a determined admin-level attacker inside the guest can often undo that.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48733","https://ubuntu.com/security/CVE-2023-48733"],"status":"curated"},{"id":"CVE-2023-49144","cve":"CVE-2023-49144","aliases":["INTEL-SA-01078"],"title":"Intel Server OpenBMC firmware (before egs-1.15-0 / bhs-0.27): An out-of-bounds read reachable by a privileged BMC user leaks memory contents across a scope boundary - CVSS…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server OpenBMC firmware (before egs-1.15-0 / bhs-0.27)","year":"2023","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An out-of-bounds read reachable by a privileged BMC user leaks memory contents across a scope boundary - CVSS v4 scores it 8.1, notably higher than the v3 6.7, because the disclosure crosses out of the vulnerable component. What comes back is BMC process memory, which on this stack means session tokens, credential material and configuration. Combined with the privilege-escalation entry from the same product line, an attacker with a modest BMC account has a path to reading things that let them keep the access permanently.","attack_vector":"Local access on the BMC with a privileged account. Requires an existing high-privilege BMC credential, so this is a post-compromise deepening tool rather than an entry point.","remediation":"Fixed in Intel Server OpenBMC egs-1.15-0 / bhs-0.27 and later - per-node out-of-band BMC firmware update via Intel platform packages, subject to OEM rebase lag on boards that derive from the same base. No config-only mitigation for the bug itself; limit the blast radius by minimizing the number of accounts holding BMC admin and by rotating BMC credentials after any suspected node compromise.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-49144","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01078.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21766","cve":"CVE-2024-21766","aliases":[],"title":"Intel oneAPI Math Kernel Library (oneMKL): An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel oneAPI Math Kernel Library (oneMKL)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21766","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01072.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-21857","cve":"CVE-2024-21857","aliases":[],"title":"Intel oneAPI compiler: An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI compiler","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21857","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01057.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-31073","cve":"CVE-2024-31073","aliases":[],"title":"Intel oneAPI Level Zero software: An uncontrolled search path in Level Zero lets an authenticated local user get code loaded at higher…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI Level Zero software","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled search path in Level Zero lets an authenticated local user get code loaded at higher privilege. Level Zero is the low-level runtime under oneAPI and under Intel GPU compute generally, so it sits in the same position in an Intel accelerator stack that the CUDA driver API occupies in an NVIDIA one.","attack_vector":"Local, authenticated. A user able to place a file on a path the runtime searches - which on a shared build host is a low bar.","remediation":"Update the oneAPI Level Zero package. Cost: package update and job restart; no driver reload or node drain. Audit writable directories on shared build and inference hosts as the standing control.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-31073"],"status":"curated"},{"id":"CVE-2024-42642","cve":"CVE-2024-42642","aliases":[],"title":"Micron Crucial MX500 series SSD, firmware M3CR046 - buffer overflow in the drive controller reachable from host ATA commands: Specially crafted ATA packets sent from the HOST to the drive…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Micron Crucial MX500 series SSD, firmware M3CR046 - buffer overflow in the drive controller reachable from host ATA…","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Specially crafted ATA packets sent from the HOST to the drive controller overflow a buffer in the firmware, with high confidentiality, integrity and availability impact. This is the shape of bug that matters most for bare-metal multi-tenancy: the attack surface is the ordinary storage command path, reachable from the operating system, not a JTAG pad or a soldering iron. A tenant with root on the machine can reach the controller's memory. TOP TENANT-HANDOFF RISK PATTERN: a departing tenant who achieves code execution on the drive controller can attempt to leave something behind that outlives your reimage entirely, because your reimage rewrites NAND contents and never touches controller firmware. Even short of a persistent implant, controller-level code execution means the drive's own reports about sanitize, lock state and encryption become worthless.","attack_vector":"A tenant with root (high privilege) on the bare-metal host, issuing crafted ATA commands down the normal storage path. No physical access, no chassis entry, no special hardware - this is reachable from a shell on the rented machine.","remediation":"Flash past M3CR046; Micron states the issue was fully remediated in December 2024 and firmware is on Crucial's MX500 support page. Drive offline for the flash, plus Crucial's own tooling. The broader point for an operator: the MX500 is a consumer SSD and has no business being tenant-writable media in a bare-metal fleet - if you find these in GPU nodes (they turn up as cheap boot/scratch drives in budget builds), the correct remediation is to replace them with datacenter SKUs, not just to patch. Where you cannot replace, deny tenants the ability to issue raw ATA/NVMe pass-through: do not hand out unfiltered block devices, and if the tenant workload does not need direct device access, put a virtualization or filesystem layer between them and the controller. Note also that Crucial Storage Executive, the management tool you would use to flash this, has its own separate installer DLL-preloading flaw (CVE-2025-71178) - patch the tool before you trust it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42642","https://www.crucial.com/support/ssd-support/mx500-support","https://github.com/VL4DR/CVE-2024-42642/tree/main"],"status":"curated"},{"id":"CVE-2024-45105","cve":"CVE-2024-45105","aliases":["LEN-165524"],"title":"Lenovo ThinkSystem UEFI/BIOS (SMM callout): A System Management Mode callout vulnerability in ThinkSystem UEFI - SMM code calls out to memory it does not…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo ThinkSystem UEFI/BIOS (SMM callout)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A System Management Mode callout vulnerability in ThinkSystem UEFI - SMM code calls out to memory it does not control, letting a local attacker with elevated privileges execute code in SMM. SMM sits above the hypervisor and is invisible to it, so this is a persistent-implant primitive: what an attacker installs there is not removed by reimaging, disk replacement or hypervisor reinstall. The affected list runs to roughly 99 platforms and explicitly includes the GPU boxes - SR670 V2 and SR675 V3 - alongside SR630/SR650/SR645/SR665 V3 and the ThinkAgile appliances.","attack_vector":"Local to the host with elevated privileges - root or administrator on the operating system. Not reachable from the management VLAN; the path is a tenant or workload that already holds privileged host access.","remediation":"UEFI/BIOS update on each affected node, per the per-model version table in LEN-165524. This is the expensive kind: the payload can be staged through XCC, but it applies only on the next host reboot, so it needs a drain of running training jobs and a maintenance window per node. For SR670 V2 / SR675 V3 GPU nodes that is real lost capacity - plan it as a rolling campaign against spare capacity, not an emergency sweep. No config-only mitigation exists for an SMM defect.","references":["https://support.lenovo.com/us/en/product_security/LEN-165524","https://nvd.nist.gov/vuln/detail/CVE-2024-45105"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-45774","cve":"CVE-2024-45774","aliases":[],"title":"GRUB2 (JPEG parser): Out-of-bounds write in GRUB's JPEG parser from a crafted image","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (JPEG parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Out-of-bounds write in GRUB's JPEG parser from a crafted image; possible Secure Boot bypass — part of the 2024-2025 GRUB2 vulnerability wave (73 issues)","attack_vector":"Local, ESP write","remediation":"Coordinated GRUB2 + shim + dbx rollout across every distro image in the fleet. The distro-by-distro fan-out is the real cost: a neocloud offering multiple guest images has to rebuild all of them","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45774"],"status":"curated"},{"id":"CVE-2024-45775","cve":"CVE-2024-45775","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (commands/extcmd): A failed allocation goes unchecked, so GRUB proceeds on a NULL pointer and its state becomes…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (commands/extcmd)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A failed allocation goes unchecked, so GRUB proceeds on a NULL pointer and its state becomes attacker-influenced. Practical outcome is the same as the rest of this batch: a foothold below the OS that reimaging does not remove.","attack_vector":"Local, through GRUB command processing on a node the attacker can already influence.","remediation":"grub2 package update + reboot. This whole 2025 batch lands as one distro update - do it as a single pass rather than 20 separate changes, and remember to rebuild the netboot image.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45775","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45778","cve":"CVE-2024-45778","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (BFS filesystem parser): Integer overflow in the BeFS parser leads to heap corruption. Nobody runs BFS in production, which is exactly…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (BFS filesystem parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Integer overflow in the BeFS parser leads to heap corruption. Nobody runs BFS in production, which is exactly the point: GRUB compiles in filesystem modules you will never mount, and each one is reachable from an attacker-supplied disk image.","attack_vector":"Attacker-supplied filesystem image on a disk or virtual media the node will parse.","remediation":"grub2 package update + reboot. The structural fix is to build GRUB with only the filesystem modules you actually boot from - it deletes a whole recurring CVE class from your fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45778","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45780","cve":"CVE-2024-45780","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (tar filesystem parser): Integer overflow in the tarfs module writes out of bounds. Tar archives are a normal part of initrd and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (tar filesystem parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Integer overflow in the tarfs module writes out of bounds. Tar archives are a normal part of initrd and provisioning workflows, so this is more reachable in practice than the exotic filesystem parsers.","attack_vector":"Attacker-supplied tar image parsed by GRUB during boot.","remediation":"grub2 package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45780","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45781","cve":"CVE-2024-45781","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (UFS filesystem parser): Symlink name length is never validated, giving a heap out-of-bounds write in the UFS parser and a route to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (UFS filesystem parser)","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Symlink name length is never validated, giving a heap out-of-bounds write in the UFS parser and a route to circumventing Secure Boot.","attack_vector":"Attacker-supplied UFS image on an attached or virtual disk.","remediation":"grub2 package update + reboot. Red Hat noted no viable mitigation short of the update, so this is a patch-or-accept decision.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45781","https://access.redhat.com/security/cve/CVE-2024-45781"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47795","cve":"CVE-2024-47795","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler: An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI DPC++/C++ compiler","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47795","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01243.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-47976","cve":"CVE-2024-47976","aliases":["Solidigm SA-000563","improper access removal handling"],"title":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620 Opal) - access not properly revoked: The firmware mishandles the REMOVAL of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSDs with TCG Opal (DC P4510/P4511/P4610 Opal, D5-P4320/P4326 Opal, D5-P5316 Opal, D7-P5510/P5520/P5620…","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The firmware mishandles the REMOVAL of access - that is, revocation does not fully take effect. An authorization that should have ended when you deprovisioned the drive can still be usable. BREAKS TENANT HANDOFF precisely at the moment of handoff: the step where you revoke the previous tenant's access to a locking range is the step that fails, so credentials or authority granted to the departing customer keep working against the drive. Distinct from the access-control-validation bug on the same SKUs, and worth tracking separately because the failure is in de-provisioning rather than in the initial check - it is invisible to any test that only verifies that locking works when you set it up.","attack_vector":"An attacker with physical access to the drive plus some low-level privilege - realistically a departing tenant who retained drive credentials from their tenancy, or anyone who later obtains the physical media through RMA or decommission.","remediation":"Firmware flash, per SKU, drive offline via Solidigm Storage Tool - same firmware trains as the sibling Opal issue (VEV10294/VDV10194/VCV10394 for P4510/P4511/P4610 Opal, 3DV10132 for D5-P4320 Opal, 8DV10564 for D5-P4326 Opal, ACV10340 for D5-P5316 Opal, JCV10404 for D7-P5510, 9CV10410 for D7-P5520/P5620 Opal). Operationally, add a verification step your reclaim pipeline probably lacks: after revoking a locking range, actively re-test that the old credential is rejected, rather than assuming revocation succeeded because the command returned success. And do not let Opal credentials be the only thing standing between two customers - a per-tenant LUKS key you destroy at reclaim is revocation you can actually prove.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47976","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-53681","cve":"CVE-2024-53681","aliases":["nvmet subsysnqn overflow","NVMe-oF discovery NQN buffer handling"],"title":"Linux kernel - NVMe-oF target configfs, drivers/nvme/target/configfs.c: TENANT ISOLATION: nvmet_root_discovery_nqn_store() treated the subsystem NQN string as a fixed-size buffer…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel - NVMe-oF target configfs, drivers/nvme/target/configfs.c","year":"2024","cvss_score":6.7,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: nvmet_root_discovery_nqn_store() treated the subsystem NQN string as a fixed-size buffer even though it is dynamically allocated to the length of the existing string, so writing a longer discovery NQN overflows the allocation. The direct exposure is local to whoever administers the target's configfs, but in an operator context that includes storage orchestration and CSI drivers running with elevated privileges - so a compromised control-plane component turns a configuration write into kernel memory corruption on the node holding every tenant's namespaces. It also matters because NQN handling is exactly the surface an attacker probes when attempting NQN spoofing.","attack_vector":"Write an over-long NQN string to the discovery subsystem's configfs attribute on the target. Requires privileged access to the target host's configfs, which orchestration and storage-management agents routinely have. Not remotely reachable on its own, but a natural second stage after compromising a storage control-plane component.","remediation":"Host reboot / kernel upgrade on nvmet target nodes; fold into the same window as the other nvmet findings rather than scheduling separately. Config-side hardening that pays off regardless: keep the target's configfs off any path an untrusted orchestration component can reach, and audit which service accounts can write nvmet configuration.","references":["https://git.kernel.org/pub/scm/linux/security/vulns.git/plain/cve/published/2024/CVE-2024-53681.json","https://nvd.nist.gov/vuln/detail/CVE-2024-53681"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-20017","cve":"CVE-2025-20017","aliases":[],"title":"Intel oneAPI toolkit and component installers: An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI toolkit and component installers","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20017","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01285.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-20087","cve":"CVE-2025-20087","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler installer: The compiler installer sets permissions that let a local user modify installed files that later execute with…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI DPC++/C++ compiler installer","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"The compiler installer sets permissions that let a local user modify installed files that later execute with higher privilege. Same practical outcome as the search-path family: local privilege escalation on shared build and training nodes.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20087","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01285.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-20627","cve":"CVE-2025-20627","aliases":[],"title":"Intel oneAPI DPC++/C++ compiler: An uncontrolled library search path: the component loads a shared library by name from a directory a non-root…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel oneAPI DPC++/C++ compiler","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"An uncontrolled library search path: the component loads a shared library by name from a directory a non-root user can write. Anyone who can drop a file in that directory gets code execution in the context of whoever next runs the tool - which on an AI node is usually a privileged installer, a service account, or root.","attack_vector":"A local authenticated user on a node that has the toolkit installed. On shared build/dev nodes and on container images built from the Intel toolkits, that is a broad set of people.","remediation":"Upgrade the affected component and, just as importantly, audit directory permissions on already-provisioned nodes and container images - upgrading the package does not remove a writable directory an earlier install created. Userspace only: no reboot, no BIOS, no microcode. Rebuild base images rather than patching running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20627","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01285.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-20629","cve":"CVE-2025-20629","aliases":[],"title":"Intel E810 NVM Update Utility: Insecure inherited permissions in the NVM update utility - the tool you run to fix the NIC firmware issues…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 NVM Update Utility","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Insecure inherited permissions in the NVM update utility - the tool you run to fix the NIC firmware issues above - let a local authenticated user escalate. Worth noting because the remediation tool being the vulnerability is a genuine operational trap when you push it fleet-wide under automation.","attack_vector":"Authenticated local user on a node where the utility is staged.","remediation":"Use NVM Update Utility 4.60 or later, and check permissions on the staging directory your automation copies it into. Userspace tool, no reboot for the tool update itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20629","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01295.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2025-23299","cve":"CVE-2025-23299","aliases":[],"title":"NVIDIA ConnectX / BlueField: Privilege escalation leading to arbitrary code execution via the management interface","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA ConnectX / BlueField","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Privilege escalation leading to arbitrary code execution via the management interface","attack_vector":"Local","remediation":"NIC/DPU firmware update","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23299"],"status":"curated","fleet":{"ubiquity":"Universal on RDMA/InfiniBand fleets - ConnectX is the standard NIC in GPU clusters","remediation_pain":"`firmware-flash` on every NIC; requires a node reboot and in some fleets a maintenance window per rack","pain_class":"firmware-flash","why_fleet_wide":"Arbitrary code execution in the NIC/DPU management interface: the NIC sees all inter-node RDMA traffic for training jobs, so a compromise reads or corrupts gradients across tenants"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"]},{"id":"CVE-2025-23337","cve":"CVE-2025-23337","aliases":[],"title":"HGX / DGX GB200, GB300, B300 (BMC -> HMC): BMC admin can pivot to the HMC as administrator","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"HGX / DGX GB200, GB300, B300 (BMC -> HMC)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"BMC admin can pivot to the HMC as administrator -> code execution and privesc on the GPU baseboard controller","attack_vector":"Operator or attacker holding BMC admin on the mgmt network","remediation":"Flash BMC + HMC firmware out-of-band; treat BMC admin as a tier-0 credential and rotate","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=NVIDIA%20DGX%20BMC"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1244"]},{"id":"CVE-2025-33190","cve":"CVE-2025-33190","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A second out-of-bounds write in SROOT firmware, reachable from a privileged local account, reaching firmware…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A second out-of-bounds write in SROOT firmware, reachable from a privileged local account, reaching firmware code execution. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33190","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33231","cve":"CVE-2025-33231","aliases":[],"title":"CUDA Toolkit: Code exec via path manipulation on library load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"Code exec via path manipulation on library load","attack_vector":"Local user","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33231","https://github.com/NVIDIA/product-security/tree/main/2026/5755"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-427"]},{"id":"CVE-2025-5187","cve":"CVE-2025-5187","aliases":[],"title":"Kubernetes (kube-apiserver): A node can delete itself, and cascade-delete other objects, by adding an OwnerReference","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2025","cvss_score":6.7,"severity":"medium","kev":false,"impact":"A node can delete itself, and cascade-delete other objects, by adding an OwnerReference","attack_vector":"A compromised node / kubelet credential","remediation":"Rolling control-plane upgrade; tighten the NodeRestriction admission plugin","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2018-7113","cve":"CVE-2018-7113","aliases":["HPESBHF03894"],"title":"HPE iLO 5 (firmware update security restriction bypass): Bypass of the security restrictions that guard iLO 5 firmware updates. The severity number undersells what…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 5 (firmware update security restriction bypass)","year":"2018","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Bypass of the security restrictions that guard iLO 5 firmware updates. The severity number undersells what this is: the firmware-update gate is the control that stops an attacker from writing their own image to the service processor. Defeat it and the attacker installs persistent BMC firmware of their choosing - the deepest and most durable implant available on the node, below the hypervisor and untouched by any host reimage. On a bare-metal cloud, a node whose iLO firmware was replaced by a previous tenant never comes clean again through normal reprovisioning.","attack_vector":"Local exploitation - an attacker who already has a foothold on the node or its iLO context, not an unauthenticated remote attacker. The realistic path in a bare-metal fleet is a tenant with host-level access using it during their tenancy to leave something behind for the next one.","remediation":"Flash iLO 5 to v1.37 or later - out-of-band, per-node, no host reboot and no job drain. Complementary control that matters more than the version number: enable and actually check the iLO firmware integrity/attestation features HPE exposes, and verify iLO firmware version and measurement as part of node reprovisioning between tenants rather than trusting that a wipe covered it.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03894en_us","https://nvd.nist.gov/vuln/detail/CVE-2018-7113"],"status":"curated"},{"id":"CVE-2021-0060","cve":"CVE-2021-0060","aliases":[],"title":"Intel SPS (HECI subsystem compartmentalisation): Insufficient compartmentalisation in the HECI interface - the host-to-management-engine channel - on Server…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SPS (HECI subsystem compartmentalisation)","year":"2021","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Insufficient compartmentalisation in the HECI interface - the host-to-management-engine channel - on Server Platform Services firmware. HECI is the door between the OS and the management engine, so weak compartmentalisation there means host-side code reaches further into the engine than it should.","attack_vector":"Local access on the host with the ability to talk to the HECI device.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0060","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00470.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1076","cve":"CVE-2021-1076","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Improper access control in the kernel-mode layer on both Windows and Linux, producing disclosure or data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Improper access control in the kernel-mode layer on both Windows and Linux, producing disclosure or data corruption. Data corruption from a driver-level access control gap is the quiet kind of failure - silently wrong training runs rather than a visible crash. Debian and Gentoo shipped it as a security update.","attack_vector":"Any local user or GPU container on the host.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1076"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1077","cve":"CVE-2021-1077","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Reference-count mishandling on a driver resource in the R450 and R460 branches. Refcount bugs of this shape…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Reference-count mishandling on a driver resource in the R450 and R460 branches. Refcount bugs of this shape typically become use-after-free with effort; NVIDIA rates the confirmed impact as denial of service.","attack_vector":"Any local user or GPU container with device access.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1077"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-3294","cve":"CVE-2022-3294","aliases":[],"title":"Kubernetes (kube-apiserver): Node address not verified when proxying","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2022","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Node address not verified when proxying; a user who can modify Node objects reaches control-plane-only endpoints","attack_vector":"Cluster user able to patch Node objects","remediation":"Rolling control-plane upgrade; restrict Node patch RBAC","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3294"],"status":"curated"},{"id":"CVE-2023-0198","cve":"CVE-2023-0198","aliases":[],"title":"GPU Display Driver: Local privesc (kernel buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Local privesc (kernel buffer overflow)","attack_vector":"Any tenant with a container; vGPU guest","remediation":"Driver + vGPU Manager upgrade; drain + reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0198","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-31015","cve":"CVE-2023-31015","aliases":[],"title":"DGX H100 BMC (REST): Privesc via auth flaw","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Privesc via auth flaw","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:H/A:L","cwe":["CWE-287"]},{"id":"CVE-2023-31034","cve":"CVE-2023-31034","aliases":[],"title":"DGX A100 SBIOS: Integer overflow in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Integer overflow in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31034","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:C/C:L/I:L/A:H","cwe":["CWE-190"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34473","cve":"CVE-2023-34473","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC hard-coded credentials): Hard-coded credentials inside the BMC firmware. Once extracted from a publicly downloadable image, they work…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC hard-coded credentials)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Hard-coded credentials inside the BMC firmware. Once extracted from a publicly downloadable image, they work on every BMC running that build regardless of what passwords the operator set, so your entire fleet shares an authentication backdoor you cannot rotate. That converts a per-node authentication story into a single fleet-wide key, and BMC access means power control, console, virtual media and firmware.","attack_vector":"Adjacent network reachability to the BMC plus a valid user session and some interaction, per AMI's vector. The credential itself is obtained offline by unpacking a firmware image - no access to your systems is needed for that half.","remediation":"Firmware flash to SPx_12.2 / SPx_13.0 or later. This one has been fixed since early SPx builds, so the operator task is an audit: enumerate the actual running BMC firmware version across the fleet and find the SKUs still on a pre-fix ODM image - typically older or white-box nodes whose vendor stopped publishing BMC updates. There is no config-only fix, because the credential is baked into the image; the only compensating control is hard network isolation of the BMC plane so the credential has nothing to authenticate against.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34473"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-4622","cve":"CVE-2023-4622","aliases":[],"title":"Linux kernel (AF_UNIX): Use-after-free in unix_stream_sendpage - local privilege escalation, no capabilities needed","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (AF_UNIX)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Use-after-free in unix_stream_sendpage - local privilege escalation, no capabilities needed","attack_vector":"Any tenant process in a container","remediation":"Livepatchable; otherwise drain + reboot. Note this one needs no CAP_NET_ADMIN, so userns hardening does not help","references":["https://access.redhat.com/security/cve/CVE-2023-4622"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52819","cve":"CVE-2023-52819","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd): An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd)","year":"2023","cvss_score":6.6,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: Fix UBSAN array-index-out-of-bounds for Polaris and Tonga","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52819","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-28224","cve":"CVE-2024-28224","aliases":[],"title":"Ollama: DNS rebinding grants a remote page full API access","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ollama","year":"2024","cvss_score":6.6,"severity":"medium","kev":false,"impact":"DNS rebinding grants a remote page full API access","attack_vector":"Browser of anyone on a network with an Ollama host","remediation":"Upgrade past 0.1.29; bind loopback","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28224"],"status":"curated"},{"id":"CVE-2025-0037","cve":"CVE-2025-0037","aliases":[],"title":"AMD Versal Adaptive SoC - PLM runtime services address validation: MULTI-TENANT ISOLATION: The Platform Loader and Manager firmware on AMD Versal devices does not validate…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Versal Adaptive SoC - PLM runtime services address validation","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The Platform Loader and Manager firmware on AMD Versal devices does not validate addresses when executing runtime services, so a caller reaches isolated or protected memory spaces. On Versal parts the PLM is the root of trust and the isolation enforcer for the device's partitions - bypassing its address checks means crossing whatever partition boundary the design relies on, which on a multi-tenant SmartNIC or accelerator card is the tenant boundary.","attack_vector":"Local to the device, via PLM runtime service calls.","remediation":"Fixed in updated PLM firmware from AMD/Xilinx, applied as a device firmware image plus a card reset. Reaches you through whoever built the card, so expect integrator lag on top of AMD's release. Track which Versal-based cards are in your fleet and who owns their firmware pipeline.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0037","https://www.amd.com/en/resources/product-security.html"],"status":"curated"},{"id":"CVE-2025-0038","cve":"CVE-2025-0038","aliases":[],"title":"AMD Zynq UltraScale+ - CSU runtime service address validation in PMU firmware: MULTI-TENANT ISOLATION: The PMU firmware on Zynq UltraScale+ devices does not validate addresses when…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zynq UltraScale+ - CSU runtime service address validation in PMU firmware","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The PMU firmware on Zynq UltraScale+ devices does not validate addresses when executing Configuration Security Unit runtime services, giving access to isolated or protected memory. Companion to the Versal PLM issue and the same shape: the firmware component that enforces isolation on the device can be steered outside its own boundaries.","attack_vector":"Local to the device, via CSU runtime service calls through the PMU firmware.","remediation":"Fixed in updated PMU firmware from AMD/Xilinx. Device firmware update plus card reset, gated on your card integrator shipping it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0038","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-21692","cve":"CVE-2025-21692","aliases":[],"title":"Linux kernel (net/sched ETS): Out-of-bounds indexing in the ETS qdisc - memory corruption from CAP_NET_ADMIN","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched ETS)","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Out-of-bounds indexing in the ETS qdisc - memory corruption from CAP_NET_ADMIN","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2025-21692"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-2884","cve":"CVE-2025-2884","aliases":[],"title":"TCG TPM 2.0 reference implementation (CryptHmacSign): Out-of-bounds read in the reference implementation's HMAC signing helper because the signature scheme is not…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TCG TPM 2.0 reference implementation (CryptHmacSign)","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the reference implementation's HMAC signing helper because the signature scheme is not validated against the key's algorithm. Reads past the buffer can disclose TPM-internal memory - the one place in the system that is supposed to be opaque. Because this is the reference code, the same bug propagates into every fTPM, software TPM and vendor TPM derived from it, which is most of them.","attack_vector":"Local user able to issue TPM commands. On a shared or bare-metal node, that is any tenant.","remediation":"Update to TPM 2.0 reference implementation 1.83 or later - in practice that arrives as a platform firmware/BIOS update, a swtpm/libtpms package update for virtualised TPMs, or nothing at all if your TPM vendor has not rebased. Virtualised TPMs are the easy case (package update plus VM restart); silicon and fTPM are a per-node firmware flash. Check both paths separately; fleets usually have some of each.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2884","https://trustedcomputinggroup.org/trusted-computing-group-vulnerability-response/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-51481","cve":"CVE-2025-51481","aliases":[],"title":"Dagster (gRPC `get_notebook_data`): Local file inclusion — read arbitrary files","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dagster (gRPC `get_notebook_data`)","year":"2025","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Local file inclusion — read arbitrary files","attack_vector":"Attacker with access to the Dagster gRPC server, i.e. a co-tenant in a flat network","remediation":"Upgrade past 1.10.14","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-51481"],"status":"curated"},{"id":"CVE-2026-47619","cve":"CVE-2026-47619","aliases":[],"title":"NVIDIA Dynamo: Improper access control on privileged operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":6.6,"severity":"medium","kev":false,"impact":"Improper access control on privileged operations","attack_vector":"Tenant with cluster network access","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47619","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-1357"]},{"id":"CVE-2018-12207","cve":"CVE-2018-12207","aliases":["iTLB multihit","Machine Check Error on Page Size Change","MCEPSC","No eXcuses"],"title":"Intel Core and Xeon CPUs - INTEL-SA-00210: This one is availability, not confidentiality, and it is the most operationally brutal item in the family. A…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Core and Xeon CPUs - INTEL-SA-00210","year":"2018","cvss_score":6.5,"severity":"medium","kev":false,"impact":"This one is availability, not confidentiality, and it is the most operationally brutal item in the family. A malicious guest changes page size in its own page tables so that the instruction TLB holds multiple hits for one address; the CPU raises an unrecoverable machine check and the ENTIRE HOST hangs. One tenant VM, using only its own memory mappings, hard-kills the physical machine - and every other tenant on it. On a GPU host that means every co-resident training job loses its in-flight state back to the last checkpoint, and the node needs a physical or BMC-driven power cycle. It is a one-guest, no-privilege denial of service against the whole box.","attack_vector":"A privileged user inside any guest VM - i.e. root in a tenant's own VM, which on a bare-metal or VM-rental product is something every customer legitimately has. Does not require SMT or core sharing. Container tenants cannot reach it (no control of page tables); VM tenants can.","remediation":"Hypervisor patch, plus microcode on some parts. KVM's fix restricts large pages to non-executable mappings and splits huge pages down to 4K the moment a guest executes from them. That is the expensive part: losing 2MB/1GB EPT pages for guest code raises TLB pressure and costs memory-access performance on large-footprint guests, and it burns extra host memory on page tables. Linux exposes kvm.nx_huge_pages=off to reclaim it - do not set that on a multi-tenant host. If you rent whole machines to one tenant at a time, the blast radius is only that tenant and you may reasonably leave huge pages on. Recent Xeon generations are not affected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12207","https://docs.kernel.org/admin-guide/hw-vuln/multihit.html","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00210.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-3979","cve":"CVE-2018-3979","aliases":[],"title":"Nouveau display driver (in-tree Linux nouveau, NV117): Remote denial of service against a workstation or node running the open-source Nouveau driver: a crafted…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Nouveau display driver (in-tree Linux nouveau, NV117)","year":"2018","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Remote denial of service against a workstation or node running the open-source Nouveau driver: a crafted pixel shader delivered through a web page wedges the GPU driver and takes the machine's graphics stack down. No code execution, but on a shared render or VDI host it is a free reboot for anyone who can get a browser to load their page.","attack_vector":"Anyone who can get a user on the host to open a web page - so effectively internet-reachable. No local account needed.","remediation":"This is the in-tree open-source Nouveau driver, not NVIDIA's proprietary stack. Update the distribution kernel (Ubuntu 18.04 shipped the vulnerable NV117 code) or, on GPU nodes, blacklist nouveau entirely and run the proprietary NVIDIA driver, which is what a compute fleet should be doing anyway. Kernel update means a node reboot.","references":["https://talosintelligence.com/vulnerability_reports/TALOS-2018-0647","https://nvd.nist.gov/vuln/detail/CVE-2018-3979"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-1000008","cve":"CVE-2019-1000008","aliases":[],"title":"Helm: Path traversal in `helm fetch --untar` writes outside the target directory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Path traversal in `helm fetch --untar` writes outside the target directory","attack_vector":"Malicious chart","remediation":"Upgrade Helm on all CI and operator machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1000008"],"status":"curated"},{"id":"CVE-2019-11135","cve":"CVE-2019-11135","aliases":["TAA","TSX Asynchronous Abort","ZombieLoad v2"],"title":"Intel CPUs supporting TSX, including Cascade Lake Xeon Scalable - INTEL-SA-00270: Same class of in-flight data leak as MDS, but reached through a TSX transaction abort - and crucially it…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CPUs supporting TSX, including Cascade Lake Xeon Scalable - INTEL-SA-00270","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Same class of in-flight data leak as MDS, but reached through a TSX transaction abort - and crucially it works on Cascade Lake, the generation Intel shipped with the MDS_NO silicon fix. That is the operator lesson: the parts you bought specifically because they were 'the fixed ones' were still vulnerable. An attacker samples data another tenant or the host kernel is moving through the fill buffers. Only the attacker needs to use TSX; the victim does not. Cross-hyperthread, so on a shared-SMT fleet it is a direct tenant-boundary break yielding keys, tokens and credentials.","attack_vector":"Unprivileged local code on a TSX-capable CPU, in any guest or container. Cross-hyperthread attacks work because the fill buffers are shared between siblings.","remediation":"Microcode/BIOS update plus kernel patch - firmware flash, host reboot, job drain. Then a real choice: (a) tsx=off disables TSX entirely and fully closes it including cross-thread - this is the Linux default and the right call for a multi-tenant operator, and the only workloads that notice are the rare ones using hardware lock elision; or (b) keep TSX and rely on VERW buffer clearing, which leaves cross-hyperthread attacks possible unless you also disable SMT. Choose tsx=off. It costs almost nothing on AI infrastructure workloads and removes both the mitigation overhead and the SMT dilemma.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11135","https://docs.kernel.org/admin-guide/hw-vuln/tsx_async_abort.html","https://zombieloadattack.com/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-11246","cve":"CVE-2019-11246","aliases":[],"title":"Kubernetes (kubectl): `kubectl cp` path traversal from a malicious container tar overwrites files on the operator's workstation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"`kubectl cp` path traversal from a malicious container tar overwrites files on the operator's workstation","attack_vector":"Malicious image, triggered when an operator runs kubectl cp","remediation":"Upgrade kubectl on all operator and CI machines; not a cluster change","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11246"],"status":"curated"},{"id":"CVE-2019-11249","cve":"CVE-2019-11249","aliases":[],"title":"Kubernetes (kubectl): Follow-up incomplete fix for the kubectl cp traversal","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Follow-up incomplete fix for the kubectl cp traversal","attack_vector":"Malicious image","remediation":"Upgrade kubectl on operator and CI machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11249"],"status":"curated"},{"id":"CVE-2019-11250","cve":"CVE-2019-11250","aliases":[],"title":"Kubernetes (client-go): Bearer tokens logged at verbosity 7+","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (client-go)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Bearer tokens logged at verbosity 7+; credential disclosure via logs","attack_vector":"Anyone with log-pipeline read access","remediation":"Lower component verbosity; rotate service-account tokens; scrub log store","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11250"],"status":"curated"},{"id":"CVE-2019-11254","cve":"CVE-2019-11254","aliases":[],"title":"Kubernetes (kube-apiserver): YAML parsing CPU exhaustion in the apiserver","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"YAML parsing CPU exhaustion in the apiserver","attack_vector":"Any authorized cluster user","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11254"],"status":"curated"},{"id":"CVE-2019-16097","cve":"CVE-2019-16097","aliases":[],"title":"Harbor: Non-admin users create admin accounts via POST /api/users","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Non-admin users create admin accounts via POST /api/users; full registry takeover","attack_vector":"Any registered registry user","remediation":"Upgrade Harbor; audit the admin user list","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16097"],"status":"curated"},{"id":"CVE-2019-1890","cve":"CVE-2019-1890","aliases":[],"title":"Cisco Nexus 9000 ACI Mode Switch Software (fabric infrastructure VLAN): TENANT ISOLATION: the earlier instance of the same ACI class of bug — security validation on…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco Nexus 9000 ACI Mode Switch Software (fabric infrastructure VLAN)","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: the earlier instance of the same ACI class of bug — security validation on infrastructure-VLAN connection setup can be bypassed, letting an adjacent device join the fabric's internal VLAN. Worth tracking separately because the fixed releases differ and clusters running older ACI trains are exposed to this one and not the 2021 variant.","attack_vector":"Unauthenticated, adjacent — a device connected to a leaf port.","remediation":"ACI fabric software upgrade (APIC + switches). Staged, multi-hour. No config-only workaround beyond locking down unused ports.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1890"],"status":"curated"},{"id":"CVE-2019-9901","cve":"CVE-2019-9901","aliases":[],"title":"Envoy: No URL path normalization, so `something/../admin` bypasses access control","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"No URL path normalization, so `something/../admin` bypasses access control","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy; enable path normalization","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9901"],"status":"curated"},{"id":"CVE-2020-15136","cve":"CVE-2020-15136","aliases":[],"title":"etcd: Gateway TLS authentication applied only to endpoints found in DNS SRV records","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Gateway TLS authentication applied only to endpoints found in DNS SRV records; unauthenticated etcd access","attack_vector":"Unauthenticated network reaching etcd","remediation":"Rolling etcd upgrade; enforce mutual TLS on all etcd peers and clients","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15136"],"status":"curated"},{"id":"CVE-2020-24501","cve":"CVE-2020-24501","aliases":[],"title":"Intel E810 Ethernet Controller firmware: Buffer overflow in early E810 firmware, triggerable by an unauthenticated adjacent attacker for denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet Controller firmware","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Buffer overflow in early E810 firmware, triggerable by an unauthenticated adjacent attacker for denial of service. Worth carrying in an operator database because E810 cards that shipped in 2020-2021 server generations and were never NVM-updated are still in production fleets — NIC firmware is the layer operators most reliably forget to patch.","attack_vector":"Unauthenticated, adjacent — same L2 segment.","remediation":"Flash E810 firmware to 1.4.1.13 or later. Cold power cycle. Practically: audit your fleet's NVM versions first (`ethtool -i` reports the firmware-version string) — most operators discover a wide spread of versions and should batch the whole update rather than chase individual CVEs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24501"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-24511","cve":"CVE-2020-24511","aliases":[],"title":"Intel processors (shared resource isolation): MULTI-TENANT ISOLATION: Improper isolation of shared processor resources allowing information disclosure to a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (shared resource isolation)","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper isolation of shared processor resources allowing information disclosure to a local authenticated user. Fixed in the same microcode drop as the associated Atom domain-bypass issue; relevant to any multi-tenant host.","attack_vector":"Local authenticated code on the host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24511","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00464.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2020-24513","cve":"CVE-2020-24513","aliases":["Vector Register Sampling on Atom","SRBDS-family"],"title":"Intel Atom processors (domain-bypass transient execution): MULTI-TENANT ISOLATION: A domain-bypass transient execution flaw on Atom parts leaking information across…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Atom processors (domain-bypass transient execution)","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A domain-bypass transient execution flaw on Atom parts leaking information across privilege domains. Matters for edge inference boxes and storage/management appliances built on Atom silicon rather than for Xeon compute nodes - but those appliances often sit inside the trusted network of an AI datacenter.","attack_vector":"Local authenticated code on an affected Atom platform.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24513","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00465.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2020-8766","cve":"CVE-2020-8766","aliases":[],"title":"Intel SGX DCAP (datacenter attestation primitives): An improper conditions check in DCAP lets an unauthenticated adjacent attacker deny service to the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX DCAP (datacenter attestation primitives)","year":"2020","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An improper conditions check in DCAP lets an unauthenticated adjacent attacker deny service to the attestation path. In a confidential-compute fleet, killing attestation means new workloads cannot start and existing ones cannot renew - an availability failure that looks like a control-plane outage.","attack_vector":"Unauthenticated attacker with adjacent network access to the attestation service - so anything on the same network segment as your PCCS/quote-generation service.","remediation":"Upgrade SGX DCAP to 1.6 or later and keep the caching service off flat networks. Userspace service update and restart; no node reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8766","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00398"],"status":"curated"},{"id":"CVE-2021-0009","cve":"CVE-2021-0009","aliases":[],"title":"Intel Ethernet 800 Series Controller firmware: Out-of-bounds read in 800-series (E810 family) adapter firmware, reachable by an unauthenticated user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet 800 Series Controller firmware","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in 800-series (E810 family) adapter firmware, reachable by an unauthenticated user, causing denial of service. Listed separately from the later E810 issues because the fixed version target is different (1.5.3.0), so a fleet standardised on a mid-2021 NVM image is exposed to this one even if it is patched for the 2023-2025 batch.","attack_vector":"Unauthenticated, over the network to the adapter.","remediation":"Flash 800-series firmware to 1.5.3.0 or later; cold power cycle. In practice, set a single fleet-wide minimum NVM version well above all of these and enforce it in provisioning, rather than tracking each CVE's individual threshold.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0009"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1115","cve":"CVE-2021-1115","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference reachable through private IOCTLs, with NVIDIA noting the denial of service lands in a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"NULL dereference reachable through private IOCTLs, with NVIDIA noting the denial of service lands in a component beyond the driver itself - the failure propagates past the GPU stack.","attack_vector":"Any local unprivileged user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1115"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-21285","cve":"CVE-2021-21285","aliases":[],"title":"Docker / moby: Malformed image manifest crashes dockerd","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Malformed image manifest crashes dockerd; node-level DoS","attack_vector":"Malicious image","remediation":"Upgrade Docker Engine; daemon restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21285"],"status":"curated"},{"id":"CVE-2021-25735","cve":"CVE-2021-25735","aliases":[],"title":"Kubernetes (kube-apiserver): Node updates bypass a validating admission webhook, defeating node-level policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Node updates bypass a validating admission webhook, defeating node-level policy","attack_vector":"Cluster user able to update Node objects","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25735"],"status":"curated"},{"id":"CVE-2021-26341","cve":"CVE-2021-26341","aliases":[],"title":"AMD processors - transient execution beyond unconditional direct branches: MULTI-TENANT ISOLATION: Some AMD CPUs transiently execute instructions past an unconditional direct branch…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - transient execution beyond unconditional direct branches","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Some AMD CPUs transiently execute instructions past an unconditional direct branch - code that should never run, running speculatively and leaving traces in the cache. That gives an attacker speculative gadgets in places the compiler and the kernel's own Spectre auditing assume are unreachable, so hardened code can still leak. The practical outcome is data disclosure across privilege and guest boundaries.","attack_vector":"Local, from an unprivileged process or a guest VM.","remediation":"Mitigated by kernel-side changes that insert INT3 speculation barriers after unconditional branches in sensitive paths. Take the distro kernel update and reboot; no firmware step for the kernel mitigation itself. Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26341","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26921","cve":"CVE-2021-26921","aliases":[],"title":"Argo CD: Tokens keep working after the user account is disabled","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Tokens keep working after the user account is disabled; offboarding does not revoke access","attack_vector":"A former user with a cached token","remediation":"Rolling Argo CD upgrade; force-rotate all tokens","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26921"],"status":"curated"},{"id":"CVE-2021-31920","cve":"CVE-2021-31920","aliases":[],"title":"Istio: Multiple or escaped slashes bypass an Istio authorization policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Multiple or escaped slashes bypass an Istio authorization policy","attack_vector":"Unauthenticated network","remediation":"Rolling istiod upgrade plus sidecar restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-31920"],"status":"curated"},{"id":"CVE-2021-3524","cve":"CVE-2021-3524","aliases":[],"title":"Ceph RGW: HTTP header injection via a newline in the CORS ExposeHeader tag","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RGW","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"HTTP header injection via a newline in the CORS ExposeHeader tag","attack_vector":"Network (remote)","remediation":"Control-plane: RGW upgrade only; no OSD disruption","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3524"],"status":"curated"},{"id":"CVE-2021-3979","cve":"CVE-2021-3979","aliases":[],"title":"Ceph: Key length incorrectly passed to the encryption algorithm","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Key length incorrectly passed to the encryption algorithm -> non-random, weak key on encrypted disks","attack_vector":"Network (remote)","remediation":"Data-plane: storage node upgrade; re-encrypt affected RBD volumes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3979"],"status":"curated"},{"id":"CVE-2021-44776","cve":"CVE-2021-44776","aliases":[],"title":"Lanner IAC-AST2500A BMC firmware: The attacker rewrites who is permitted to use KVM and virtual media on the BMC - which is to say they grant…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lanner IAC-AST2500A BMC firmware","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The attacker rewrites who is permitted to use KVM and virtual media on the BMC - which is to say they grant themselves console access and the ability to attach a boot image. Even at a medium score this is a direct route to the two BMC capabilities that matter most to an operator: watching a tenant's console, and booting the node into attacker-supplied media. It is also a quieter attack than the overflow bugs, because it leaves the BMC running normally with an altered permission set rather than crashing anything. Broken access control in the SubNet_handler_func function of spx_restservice, allowing an attacker to change the security access rights governing KVM and virtual media.","attack_vector":"Network access to the BMC REST service with sufficient standing to invoke the subnet handler. Reachable from the out-of-band management network.","remediation":"Firmware flash, subject to the same Lanner sourcing problem as the rest of the cluster. Because this bug manipulates configuration rather than corrupting memory, it leaves auditable state: check the KVM and virtual-media permission settings on every IAC-AST2500A BMC against your intended baseline, and alert on changes. Where fixed firmware is unobtainable, disabling virtual media and KVM outright on the BMC - where the platform permits it - removes the capability the attacker is trying to grant themselves.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-44776","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-44776/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-46744","cve":"CVE-2021-46744","aliases":["CipherLeaks","ciphertext side channel"],"title":"AMD SEV / SEV-ES / SEV-SNP - ciphertext observability: MULTI-TENANT ISOLATION: SEV encrypts guest memory deterministically per physical address, so the same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV / SEV-ES / SEV-SNP - ciphertext observability","year":"2021","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: SEV encrypts guest memory deterministically per physical address, so the same plaintext at the same address always produces the same ciphertext. A malicious hypervisor that can read the encrypted pages watches ciphertext blocks change over time and infers the plaintext values underneath - the CipherLeaks researchers recovered full RSA and ECDSA private keys out of the SEV-protected VMSA this way. This is the deepest structural problem in the SEV family: it is a consequence of the encryption mode, not a coding bug, so it cannot simply be patched away.","attack_vector":"Requires a malicious or compromised hypervisor with the ability to read guest ciphertext - the exact adversary SEV exists to defeat. No guest bug needed.","remediation":"Partially mitigated: AMD added ciphertext-hiding for the VMSA in SEV-SNP firmware, delivered through AGESA/SEV firmware as an OEM SBIOS package with the usual **one to six month lag** and a drained-node power cycle, plus a TCB bump that requires refreshing VCEK certificates. On newer parts, SEV-SNP Ciphertext Hiding can be enabled to block host reads of guest ciphertext outright - check whether your platform and firmware support it and turn it on. The residual risk on older EPYC generations is **not fully fixable**: guest software must avoid keeping secrets in memory patterns an observer can correlate, which is a burden you cannot impose on a tenant's workload. Be honest with confidential-computing customers about which EPYC generation their VM lands on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46744","https://arxiv.org/abs/2204.11669","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-0001","cve":"CVE-2022-0001","aliases":["BHI","Spectre-BHB","Branch History Injection"],"title":"Intel processors (branch history injection): MULTI-TENANT ISOLATION: BHI / Spectre-BHB: even with eIBRS enabled, the branch history buffer is shared…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (branch history injection)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: BHI / Spectre-BHB: even with eIBRS enabled, the branch history buffer is shared across privilege levels, so unprivileged code can steer kernel-side speculation and read kernel memory. This is the attack that showed hardware Spectre-v2 mitigations were not the end of the story, and it is directly a container-to-host and guest-to-host read primitive.","attack_vector":"Local unprivileged code - any container or VM on the node.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. Linux additionally offers unprivileged-eBPF disabling and BHB-clearing sequences; check the spectre_v2 sysfs file after patching to see which mitigation actually engaged.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0001","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00598.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-0002","cve":"CVE-2022-0002","aliases":["Intra-mode BTI"],"title":"Intel processors (intra-mode branch target injection): The intra-mode sibling of BHI: branch predictor state is shared within a privilege level, so one sandboxed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (intra-mode branch target injection)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"The intra-mode sibling of BHI: branch predictor state is shared within a privilege level, so one sandboxed context can steer another's speculation without crossing rings. The concern is sandbox escape inside a single process - JIT tenants sharing a runtime.","attack_vector":"Local code inside the same privilege level as the victim, e.g. a sandboxed JIT tenant.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0002","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00598.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-23816","cve":"CVE-2022-23816","aliases":["RetBleed (AMD)","Branch Type Confusion"],"title":"AMD processors - branch predictor aliasing causing wrong branch type prediction (AMD-SB-1037): MULTI-TENANT ISOLATION: Aliases in the branch predictor cause some AMD processors to predict the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - branch predictor aliasing causing wrong branch type prediction (AMD-SB-1037)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Aliases in the branch predictor cause some AMD processors to predict the wrong branch type, which is AMD's form of the RetBleed / Branch Type Confusion problem. An attacker in one security domain trains the predictor so the victim's return or branch speculates to an attacker-chosen target, leaking data across process, kernel and guest boundaries. This is the AMD analogue people look for when they ask about Spectre-BHB, and it is the real answer.","attack_vector":"Local, cross-privilege and cross-guest. Reachable from any tenant workload on affected silicon.","remediation":"Needs both halves: AGESA/microcode from the OEM SBIOS package (**one to six months of lag**, drained node, power cycle) **and** an OS update that issues IBPB on context switch. Zen 1 and Zen 2 additionally depend on the LFENCE/JMP construct, which was itself found insufficient - so check that your kernel is using retpoline or IBRS rather than LFENCE/JMP by reading /sys/devices/system/cpu/vulnerabilities/spectre_v2 on the fleet. Patch alongside the other AMD-SB-1037 CVEs; they are one disclosure.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23816","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-1037.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-28199","cve":"CVE-2022-28199","aliases":[],"title":"NVIDIA MLNX_DPDK: Improper error recovery in NVIDIA's DPDK distribution lets a remote attacker cause denial of service with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA MLNX_DPDK","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper error recovery in NVIDIA's DPDK distribution lets a remote attacker cause denial of service with some integrity and confidentiality impact. Relevant on DPDK-based storage and network dataplanes fronting GPU clusters.","attack_vector":"Network - the attacker sends traffic that the DPDK dataplane mishandles. No authentication involved; this is packet-level reachability.","remediation":"Update MLNX_DPDK per bulletin 5389 and restart the dataplane application. Cost: a dataplane restart drops in-flight connections; on a storage path that means an I/O stall visible to running jobs, so drain or fail over first.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28199","https://github.com/NVIDIA/product-security/tree/main/2022/5389"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-1284"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-3162","cve":"CVE-2022-3162","aliases":[],"title":"Kubernetes (kube-apiserver): Users authorized to list/watch one namespaced CR type can read other CR types in the same API group…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Users authorized to list/watch one namespaced CR type can read other CR types in the same API group cluster-wide; cross-tenant data exposure","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain. Audit CRD RBAC","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3162"],"status":"curated"},{"id":"CVE-2022-3287","cve":"CVE-2022-3287","aliases":[],"title":"fwupd's Redfish plugin: Any unprivileged local user on the host can read a working BMC credential out of a config file. That is a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"fwupd's Redfish plugin","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Any unprivileged local user on the host can read a working BMC credential out of a config file. That is a direct escalation from 'has a shell on the node' to 'has an account on the node's out-of-band controller', which is the boundary a bare-metal operator is selling. From there the attacker reaches the BMC's Redfish surface with a legitimate operator account and can chain into any of the authenticated BMC bugs elsewhere in this list. If your provisioning tooling deploys fwupd with the Redfish plugin across the fleet, the same class of credential is sitting on every node. When it creates an OPERATOR account on the BMC, it writes the auto-generated password into /etc/fwupd/redfish.conf without restricting the file's permissions, leaving BMC credentials world-readable on the host.","attack_vector":"Any local unprivileged account on a host running fwupd with the Redfish plugin enabled. No root, no network position on the management VLAN, no exploit - just file read. On rented bare metal, the tenant is that local account.","remediation":"Update fwupd to a version carrying the permissions fix. Because the credential has already been written in the clear on existing installs, updating the package is not sufficient: you must also rotate the BMC operator account fwupd created on every affected node, and check the permissions of /etc/fwupd/redfish.conf directly rather than trusting the package version. This is config-and-credential work rather than a firmware flash, so it is cheap to remediate but easy to leave half-done.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3287","https://github.com/fwupd/fwupd/commit/ea676855f2119e36d433fbd2ed604039f53b2091"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-34665","cve":"CVE-2022-34665","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): A null-pointer dereference in the kernel mode layer with a changed CVSS scope crashes the node from an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the kernel mode layer with a changed CVSS scope crashes the node from an unprivileged account. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5383. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34665","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-34666","cve":"CVE-2022-34666","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): A second null-pointer dereference path in the kernel mode layer, same unprivileged local reach. Both the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A second null-pointer dereference path in the kernel mode layer, same unprivileged local reach. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5383. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34666","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-34678","cve":"CVE-2022-34678","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An unprivileged user null-pointer-dereferences the kernel mode layer, scope-changed, and panics the node.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unprivileged user null-pointer-dereferences the kernel mode layer, scope-changed, and panics the node. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5415. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34678","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-36055","cve":"CVE-2022-36055","aliases":[],"title":"Helm: OOM panic in the strvals package from crafted `--set` input","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"OOM panic in the strvals package from crafted `--set` input","attack_vector":"Anyone who can supply chart values, e.g. through a self-service portal","remediation":"Upgrade Helm; validate tenant-supplied values","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36055"],"status":"curated"},{"id":"CVE-2022-40716","cve":"CVE-2022-40716","aliases":[],"title":"HashiCorp Consul: Internal RPC endpoint does not check multiple SAN URIs in a CSR","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Consul","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Internal RPC endpoint does not check multiple SAN URIs in a CSR -> bypass service-mesh intentions","attack_vector":"Network (remote)","remediation":"Control-plane: Consul server upgrade then clients; re-verify mesh intentions","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40716"],"status":"curated"},{"id":"CVE-2022-40982","cve":"CVE-2022-40982","aliases":[],"title":"Intel CPU (Downfall / GDS): Downfall: Gather Data Sampling leaks AVX gather-instruction data across SMT siblings, containers and VMs…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Downfall / GDS)","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Downfall: Gather Data Sampling leaks AVX gather-instruction data across SMT siblings, containers and VMs - directly breaks multi-tenant isolation","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode update + reboot, standing perf cost (reported up to ~50% on gather-heavy vector code). Alternative is disabling AVX gather, which is worse for AI workloads. On shared-GPU nodes with co-tenanted CPUs this is a must-fix","references":["https://access.redhat.com/security/cve/CVE-2022-40982"],"status":"curated","fleet":{"ubiquity":"very common - Skylake through Tiger Lake era Xeons, still the host CPU under a large installed base of GPU nodes","remediation_pain":"microcode+reboot - microcode is loaded at boot, so every node drains and reboots; the mitigation carries a measurable AVX2/AVX-512 gather slowdown, i.e. a permanent throughput tax on the fleet","pain_class":"microcode + reboot","why_fleet_wide":"Cross-tenant data leakage from stale vector registers on shared hardware - exactly the isolation property a multi-tenant GPU cloud sells - so it forces a fleet-wide reboot campaign regardless of workload."}},{"id":"CVE-2022-42282","cve":"CVE-2022-42282","aliases":[],"title":"DGX-2 BMC: Info disclosure via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Info disclosure via path traversal","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42282","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-0204","cve":"CVE-2023-0204","aliases":[],"title":"ConnectX-5/6/6-DX NIC firmware: NIC DoS (improper exception handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX-5/6/6-DX NIC firmware","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"NIC DoS (improper exception handling)","attack_vector":"Any unprivileged tenant with the NIC exposed (SR-IOV VF)","remediation":"Flash NIC firmware to 35.1012+; requires node reboot, evict tenants","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5459/5459.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-703"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-20575","cve":"CVE-2023-20575","aliases":[],"title":"AMD processors - power reporting side channel against SEV VMs: MULTI-TENANT ISOLATION: An authenticated attacker uses the platform's power reporting functionality to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - power reporting side channel against SEV VMs","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An authenticated attacker uses the platform's power reporting functionality to monitor execution inside an AMD SEV VM. The whole promise of SEV is that the host cannot see what the confidential guest is doing; power telemetry is a channel the memory encryption does not cover, so a host operator watches the guest's execution profile through the power meter. For anyone selling confidential computing on EPYC this is a direct hole in the product claim.","attack_vector":"Local, authenticated, with access to power reporting interfaces on a host running SEV guests.","remediation":"Mitigated by restricting access to power reporting interfaces and by AGESA-level changes to reduce telemetry resolution. Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. The immediately actionable step is to lock down who can read host power telemetry on confidential-computing nodes - that is a permissions change, not a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20575","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20591","cve":"CVE-2023-20591","aliases":[],"title":"AMD IOMMU - not re-initialized during DRTM (AMD-SB-3003): MULTI-TENANT ISOLATION: The IOMMU is not re-initialized during a Dynamic Root of Trust for Measurement…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU - not re-initialized during DRTM (AMD-SB-3003)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The IOMMU is not re-initialized during a Dynamic Root of Trust for Measurement launch, so stale DMA mappings survive into the measured environment and an attacker can read or modify **hypervisor** memory via a device. DRTM exists precisely to establish a clean, measured starting state; leaving the IOMMU stale means the thing you are measuring can already be under device-level attack. On a GPU host, where accelerators and RDMA NICs are DMA-capable and numerous, that is a large set of devices to leave pointed at hypervisor memory.","attack_vector":"Local, requires the ability to set up DMA mappings before a DRTM launch and to drive a device afterwards.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20591","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-20593","cve":"CVE-2023-20593","aliases":[],"title":"AMD CPU (Zenbleed): Zenbleed: cross-process/cross-VM register-file data leak on Zen 2 at ~30 kB/s per core, no special privileges","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Zenbleed)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Zenbleed: cross-process/cross-VM register-file data leak on Zen 2 at ~30 kB/s per core, no special privileges","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"AMD microcode (AGESA) update + reboot; kernel chicken-bit workaround (DE_CFG[9]) available with a measured perf cost. Zen 2 EPYC is still common as the CPU side of A100/L40S nodes","references":["https://access.redhat.com/security/cve/CVE-2023-20593"],"status":"curated","fleet":{"ubiquity":"common - Zen 2 EPYC (Rome) still hosts a large installed base of GPU nodes and rental fleets","remediation_pain":"microcode+reboot for the real fix; the interim DE_CFG MSR chicken-bit workaround is a kernel change with a measurable FP/vector performance cost - so the fleet either reboots for microcode or eats a permanent tax","pain_class":"microcode + reboot","why_fleet_wide":"Leaks ~30 KB/s/core of stale vector-register data across any privilege boundary including cross-process and cross-VM, i.e. exactly the co-tenancy isolation a GPU cloud sells, on every Rome host at once."}},{"id":"CVE-2023-22276","cve":"CVE-2023-22276","aliases":[],"title":"Intel Ethernet Controller E810 Series firmware: A race condition in E810 firmware lets an authenticated local user cause denial of service. In a bare-metal…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 Series firmware","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A race condition in E810 firmware lets an authenticated local user cause denial of service. In a bare-metal multi-tenant cluster 'authenticated local' is the tenant, so this is a tenant able to kill its own node's NIC — and on shared-NIC designs, potentially the NIC serving other functions on the same host.","attack_vector":"Authenticated local access to the host. In a bare-metal GPU rental model, that is the customer.","remediation":"Flash E810 firmware to 1.7.2.4 or later; cold power cycle. This one is a good argument for reflashing NIC firmware as part of tenant handoff rather than only on a CVE-driven schedule.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22276"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-22497","cve":"CVE-2023-22497","aliases":[],"title":"Netdata: Agent MACHINE GUID is readable and reusable","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Netdata","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Agent MACHINE GUID is readable and reusable -> impersonate an agent to the registry/parent","attack_vector":"Network (remote)","remediation":"Data-plane: Netdata agents run on GPU nodes - fleet-wide agent upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22497"],"status":"curated"},{"id":"CVE-2023-25526","cve":"CVE-2023-25526","aliases":[],"title":"Cumulus Linux (neighmgrd/nlmanager): Switch DoS via crafted packet","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux (neighmgrd/nlmanager)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Switch DoS via crafted packet","attack_vector":"Network-adjacent attacker (any tenant on the L2 domain)","remediation":"Upgrade Cumulus Linux to 5.5.0+; rolling switch upgrade","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5480/5480.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-248"]},{"id":"CVE-2023-25532","cve":"CVE-2023-25532","aliases":[],"title":"DGX H100 BMC (IPMI): Info disclosure of credentials","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Info disclosure of credentials","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18; rotate credentials","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-522"]},{"id":"CVE-2023-26054","cve":"CVE-2023-26054","aliases":[],"title":"BuildKit: Git URL credentials in a build request are persisted into the build cache and can be read by other builds","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Git URL credentials in a build request are persisted into the build cache and can be read by other builds","attack_vector":"Anyone sharing a multi-tenant builder","remediation":"Upgrade BuildKit; purge shared cache; rotate git credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-26054"],"status":"curated"},{"id":"CVE-2023-2727","cve":"CVE-2023-2727","aliases":[],"title":"Kubernetes (kube-apiserver): Ephemeral containers bypass the ImagePolicyWebhook, so unapproved images run","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Ephemeral containers bypass the ImagePolicyWebhook, so unapproved images run","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2727"],"status":"curated"},{"id":"CVE-2023-2728","cve":"CVE-2023-2728","aliases":[],"title":"Kubernetes (kube-apiserver): Ephemeral containers bypass the ServiceAccount mountable-secrets policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Ephemeral containers bypass the ServiceAccount mountable-secrets policy; a tenant mounts secrets they were denied","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2728"],"status":"curated"},{"id":"CVE-2023-27595","cve":"CVE-2023-27595","aliases":[],"title":"Cilium: On agent start, eBPF programs are briefly detached, so traffic bypasses NetworkPolicy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"On agent start, eBPF programs are briefly detached, so traffic bypasses NetworkPolicy","attack_vector":"Any pod on the cluster network during an agent restart","remediation":"Upgrade Cilium; note that every agent restart, including your own upgrades, is a policy-gap window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27595"],"status":"curated"},{"id":"CVE-2023-28376","cve":"CVE-2023-28376","aliases":["INTEL-SA-00869"],"title":"Intel E810 Ethernet Controller firmware: Out-of-bounds read in E810 firmware reachable from an adjacent, unauthenticated attacker, causing denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet Controller firmware","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in E810 firmware reachable from an adjacent, unauthenticated attacker, causing denial of service. Adjacent means the same L2 segment — in a leaf/spine cluster that is every other node in the rack, including other tenants' nodes if you share a VLAN. One compromised tenant machine can walk the rack knocking NICs offline.","attack_vector":"Unauthenticated, adjacent — a host on the same layer-2 segment as the target NIC.","remediation":"Flash E810 firmware to 1.7.1 or later. Cold power cycle required for the NVM image to activate; plan a per-node drain. If you cannot patch immediately, keeping tenants on separate VLANs limits who is 'adjacent' — a switch config change that meaningfully shrinks the exposed set.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00869.html","https://nvd.nist.gov/vuln/detail/CVE-2023-28376"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-28746","cve":"CVE-2023-28746","aliases":["RFDS","Register File Data Sampling"],"title":"Intel processors (register file data sampling): MULTI-TENANT ISOLATION: RFDS: stale data left in the integer, floating-point and vector register files after…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (register file data sampling)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: RFDS: stale data left in the integer, floating-point and vector register files after transient execution can be sampled by a local attacker, crossing the process, VM and enclave boundaries. Vector register files are where model activations and weights live during compute, so on an AI host this leaks the workload's actual data, not just pointers.","attack_vector":"Local authenticated code on an affected processor, including a co-tenant VM or container.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The kernel-side mitigation reuses the VERW buffer-clearing path, so a current kernel plus current microcode is the whole story; check the reg_file_data_sampling sysfs entry after reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28746","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00898.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2023-2878","cve":"CVE-2023-2878","aliases":[],"title":"secrets-store-csi-driver: Service account tokens written to driver logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"secrets-store-csi-driver","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Service account tokens written to driver logs","attack_vector":"Anyone with log-pipeline read access","remediation":"Upgrade the driver via DaemonSet rollout; rotate the exposed service-account tokens; scrub logs","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2023-30575","cve":"CVE-2023-30575","aliases":[],"title":"Apache Guacamole: Miscalculated instruction lengths during the Guacamole handshake","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Guacamole","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Miscalculated instruction lengths during the Guacamole handshake -> protocol instruction injection","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade guacd and the webapp; console gateways are high-value","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-30575"],"status":"curated"},{"id":"CVE-2023-31018","cve":"CVE-2023-31018","aliases":[],"title":"GPU Display Driver (Linux + Windows KMD): Host DoS (null deref from unprivileged user)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Linux + Windows KMD)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Host DoS (null deref from unprivileged user)","attack_vector":"Any tenant with a container holding /dev/nvidia*","remediation":"Driver upgrade; rolling node reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-31025","cve":"CVE-2023-31025","aliases":[],"title":"DGX A100 BMC: LDAP injection in BMC auth","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 BMC","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"LDAP injection in BMC auth","attack_vector":"Network-adjacent attacker","remediation":"Flash BMC 00.22.05+","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31025","https://github.com/NVIDIA/product-security/tree/main/2024/5510"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-90"]},{"id":"CVE-2023-31419","cve":"CVE-2023-31419","aliases":[],"title":"Elasticsearch: Crafted _search query string","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Elasticsearch","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Crafted _search query string -> stack overflow and node denial of service","attack_vector":"Network (remote)","remediation":"Control-plane: rolling upgrade of the log cluster","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31419"],"status":"curated"},{"id":"CVE-2023-34345","cve":"CVE-2023-34345","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (SPX REST API): Path traversal in the BMC REST API letting a low-privilege user read arbitrary files off the BMC filesystem.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (SPX REST API)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Path traversal in the BMC REST API letting a low-privilege user read arbitrary files off the BMC filesystem. That is where the interesting material lives: the shadow file, TLS private keys, SSH host keys, IPMI user database and stored configuration. It is a credential-harvesting step - the attacker takes what they read here and uses it to escalate on this BMC and, because BMC credentials are usually cloned across a fleet, on every sibling node.","attack_vector":"Network-reachable REST API with only a low-privilege BMC account. That is a much lower bar than the admin-required bugs in the same advisory - a read-only telemetry or monitoring account is sufficient, and those are exactly the accounts operators hand out widely and rarely rotate.","remediation":"Firmware flash to SPx_12.5 / SPx_13.3 or later, out-of-band per node, ODM-gated. Because the exploit path starts from a low-privilege account, the immediate config-only step is to inventory BMC accounts and delete or rotate every stale, shared or monitoring credential - and to assume any credential material stored on an unpatched BMC has already been read, and rotate it rather than leaving it in place.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34345"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40285","cve":"CVE-2023-40285","aliases":[],"title":"Supermicro BMC (IPMI web interface, XSS): Another injection point in the BMC web interface, lower-impact than its siblings but usable for the same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface, XSS)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Another injection point in the BMC web interface, lower-impact than its siblings but usable for the same session-hijack chain toward virtual media and firmware flash.","attack_vector":"Network reach to the BMC web UI plus an operator loading the affected page.","remediation":"BMC firmware flash per board; same batch as the rest of the 2023 Supermicro BMC advisories, so fix them together rather than one at a time.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40285"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-40584","cve":"CVE-2023-40584","aliases":[],"title":"Argo CD: repo-server extracts a user-controlled tar.gz without size validation","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Argo CD","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"repo-server extracts a user-controlled tar.gz without size validation -> denial of service","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; isolate and resource-cap the repo-server","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40584"],"status":"curated"},{"id":"CVE-2023-4091","cve":"CVE-2023-4091","aliases":[],"title":"Samba: SMB client can truncate files despite read-only permissions when acl_xattr ignores system ACLs","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Samba","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"SMB client can truncate files despite read-only permissions when acl_xattr ignores system ACLs","attack_vector":"Network (remote)","remediation":"Data-plane: smbd upgrade + review acl_xattr share config","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4091"],"status":"curated"},{"id":"CVE-2023-43040","cve":"CVE-2023-43040","aliases":[],"title":"Ceph RGW (IBM Spectrum Fusion HCI): Improper bucket access lets an actor perform unauthorized actions in RGW","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Ceph RGW (IBM Spectrum Fusion HCI)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper bucket access lets an actor perform unauthorized actions in RGW","attack_vector":"Network (remote)","remediation":"Control-plane: RGW/appliance firmware upgrade; re-audit bucket policies","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-43040"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45229","cve":"CVE-2023-45229","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (DHCPv6 Advertise, IA_NA/IA_TA option parsing): An integer underflow when parsing the identity-association options in a DHCPv6 Advertise causes the firmware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (DHCPv6 Advertise, IA_NA/IA_TA option parsing)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An integer underflow when parsing the identity-association options in a DHCPv6 Advertise causes the firmware to read outside its buffer. On its own this leaks firmware memory contents or crashes the boot; chained with the overflow bugs in the same advertisement path it is the information-leak half of a reliable pre-OS exploit (defeating whatever address-layout guesswork the attacker would otherwise need).","attack_vector":"Anyone able to send DHCPv6 Advertise messages on the segment the node PXE-boots from. Unauthenticated, pre-OS.","remediation":"Firmware flash via the server OEM's BIOS package - the fix is in upstream edk2 but only reaches you after the IBV rebase and the OEM's own validation cycle. Reboot per node. Config-only stopgap: disable IPv6 network boot, or PXE entirely, and treat the provisioning VLAN as a trust boundary that tenant workloads must never reach.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45229","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-45231","cve":"CVE-2023-45231","aliases":["PixieFail","VU#132380"],"title":"EDK II NetworkPkg (IPv6 Neighbor Discovery Redirect handling): A truncated ND Redirect message drives an out-of-bounds read in the firmware IPv6 stack. Practical outcome is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IPv6 Neighbor Discovery Redirect handling)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A truncated ND Redirect message drives an out-of-bounds read in the firmware IPv6 stack. Practical outcome is firmware memory disclosure or a wedged boot; the more interesting operational consequence is that the Redirect path itself lets an on-link attacker steer where the booting node sends its traffic, so this is both a leak and a foothold for redirecting the netboot fetch.","attack_vector":"On-link IPv6 attacker on the provisioning segment - any host that can emit ICMPv6 Neighbor Discovery to the booting node. Unauthenticated, pre-OS.","remediation":"OEM BIOS update, flash + reboot per node; the IBV-to-OEM rebase lag applies. Interim: enable IPv6 RA Guard / ND inspection on the provisioning switches, and disable the UEFI IPv6 network stack on nodes that boot locally. No OS-level or config-in-firmware toggle short of turning network boot off.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-45231","https://kb.cert.org/vuls/id/132380","https://github.com/tianocore/edk2/security/advisories/GHSA-hc6x-cw6p-gj7h"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-4969","cve":"CVE-2023-4969","aliases":["LeftoverLocals","VU#446598"],"title":"GPU local/shared memory not cleared between kernels (AMD, Apple, Qualcomm, Imagination): MULTI-TENANT ISOLATION: a GPU kernel reads whatever the previous kernel left in local (shared/scratchpad)…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"GPU local/shared memory not cleared between kernels (AMD, Apple, Qualcomm, Imagination)","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: a GPU kernel reads whatever the previous kernel left in local (shared/scratchpad) memory, including a kernel belonging to a different user, process or container. Trail of Bits recovered another process's LLM inference output token by token - on an AMD Radeon RX 7900 XT the leak was around 5.5 MB per GPU invocation, roughly 181 MB per query against a 7B model on llama.cpp. That is enough to reconstruct prompts, activations and responses, not just fragments. Affected vendors are AMD, Apple, Qualcomm and Imagination; NVIDIA, Intel and Arm tested clean and told CERT/CC they were not impacted. If you run AMD Instinct or ROCm, this is your bug and AMD was still investigating mitigations at disclosure.","attack_vector":"A co-tenant. The attacker needs only to run an ordinary OpenCL/Vulkan/Metal compute kernel on the same physical GPU - no privileges, no kernel exploit, no driver bug in the usual sense. Time-sliced sharing, MPS-style sharing and sequential job scheduling on the same device all qualify.","remediation":"Qualcomm shipped firmware v2.07 (January 2024) and Imagination fixed it in DDK 23.3 (December 2023); Apple fixed it in silicon from A17/M3 onward, leaving older Apple GPUs UNPATCHABLE; ChromeOS shipped AMD and Qualcomm mitigations in stable 120 / LTS 114. AMD's position at disclosure was that devices remained vulnerable pending mitigation work - track AMD-SB-6010 for your specific Instinct parts rather than assuming a fix exists. Cost where a driver fix does exist: driver upgrade plus node drain to reload the kernel module. Where it does not, the only controls are refusing to share a GPU across trust boundaries, or having your runtime explicitly zero local memory at kernel entry - which costs measurable throughput on small kernels and has to be done by whoever compiles the kernels, not by you.","references":["https://kb.cert.org/vuls/id/446598","https://blog.trailofbits.com/2024/01/16/leftoverlocals-listening-to-llm-responses-through-leaked-gpu-local-memory/","https://nvd.nist.gov/vuln/detail/CVE-2023-4969","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-6010"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-6570","cve":"CVE-2023-6570","aliases":[],"title":"Kubeflow: SSRF","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"SSRF","attack_vector":"Authenticated notebook/pipeline user","remediation":"Upgrade; block IMDS egress from Kubeflow pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6570"],"status":"curated"},{"id":"CVE-2024-0078","cve":"CVE-2024-0078","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0078","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0079","cve":"CVE-2024-0079","aliases":[],"title":"vGPU Manager: Guest-triggered host DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Guest-triggered host DoS","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0079","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"]},{"id":"CVE-2024-0083","cve":"CVE-2024-0083","aliases":[],"title":"ChatRTX: XSS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"XSS","attack_vector":"Web user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0083","https://github.com/NVIDIA/product-security/tree/main/2024/5532"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L","cwe":["CWE-79"]},{"id":"CVE-2024-0093","cve":"CVE-2024-0093","aliases":[],"title":"vGPU Manager: Cross-tenant info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Cross-tenant info disclosure","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0093","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:H/I:N/A:N","cwe":["CWE-200"]},{"id":"CVE-2024-0100","cve":"CVE-2024-0100","aliases":[],"title":"Triton Inference Server: Info disclosure via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Info disclosure via path traversal","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0100","https://github.com/NVIDIA/product-security/tree/main/2024/5535"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-73"]},{"id":"CVE-2024-10270","cve":"CVE-2024-10270","aliases":[],"title":"Keycloak: Regex complexity in SearchQueryUtils","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Regex complexity in SearchQueryUtils -> resource exhaustion DoS of the auth service","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; rate-limit the token/search endpoints","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10270"],"status":"curated"},{"id":"CVE-2024-11185","cve":"CVE-2024-11185","aliases":["Arista Security Advisory 0118"],"title":"Arista EOS (L2 forwarding / VLAN isolation): TENANT ISOLATION: ingress traffic on a layer-2 port is forwarded out ports belonging to a different VLAN.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (L2 forwarding / VLAN isolation)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: ingress traffic on a layer-2 port is forwarded out ports belonging to a different VLAN. VLAN separation is the primary tenant boundary in most GPU-cluster builds, so this is one tenant's frames landing in another tenant's broadcast domain. CVSS 6.5 understates the operator consequence — for a neocloud selling isolated tenancy this is a contractual failure, not a medium-severity bug.","attack_vector":"An attacker on any L2 port under the conditions the advisory describes. No credentials — this is a forwarding-plane defect, not an access-control one.","remediation":"EOS upgrade plus switch reload on every affected leaf. No config workaround — you cannot ACL your way out of a forwarding-plane leak. Plan a rolling upgrade across the leaf layer; on MLAG pairs you can do one side at a time and keep the rack up.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-11185"],"status":"curated"},{"id":"CVE-2024-1725","cve":"CVE-2024-1725","aliases":[],"title":"KubeVirt: kubevirt-csi in OpenShift Virtualization HCP grants access to the root HCP worker node's volume via a crafted…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"kubevirt-csi in OpenShift Virtualization HCP grants access to the root HCP worker node's volume via a crafted PV","attack_vector":"Authenticated cluster user","remediation":"Upgrade the kubevirt-csi component","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-1725"],"status":"curated"},{"id":"CVE-2024-20365","cve":"CVE-2024-20365","aliases":[],"title":"Redfish API implementation on Cisco UCS B-Series, UCS Managed C-Series and UCS X-Series servers: An administrator-level Redfish user escapes the API's intended command surface and executes…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Redfish API implementation on Cisco UCS B-Series, UCS Managed C-Series and UCS X-Series servers","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An administrator-level Redfish user escapes the API's intended command surface and executes commands on the underlying management controller. The operator-relevant point is the boundary being crossed: Redfish administrative privilege is supposed to mean 'can configure the server', not 'can run code on the controller', and that distinction is what lets an operator delegate Redfish admin to automation without handing over the box. When it collapses, every service account your provisioning system holds becomes a shell on every BMC it manages. Command injection reachable through the Redfish interface by a user who already holds administrative privileges in the management software.","attack_vector":"An authenticated remote user with administrative privileges on the UCS management software, reaching the Redfish API. In practice this is your automation's own credentials - Terraform providers, Ansible modules and inventory collectors all hold Redfish admin, and any compromise of the automation host inherits it.","remediation":"Firmware/software update to the fixed UCS release per Cisco's advisory; this is a controller firmware update, so plan a per-chassis maintenance window rather than a rolling config push. The durable lesson is architectural: stop treating Redfish administrative accounts as a safe delegation boundary. Scope automation credentials to the narrowest Redfish role that works, keep them out of shared secret stores that tenant-facing systems can read, and log Redfish administrative calls centrally so an anomalous command sequence is visible.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20365","https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-cimc-redfish-cominj-sbkv5ZZ"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-21121","cve":"CVE-2024-21121","aliases":[],"title":"Oracle VirtualBox: Easily exploitable Core flaw allowing unauthorised access to VirtualBox-accessible data","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Oracle VirtualBox","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Easily exploitable Core flaw allowing unauthorised access to VirtualBox-accessible data","attack_vector":"Tenant VM guest","remediation":"VirtualBox update + VM restart","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21121"],"status":"curated"},{"id":"CVE-2024-2206","cve":"CVE-2024-2206","aliases":[],"title":"Gradio: SSRF in the `/proxy` route","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Gradio","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"SSRF in the `/proxy` route","attack_vector":"Unauthenticated network","remediation":"Upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-2206"],"status":"curated"},{"id":"CVE-2024-23499","cve":"CVE-2024-23499","aliases":[],"title":"Intel ice driver (Ethernet 800 Series, Linux kernel mode): A protection-mechanism failure in the E810 Linux kernel driver reachable by an unauthenticated attacker.…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel ice driver (Ethernet 800 Series, Linux kernel mode)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A protection-mechanism failure in the E810 Linux kernel driver reachable by an unauthenticated attacker. Driver-side counterpart to the E810 firmware protection failures in the same advisory.","attack_vector":"Unauthenticated attacker able to present traffic to the interface.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes. Target ice 28.3 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23499","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-24580","cve":"CVE-2024-24580","aliases":[],"title":"Intel Data Center GPU Max Series 1100 / 1550: A second improper conditions check in the Max Series allowing a privileged local user to cause denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center GPU Max Series 1100 / 1550","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A second improper conditions check in the Max Series allowing a privileged local user to cause denial of service on the accelerator.","attack_vector":"Local, privileged. Host root on the node.","remediation":"Apply the Intel update for the Max Series. Cost: driver reload, or drain and reboot if the fix is in firmware.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-24580"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-24983","cve":"CVE-2024-24983","aliases":["INTEL-SA-00918"],"title":"Intel Ethernet Controller E810 firmware: An unauthenticated attacker on the network can take an E810 NIC out of service through a protection-mechanism…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 firmware","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated attacker on the network can take an E810 NIC out of service through a protection-mechanism failure in the adapter firmware. E810 is the standard 100/200GbE front-end NIC on a large share of AI servers, and it is often also the storage-network NIC — so knocking it down removes the node from both the job and its dataset. Because the fault is in firmware, the host OS sees a dead link, not a driver problem, and normal remediation (restart the driver) does not recover it.","attack_vector":"Unauthenticated, remote — traffic arriving at the NIC over the network. No host credentials and no adjacency requirement.","remediation":"Flash E810 adapter firmware to 4.4 or later using Intel's NVM Update Tool (or the OEM-repackaged version — Dell DSA-2025-236, HPE and Lenovo ship their own). NVM updates require a **cold power cycle**, not a warm reboot, for the new image to take effect — so this is a full node drain per server, which across a GPU fleet is the dominant cost. Sequence it with your normal node-maintenance rotation rather than as an emergency.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00918.html","https://nvd.nist.gov/vuln/detail/CVE-2024-24983"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-25742","cve":"CVE-2024-25742","aliases":["#VC injection","WeSee-class"],"title":"SEV-ES / SEV-SNP guest kernel - unsolicited #VC (vector 29) injection: MULTI-TENANT ISOLATION: An untrusted hypervisor can inject the #VC exception (vector 29) into an SEV-ES or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SEV-ES / SEV-SNP guest kernel - unsolicited #VC (vector 29) injection","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An untrusted hypervisor can inject the #VC exception (vector 29) into an SEV-ES or SEV-SNP guest at a moment of its choosing and force the guest's #VC handler to run in an unexpected context. The handler is privileged guest code that touches the GHCB shared page, so a host that controls when it fires gains a lever on guest execution - the WeSee research turned this class into reading and writing confidential guest memory and executing code inside the VM. The guest's own hardware isolation is what gets used against it.","attack_vector":"Malicious or compromised hypervisor against its own guest. No guest bug required and no tenant cooperation - the host simply injects.","remediation":"Fixed in the **guest** kernel, not the host - the hardening lives in the SEV-ES/SNP guest's #VC handler and interrupt entry code. That inverts the usual rollout: you can patch every hypervisor you own and still be exposed, because the protection has to be in the tenant's own VM image. As an operator your job is to ship updated confidential-guest images (or tell tenants which minimum kernel to run) and, where you can, enforce it as an admission requirement. Each guest picks the fix up on its next boot; no host reboot, no firmware update. Fixed in Linux 6.9 and backported; guests must run a kernel that hardens the #VC entry path. As the operator you cannot fix this for a tenant who brings their own image - the honest control is to document a minimum guest kernel and enforce it at admission for confidential workloads.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25742","https://ahoi-attacks.github.io/wesee/","https://www.amd.com/en/resources/product-security.html"],"status":"curated"},{"id":"CVE-2024-26808","cve":"CVE-2024-26808","aliases":[],"title":"Linux kernel (netfilter): nft_chain_filter NETDEV_UNREGISTER mishandling for inet/ingress basechains - UAF","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (netfilter)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"nft_chain_filter NETDEV_UNREGISTER mishandling for inet/ingress basechains - UAF","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-26808"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-29893","cve":"CVE-2024-29893","aliases":[],"title":"Argo CD: Repo-server DoS, halting all GitOps reconciliation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Repo-server DoS, halting all GitOps reconciliation","attack_vector":"Any authenticated Argo CD user","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29893"],"status":"curated"},{"id":"CVE-2024-31141","cve":"CVE-2024-31141","aliases":[],"title":"Apache Kafka (client): ConfigProvider plugins let an untrusted app read files/env of the Kafka client host","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Apache Kafka (client)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"ConfigProvider plugins let an untrusted app read files/env of the Kafka client host","attack_vector":"Network (remote)","remediation":"Control-plane: dependency upgrade in telemetry/billing pipelines","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31141"],"status":"curated"},{"id":"CVE-2024-32048","cve":"CVE-2024-32048","aliases":[],"title":"Intel Distribution of OpenVINO Model Server: An unauthenticated user can reach an input-validation flaw in OpenVINO Model Server. Model Server is a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Distribution of OpenVINO Model Server","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated user can reach an input-validation flaw in OpenVINO Model Server. Model Server is a production inference endpoint, so this is on the request path of a live service.","attack_vector":"Anything that can send a request to the model server - which for most deployments is the whole cluster network, and sometimes the internet.","remediation":"Upgrade OpenVINO Model Server to 2024.0 or later. Container image swap and rolling restart; no node reboot or firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32048","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01158.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-36293","cve":"CVE-2024-36293","aliases":[],"title":"Intel processors with SGX (EDECCSSA leaf): Improper access control on the EDECCSSA user leaf function lets an authenticated local user deny SGX service.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors with SGX (EDECCSSA leaf)","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper access control on the EDECCSSA user leaf function lets an authenticated local user deny SGX service. EDECCSSA is the SGX2 instruction used for in-enclave exception handling, so this is reachable from ordinary enclave-adjacent code.","attack_vector":"Local authenticated user on an SGX-enabled host.","remediation":"Microcode update and reboot; late-loadable at boot without an OEM BIOS release. Re-attest afterwards.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36293","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01213.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2024-36620","cve":"CVE-2024-36620","aliases":[],"title":"Docker / moby: NULL pointer dereference in image_history crashes the daemon","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"NULL pointer dereference in image_history crashes the daemon","attack_vector":"Any tenant workload / API client","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36620"],"status":"curated"},{"id":"CVE-2024-36621","cve":"CVE-2024-36621","aliases":[],"title":"Docker / moby: Race in the buildkit snapshot adapter","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Race in the buildkit snapshot adapter; concurrent builds leak resources and exhaust the host","attack_vector":"Anyone who can submit builds to a shared builder","remediation":"Upgrade moby; isolate tenant builders","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36621"],"status":"curated"},{"id":"CVE-2024-3744","cve":"CVE-2024-3744","aliases":[],"title":"azure-file-csi-driver: Service account tokens disclosed in driver logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"azure-file-csi-driver","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Service account tokens disclosed in driver logs","attack_vector":"Anyone with log-pipeline read access","remediation":"DaemonSet rollout; rotate tokens; scrub logs","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2024-45806","cve":"CVE-2024-45806","aliases":[],"title":"Envoy: External clients manipulate Envoy internal headers, reaching unauthorized behaviour","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Envoy","year":"2024","cvss_score":6.5,"severity":"medium","kev":false,"impact":"External clients manipulate Envoy internal headers, reaching unauthorized behaviour","attack_vector":"Unauthenticated network","remediation":"Upgrade Envoy","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45806"],"status":"curated"},{"id":"CVE-2025-14269","cve":"CVE-2025-14269","aliases":[],"title":"Headlamp: Credential caching in Headlamp when Helm is enabled","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Headlamp","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Credential caching in Headlamp when Helm is enabled","attack_vector":"Cluster user with dashboard access","remediation":"Upgrade Headlamp; rotate cached credentials","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-1767","cve":"CVE-2025-1767","aliases":[],"title":"Kubernetes (kubelet): gitRepo volume grants inadvertent access to local repositories on the node","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"gitRepo volume grants inadvertent access to local repositories on the node","attack_vector":"Cluster user able to create a gitRepo volume","remediation":"Rolling kubelet upgrade; block gitRepo volumes","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-1944","cve":"CVE-2025-1944","aliases":[],"title":"picklescan: ZIP manipulation crashes the scanner (scan bypass by DoS)","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"picklescan","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"ZIP manipulation crashes the scanner (scan bypass by DoS)","attack_vector":"Customer-supplied model archive","remediation":"Upgrade to 0.0.23+; fail-closed on scanner crash","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1944"],"status":"curated"},{"id":"CVE-2025-22892","cve":"CVE-2025-22892","aliases":[],"title":"OpenVINO Model Server: An unauthenticated request can drive OpenVINO Model Server into unbounded resource consumption and take the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenVINO Model Server","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated request can drive OpenVINO Model Server into unbounded resource consumption and take the endpoint down. Cheap remote DoS against a serving tier.","attack_vector":"Anyone who can send a request to the model server.","remediation":"Upgrade OpenVINO Model Server to 2024.4 or later, and put request-size and rate limits in front of the endpoint. Rolling container restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22892","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01272.html"],"status":"curated"},{"id":"CVE-2025-23047","cve":"CVE-2025-23047","aliases":[],"title":"Cilium: Insecure default Access-Control-Allow-Origin in Hubble UI exposes sensitive observability data","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Insecure default Access-Control-Allow-Origin in Hubble UI exposes sensitive observability data","attack_vector":"Anyone who can get an operator's browser to a hostile page","remediation":"Upgrade Cilium; put Hubble UI behind auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23047"],"status":"curated"},{"id":"CVE-2025-23243","cve":"CVE-2025-23243","aliases":[],"title":"NVIDIA Riva: Weak authentication","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Riva","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Weak authentication -> unauthorized use / info disclosure","attack_vector":"Network client","remediation":"Upgrade Riva containers; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23243","https://github.com/NVIDIA/product-security/tree/main/2025/5625"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-284"]},{"id":"CVE-2025-23259","cve":"CVE-2025-23259","aliases":[],"title":"Mellanox DPDK: DoS / data tampering via race condition","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox DPDK","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS / data tampering via race condition","attack_vector":"Tenant with a DPDK-attached VF","remediation":"Bump DPDK packages; rebuild dataplane images; restart dataplane","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23259","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-362"]},{"id":"CVE-2025-24323","cve":"CVE-2025-24323","aliases":["INTEL-SA-01339"],"title":"Intel PCIe Switch firmware package and LED mode toggle tool before version MR4_1.0b1: Improper access control in the PCIe switch firmware package and its management tool lets a privileged local…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel PCIe Switch firmware package and LED mode toggle tool before version MR4_1.0b1","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Improper access control in the PCIe switch firmware package and its management tool lets a privileged local user escalate. The PCIe switch is the fabric between the CPU complex and the GPUs, NVMe and NICs inside a node - it decides which device can DMA where. Firmware-level control of the switch means an attacker can re-map or observe traffic between accelerators and the host, which is below every isolation boundary the OS or hypervisor enforces, and it persists in the switch's own flash across any host reimage and across tenant handoff. This is the kind of component operators rarely inventory at all, so exposure tends to be unmeasured rather than accepted.","attack_vector":"A privileged local user on the host running the vendor firmware/management tooling against the switch.","remediation":"Update the PCIe switch firmware package to MR4_1.0b1 or later. In practice this is delivered by the system builder (Supermicro, Gigabyte, Quanta, Wiwynn and the GPU-system ODMs) rather than by Intel directly, and it is one of the least reliably shipped firmware components in the stack - you will often have to ask the integrator for it by name. Requires the node quiesced and power-cycled. Remove the LED mode toggle tool and other switch management binaries from tenant-visible host images.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24323","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01339.html"],"status":"curated"},{"id":"CVE-2025-32003","cve":"CVE-2025-32003","aliases":[],"title":"Intel Ethernet Network Adapter E810 (100GbE) firmware: Out-of-bounds read in 100GbE E810 firmware reachable from privileged host software, causing denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Network Adapter E810 (100GbE) firmware","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in 100GbE E810 firmware reachable from privileged host software, causing denial of service. The version threshold here (cvl fw 1.7.6, cpk 1.3.7) is different again from the other E810 advisories — the E810 firmware CVE trail is long enough that version-by-CVE tracking is not workable and operators should treat it as a rolling minimum-version policy.","attack_vector":"Privileged local software on the host (Ring 0).","remediation":"Flash E810 firmware to at least cvl fw 1.7.6 / cpk 1.3.7 — and given the later advisories, go straight to 1.7.8.x or newer. Cold power cycle, per node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32003"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-32386","cve":"CVE-2025-32386","aliases":[],"title":"Helm: Decompression bomb chart exhausts memory on the rendering host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Decompression bomb chart exhausts memory on the rendering host","attack_vector":"Malicious chart","remediation":"Upgrade Helm","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32386"],"status":"curated"},{"id":"CVE-2025-32387","cve":"CVE-2025-32387","aliases":[],"title":"Helm: Deeply nested JSON Schema references cause stack overflow","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Deeply nested JSON Schema references cause stack overflow","attack_vector":"Malicious chart","remediation":"Upgrade Helm","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32387"],"status":"curated"},{"id":"CVE-2025-36608","cve":"CVE-2025-36608","aliases":[],"title":"Dell SmartFabric OS10 (XML external entity): XXE in SmartFabric OS10 before 10.6.0.5, reachable remotely by a low-privileged attacker. XXE on a switch…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell SmartFabric OS10 (XML external entity)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"XXE in SmartFabric OS10 before 10.6.0.5, reachable remotely by a low-privileged attacker. XXE on a switch typically yields file read from the switch's filesystem — which holds the running configuration, and therefore the fabric's secrets and its full topology.","attack_vector":"Low-privileged attacker with remote access to the OS10 management interface.","remediation":"Upgrade OS10 to 10.6.0.5 or later plus reload. Related file-exposure issue in the same release: CVE-2025-30103.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-36608","https://nvd.nist.gov/vuln/detail/CVE-2025-30103"],"status":"curated"},{"id":"CVE-2025-48942","cve":"CVE-2025-48942","aliases":[],"title":"vLLM (`/v1/completions` guided decoding): Invalid `json_schema` kills the server","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (`/v1/completions` guided decoding)","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Invalid `json_schema` kills the server","attack_vector":"Unauthenticated network to an exposed serving port","remediation":"Upgrade to 0.9.0+; single-request DoS against a shared serving tier","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48942"],"status":"curated"},{"id":"CVE-2025-55199","cve":"CVE-2025-55199","aliases":[],"title":"Helm: Crafted JSON Schema causes OOM termination of the renderer","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Crafted JSON Schema causes OOM termination of the renderer","attack_vector":"Malicious chart","remediation":"Upgrade Helm to 3.18.5+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55199"],"status":"curated"},{"id":"CVE-2025-64433","cve":"CVE-2025-64433","aliases":[],"title":"KubeVirt: A VM reads arbitrary files from the virt-launcher pod filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"A VM reads arbitrary files from the virt-launcher pod filesystem","attack_vector":"Any tenant VM","remediation":"Upgrade KubeVirt to 1.5.3/1.6.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64433"],"status":"curated"},{"id":"CVE-2025-7445","cve":"CVE-2025-7445","aliases":[],"title":"secrets-store-sync-controller: Service account tokens disclosed in controller logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"secrets-store-sync-controller","year":"2025","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Service account tokens disclosed in controller logs","attack_vector":"Anyone with log-pipeline read access","remediation":"Controller rollout, no GPU drain; rotate exposed tokens; scrub logs","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-13208","cve":"CVE-2026-13208","aliases":[],"title":"KubeVirt: virt-handler notify server derives VMI identity from the request body without validating the connection","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"virt-handler notify server derives VMI identity from the request body without validating the connection; cross-VM event spoofing","attack_vector":"An attacker with virt-launcher access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13208"],"status":"curated"},{"id":"CVE-2026-18727","cve":"CVE-2026-18727","aliases":[],"title":"open-iscsi iscsiuio (DHCPv6 handling): Integer underflow and out-of-bounds read in iscsiuio's DHCPv6 handling. iscsiuio is the userspace daemon that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"open-iscsi iscsiuio (DHCPv6 handling)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Integer underflow and out-of-bounds read in iscsiuio's DHCPv6 handling. iscsiuio is the userspace daemon that drives iSCSI offload on Broadcom and QLogic adapters, and it processes DHCPv6 from the network to bring up the offload interface — so this is unauthenticated network input reaching a privileged storage daemon. Relevant to GPU clusters that boot or mount datasets over iSCSI, which is still common on the cheaper storage tiers.","attack_vector":"Unauthenticated, adjacent — a rogue DHCPv6 responder on the storage network. DHCPv6 has no authentication and responds fastest-wins, so this needs only presence on the segment.","remediation":"Upgrade open-iscsi and restart iscsiuio — package upgrade with a service restart; iSCSI sessions may briefly drop, so drain storage-dependent workloads first. Independently: disable IPv6 on storage networks that do not need it, or enforce DHCPv6 guard on the storage VLAN at the switch, both live config changes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-18727"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-24182","cve":"CVE-2026-24182","aliases":[],"title":"GPU Display Driver: DoS / privesc (race in GPU memory management)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS / privesc (race in GPU memory management)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24182","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-667"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24197","cve":"CVE-2026-24197","aliases":[],"title":"GPU Display Driver: Cross-tenant info disclosure via GPU memory leakage","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Cross-tenant info disclosure via GPU memory leakage","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; drain + reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24197","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-1188"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24204","cve":"CVE-2026-24204","aliases":[],"title":"NVIDIA FLARE SDK: Auth bypass via insufficient input validation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA FLARE SDK","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Auth bypass via insufficient input validation","attack_vector":"Network peer","remediation":"Upgrade FLARE; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24204","https://github.com/NVIDIA/product-security/tree/main/2026/5819"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-20"]},{"id":"CVE-2026-24514","cve":"CVE-2026-24514","aliases":[],"title":"ingress-nginx: Admission controller denial of service","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Admission controller denial of service","attack_vector":"Any pod on the cluster network","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-3864","cve":"CVE-2026-3864","aliases":[],"title":"CSI Driver NFS: Path traversal via `subDir` lets a tenant delete unintended directories on the shared NFS server","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CSI Driver NFS","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Path traversal via `subDir` lets a tenant delete unintended directories on the shared NFS server; cross-tenant data destruction","attack_vector":"Cluster user able to create a PV/PVC with a crafted subDir","remediation":"DaemonSet/controller rollout; add admission validation on subDir. High priority for neoclouds sharing one NFS backend across tenants","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-3865","cve":"CVE-2026-3865","aliases":[],"title":"CSI Driver SMB: Same `subDir` path traversal against a shared SMB server","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CSI Driver SMB","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Same `subDir` path traversal against a shared SMB server","attack_vector":"Cluster user able to create a PV/PVC with a crafted subDir","remediation":"DaemonSet/controller rollout; validate subDir at admission","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-46433","cve":"CVE-2026-46433","aliases":[],"title":"lldpd (802.1Q VLAN tag stripping in lldpd_decode): lldpd strips 802.1Q VLAN tags by memmove-ing the frame payload four bytes left, and the byte count is not…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"lldpd (802.1Q VLAN tag stripping in lldpd_decode)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"lldpd strips 802.1Q VLAN tags by memmove-ing the frame payload four bytes left, and the byte count is not correctly bounded — so a crafted tagged frame drives a bad copy. Worth noting for anyone running lldpd on trunk ports, which in a leaf/spine fabric is most of them.","attack_vector":"Unauthenticated, adjacent — a crafted VLAN-tagged Ethernet frame on a port where lldpd is listening.","remediation":"Upgrade lldpd to 1.0.22 or later and restart the daemon. Package upgrade plus service restart; on switch NOSes it comes with the image update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46433"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-47155","cve":"CVE-2026-47155","aliases":[],"title":"vLLM (revision pinning): Revision pinning does not apply to all model artifacts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (revision pinning)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Revision pinning does not apply to all model artifacts → supply-chain drift","attack_vector":"Poisoned Hub repo where a pinned revision is silently not enforced","remediation":"Upgrade to 0.22.0+. Undermines model-supply-chain controls the provider may be advertising","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47155"],"status":"curated"},{"id":"CVE-2026-47481","cve":"CVE-2026-47481","aliases":[],"title":"Triton Inference Server: MITM via weak TLS verification","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"MITM via weak TLS verification","attack_vector":"Network attacker on the serving path","remediation":"Upgrade Triton; verify TLS config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47481","https://github.com/NVIDIA/product-security/tree/main/2026/5853"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-288"]},{"id":"CVE-2026-47606","cve":"CVE-2026-47606","aliases":[],"title":"NVIDIA Triton Inference Server: An absolute path traversal reaches code execution and information disclosure - the attacker reads or writes…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An absolute path traversal reaches code execution and information disclosure - the attacker reads or writes outside the model repository. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47606","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-36"]},{"id":"CVE-2026-47620","cve":"CVE-2026-47620","aliases":[],"title":"NVIDIA Dynamo: DoS / corruption via race in resource sync","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"DoS / corruption via race in resource sync","attack_vector":"Concurrent inference clients","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47620","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-362"]},{"id":"CVE-2026-47621","cve":"CVE-2026-47621","aliases":[],"title":"NVIDIA Dynamo: TOCTOU race","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"TOCTOU race","attack_vector":"Local/network client","remediation":"Bump Dynamo; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47621","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-367"]},{"id":"CVE-2026-54235","cve":"CVE-2026-54235","aliases":[],"title":"vLLM - sampling parameter validation: Temperature validation uses strict comparison operators, so boundary values slip through the gate silently.…","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM - sampling parameter validation","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"Temperature validation uses strict comparison operators, so boundary values slip through the gate silently. In a multi-tenant serving deployment this is a request-parameter validation gap rather than memory corruption, but it lets a caller push the engine into a sampling state the operator believed was fenced off - which matters where sampling parameters are part of a tenant-facing safety or cost control.","attack_vector":"Network. Any client able to submit a request with sampling parameters to the vLLM endpoint.","remediation":"Upgrade to vLLM 0.23.1rc0 or later and roll the serving deployment. Cost: rolling restart of the inference tier; no driver or firmware change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-54235"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-63457","cve":"CVE-2026-63457","aliases":["HPESBHF05090"],"title":"HPE iLO 6 (denial of service): An unauthenticated attacker on an adjacent network can knock out iLO 6 availability. No data is read or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 6 (denial of service)","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"An unauthenticated attacker on an adjacent network can knock out iLO 6 availability. No data is read or written - the damage is losing the out-of-band plane itself. For a GPU fleet that is a real operational hit rather than a cosmetic one: when iLO is down you cannot power-cycle a hung training node, cannot console into it to see why the job died, and cannot push firmware or re-provision it. A coordinated version of this against a rack turns a recoverable incident into a physical dispatch. Affects iLO 6 before v1.78, so it covers the Gen11 fleet.","attack_vector":"Adjacent network, unauthenticated - anything sharing the management segment with the iLOs. No account, no host access, no user interaction.","remediation":"Flash iLO 6 to v1.78 or later. Out-of-band, per-node, no host reboot and no job drain - which makes this a cheap fix relative to the availability risk it removes. Because the vector is adjacent-network and unauthenticated, network segmentation is the meaningful compensating control while the rollout runs: keep BMCs off any segment shared with general-purpose hosts.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbhf05090en_us&docLocale=en_US","https://nvd.nist.gov/vuln/detail/CVE-2026-63457"],"status":"curated"},{"id":"CVE-2026-8045","cve":"CVE-2026-8045","aliases":[],"title":"Schneider Electric Data Center Expert - SOAP service endpoints: XML external entity processing on DCE SOAP endpoints lets an authenticated user read server-side files. On a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric Data Center Expert - SOAP service endpoints","year":"2026","cvss_score":6.5,"severity":"medium","kev":false,"impact":"XML external entity processing on DCE SOAP endpoints lets an authenticated user read server-side files. On a DCIM appliance that means configuration and credential material for the power and cooling estate - the recurring theme with DCE is that any read primitive is a facility-wide credential leak.","attack_vector":"Any user with a DCE account submitting crafted XML to the SOAP endpoints.","remediation":"Apply the Schneider fix for DCE. Audit DCE account holders in the same pass. Cheap software update; the credential rotation afterwards is the expensive half.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-8045"],"status":"curated"},{"id":"NCVD-2019-003-rnic-on-board-sram-metadata-cach","cve":null,"aliases":["Pythia","RDMA remote side channel","RNIC SRAM/PTE cache timing attack","Tsai, Payer, Zhang - USENIX Security 2019"],"title":"RNIC on-board SRAM metadata cache (page table entries, QP context) - most widely deployed RDMA NIC: TENANT ISOLATION: RNICs cache page-table entries and connection context in a small on-board SRAM…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RNIC on-board SRAM metadata cache (page table entries, QP context) - most widely deployed RDMA NIC","year":"2019","cvss_score":6.5,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: RNICs cache page-table entries and connection context in a small on-board SRAM and spill to host memory over PCIe when it overflows. Pythia turns the resulting timing difference into a remote side channel: an attacker on one client machine learns the memory access patterns of victims on other client machines against a shared in-memory data service. No memory contents are read directly, but access patterns over a key-value store or a shared embedding table are often enough to recover which records a victim touched. In an AI cluster the same primitive applies to shared parameter servers and RDMA-backed caches, where access pattern equals query content.","attack_vector":"The attacker is an ordinary RDMA client of the same server - no special privilege, no injection needed. They issue their own RDMA reads to addresses chosen to contend for specific RNIC SRAM cache sets, then measure completion latency to infer whether a victim's access evicted their entry. The authors reverse-engineered the memory architecture of the most widely deployed RNIC to make the eviction sets precise, raising the channel's efficiency substantially.","remediation":"No patch. Mitigations are all structural: do not let mutually untrusted tenants share an RNIC or a server-side RDMA data service; partition the server's registered memory so different tenants' regions do not share cache sets; or add deliberate noise/padding to server-side access patterns (application change with a throughput cost). Where a DPU fronts the fabric, terminating tenant connections on separate DPU cores reduces sharing. Scheduling policy - not co-locating untrusted tenants on the same RDMA service - is the realistic control and costs bin-packing efficiency, not downtime.","references":["https://www.usenix.org/conference/usenixsecurity19/presentation/tsai","https://www.usenix.org/conference/usenixsecurity22/presentation/xing"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2022-003-rocev2-congestion-control-dcqcn","cve":null,"aliases":["DCQCN manipulation","RoCE congestion-control abuse","CNP spoofing","arXiv:2207.10898"],"title":"RoCEv2 congestion control - DCQCN, ECN marking and Congestion Notification Packets: FABRIC DOS: DCQCN reacts to ECN marks by having the receiver send Congestion Notification Packets that make…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RoCEv2 congestion control - DCQCN, ECN marking and Congestion Notification Packets","year":"2022","cvss_score":6.5,"severity":"medium","kev":false,"impact":"FABRIC DOS: DCQCN reacts to ECN marks by having the receiver send Congestion Notification Packets that make the sender cut its rate. CNPs are unauthenticated RoCE packets like any other, so an attacker who can spoof onto the fabric can forge CNPs at a victim's sender and drive its rate to the floor while their own traffic is unaffected - a targeted bandwidth-theft and starvation primitive. Conversely, a tenant whose NIC simply ignores CNPs takes an unfair share and pushes everyone else into PFC. Published measurement shows the native PFC-based scheme already suffers unfairness and head-of-line blocking, and that congestion-control choice materially changes distributed DNN training time, so the manipulation lands directly on job completion times in a GPU cluster.","attack_vector":"Forged CNPs need only a spoofed source GID and the victim's QP number - the same predictability that makes packet injection work. Rate-ignoring is even simpler: run a NIC configuration or a custom firmware/driver that under-responds to congestion notifications, which looks like a tuning choice rather than an attack. Both are invisible to host-level monitoring; they show up only as unexplained throughput asymmetry between tenants.","remediation":"Config change: enforce switch-side per-tenant rate limiting and ECN marking policy rather than trusting endpoint congestion response, apply source-address filtering so CNPs cannot be spoofed across tenants, and keep tenants in separate traffic classes so an unresponsive one cannot starve others. Standardise and lock the DCQCN parameter set through the NIC driver configuration (mlxconfig / sysfs) so tenants cannot retune their own NICs - a driver-level config change applied at provisioning, no reboot. Per-tenant switch queues are the durable fix and may require a QoS profile change plus a switch reload on constrained platforms. Monitor per-QP CNP counts as a detection signal.","references":["https://arxiv.org/abs/2207.10898","https://arxiv.org/abs/1806.08159"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"NCVD-2023-006-rnic-microarchitectural-resource","cve":null,"aliases":["Husky","RDMA performance isolation test suite","RNIC microarchitecture resource contention","Kong et al., NSDI 2023"],"title":"RNIC microarchitectural resources (NIC cache, processing units) under multi-tenant RDMA: TENANT ISOLATION: This is the paper that established RDMA performance isolation in the cloud is not solved.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RNIC microarchitectural resources (NIC cache, processing units) under multi-tenant RDMA","year":"2023","cvss_score":6.5,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: This is the paper that established RDMA performance isolation in the cloud is not solved. The authors built a test suite (released as host-bench/husky) that models how RDMA operations consume RNIC microarchitecture resources, and report that it breaks every existing performance-isolation solution in various scenarios - a result acknowledged and reproduced by one of the largest RDMA NIC vendors. For an operator selling RDMA into guest VMs or containers, this means a neighbouring tenant can degrade your customer's collective-communication bandwidth at will, and no NIC-level QoS knob currently on the market reliably prevents it. Jobs miss their step-time SLOs with no attributable cause in host metrics.","attack_vector":"A co-resident tenant issues RDMA verb patterns chosen to thrash specific RNIC resources - for example many small operations across many queue pairs and memory regions to blow out the NIC's address-translation cache, or operation mixes that monopolise particular NIC processing stages. The attacker needs nothing more than normal RDMA access on a shared NIC; the damage is done inside the NIC where host-side rate limiters and cgroups have no visibility.","remediation":"No patch. Run the Husky suite against your own NIC/firmware/isolation configuration before promising RDMA SLAs - that is a test-harness exercise, not a change window. Practical controls: give each tenant a dedicated VF with vendor rate limiters plus caps on QP and MR counts (driver config change, applied at VF creation), keep per-tenant working sets small enough to stay resident in NIC cache, and where the risk is unacceptable, dedicate physical NICs. Upgrading to newer RNIC generations with larger caches and better per-VF quotas helps but does not close it - firmware flash plus driver upgrade, rolling host reboots.","references":["https://www.usenix.org/conference/nsdi23/presentation/kong","https://github.com/host-bench/husky"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2018-6622","cve":"CVE-2018-6622","aliases":[],"title":"TPM 2.0 (S3 sleep PCR reset): Platform Configuration Registers can be reset without a full platform restart by abusing the S3 sleep path…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TPM 2.0 (S3 sleep PCR reset)","year":"2018","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Platform Configuration Registers can be reset without a full platform restart by abusing the S3 sleep path, letting an attacker replay chosen measurements into a TPM that should only ever accumulate them. The effect is that a machine which booted a tampered image can present PCR values identical to a clean boot, defeating sealed-storage unlock policies and remote attestation. Any control you built on 'the PCRs cannot lie' stops holding.","attack_vector":"Local attacker with the ability to put the system into and out of S3 sleep - so a tenant with root on a bare-metal node, or anyone with console access.","remediation":"Platform firmware/BIOS update from the OEM (per node, reboot required) that correctly re-establishes the static root of trust across sleep. Practical compensating control on servers: disable S3 suspend entirely in BIOS, which most datacenter nodes never use anyway - a config-only change that eliminates the trigger. Do that first, then patch on the normal cycle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6622","https://kb.cert.org/vuls/id/922681"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-14308","cve":"CVE-2020-14308","aliases":["BootHole family"],"title":"GRUB2 (grub_malloc allocator): GRUB's allocator never checks the requested size for arithmetic overflow, so a tenant who can influence any…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (grub_malloc allocator)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"GRUB's allocator never checks the requested size for arithmetic overflow, so a tenant who can influence any parsed structure gets a heap that is smaller than GRUB believes it is. The payoff is arbitrary write inside the bootloader, which means loading an unsigned kernel with Secure Boot still reporting green. On a bare-metal GPU node that is the difference between 'the next tenant gets a clean box' and 'the next tenant gets the last tenant's rootkit'.","attack_vector":"A tenant with root on a node they rented, or anyone who can write the EFI System Partition (including via BMC virtual media). Not remote on its own - it is the persistence half of a two-stage attack.","remediation":"grub2 package update plus reboot on every node. The package update alone does not close it: the old signed GRUB binary stays trusted until the UEFI revocation list (dbx) is updated, and pushing dbx before every node is on the new shim/GRUB will brick nodes at next boot. Plan it as two passes with a verification gate between them.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14308","https://access.redhat.com/security/cve/CVE-2020-14308"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-14309","cve":"CVE-2020-14309","aliases":["BootHole family"],"title":"GRUB2 (squashfs symlink parser): Integer overflow in grub_squash_read_symlink lets a crafted squashfs image drive a heap overflow inside GRUB.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (squashfs symlink parser)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow in grub_squash_read_symlink lets a crafted squashfs image drive a heap overflow inside GRUB. Attacker-chosen code runs before the kernel and before any measured-boot evidence the operator would trust, so an implant planted here is invisible to every agent running in the tenant OS.","attack_vector":"Requires control of a filesystem image GRUB will read - the boot partition on a node the attacker already had, or an image served over the provisioning path.","remediation":"grub2 package update + reboot per node. Real closure needs the dbx revocation of the old signed GRUB, which is a separate and riskier rollout. On GPU nodes the reboot means draining running training jobs, so batch it with an existing maintenance window rather than doing it alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14309","https://access.redhat.com/security/cve/CVE-2020-14309"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-14310","cve":"CVE-2020-14310","aliases":["BootHole family"],"title":"GRUB2 (read_section_from_string): Integer overflow while reading a section string overflows the heap and gives control of GRUB before the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (read_section_from_string)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow while reading a section string overflows the heap and gives control of GRUB before the kernel loads. Practical outcome is a Secure Boot bypass that survives an OS reinstall, because the compromise lives in the boot partition rather than the root filesystem.","attack_vector":"Local write access to boot-time data on the node - a previous tenant, an operator with remote-hands, or a BMC-mounted virtual disk.","remediation":"grub2 package update + reboot. Follow with a dbx update, sequenced after every node is confirmed on the fixed binary. Config-only mitigation does not exist for this class.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14310","https://access.redhat.com/security/cve/CVE-2020-14310"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-14311","cve":"CVE-2020-14311","aliases":["BootHole family"],"title":"GRUB2 (ext2/ext4 symlink reader): Integer overflow in grub_ext2_read_link on a crafted ext filesystem yields a heap overflow in the bootloader.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (ext2/ext4 symlink reader)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow in grub_ext2_read_link on a crafted ext filesystem yields a heap overflow in the bootloader. Because ext4 is what almost every Linux GPU node actually boots from, this one is reachable on stock images rather than exotic filesystems.","attack_vector":"Attacker controls the boot filesystem - realistically a tenant who had root on the box, or anyone who can attach media over the BMC.","remediation":"grub2 package update + reboot per node; then the dbx revocation pass. Note that a node that PXE-boots a fresh image every provisioning cycle is not automatically safe - the vulnerable GRUB is in the image you serve, so fix the golden image too, not just running nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-14311","https://access.redhat.com/security/cve/CVE-2020-14311"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-15706","cve":"CVE-2020-15706","aliases":["BootHole family"],"title":"GRUB2 (script function redefinition): Use-after-free when a GRUB script redefines a function while that function is executing. Gives arbitrary code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (script function redefinition)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Use-after-free when a GRUB script redefines a function while that function is executing. Gives arbitrary code execution in the bootloader from nothing more than a modified grub.cfg, which on most distros is not itself signature-checked.","attack_vector":"Anyone who can write grub.cfg - local root, or a previous bare-metal tenant. This is the classic BootHole shape: config file trusted more than it deserves.","remediation":"grub2 package update + reboot per node. Also worth checking that grub.cfg is not writable from a tenant-reachable partition on your image layout.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15706","https://ubuntu.com/security/CVE-2020-15706"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-8559","cve":"CVE-2020-8559","aliases":[],"title":"Kubernetes (kube-apiserver): Unvalidated redirect on proxied upgrade requests lets a compromised node escalate to other nodes","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Unvalidated redirect on proxied upgrade requests lets a compromised node escalate to other nodes","attack_vector":"An attacker who already owns one node, pivoting cluster-wide","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8559"],"status":"curated"},{"id":"CVE-2021-3418","cve":"CVE-2021-3418","aliases":[],"title":"GRUB2 (grub-install shim_lock regression): GRUB 2.06~rc1 reintroduced the earlier direct-boot flaw: grub-install could produce an installation that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (grub-install shim_lock regression)","year":"2021","cvss_score":6.4,"severity":"medium","kev":false,"impact":"GRUB 2.06~rc1 reintroduced the earlier direct-boot flaw: grub-install could produce an installation that skips shim and therefore skips kernel signature verification. Nodes you believed you had already fixed silently regress when they are rebuilt with a newer GRUB.","attack_vector":"Not directly attacker-triggered - it is a build/provisioning regression that reopens the earlier bypass. The attacker then needs only local root.","remediation":"grub2 package update + reboot, and re-verify the boot chain on any node reimaged between the original BootHole fix and this one. Worth a fleet-wide audit script that asserts shim is in the chain, not a one-off check.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3418","https://access.redhat.com/security/cve/CVE-2021-3418"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-26362","cve":"CVE-2022-26362","aliases":["XSA-401"],"title":"Xen (x86 PV): Race condition in typeref acquisition - PV guest escalates to host privilege","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 PV)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Race condition in typeref acquisition - PV guest escalates to host privilege","attack_vector":"Tenant VM guest (PV)","remediation":"Hypervisor patch + host reboot with guest evacuation, or use Xen livepatch if the deployment supports it. Simplest structural fix: stop offering PV guests, run PVH/HVM only","references":["https://xenbits.xen.org/xsa/advisory-401.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-30773","cve":"CVE-2022-30773","aliases":["INSYDE-SA-2022042"],"title":"Insyde InsydeH2O (IhisiSmm parameter buffer, DMA TOCTOU): IHISI is Insyde's own firmware-services interface - the channel BIOS update and configuration tooling talks…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IhisiSmm parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"IHISI is Insyde's own firmware-services interface - the channel BIOS update and configuration tooling talks to. Its parameter buffer sits outside SMRAM, so a device can rewrite the parameters after SMM has validated them and before SMM uses them. What the attacker reaches through this particular driver is the firmware update path itself, which is the shortest route from a peripheral to a permanent SPI implant on a GPU node.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.4 / 05.44.23 and 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30773","https://www.insyde.com/security-pledge/SA-2022042"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-30774","cve":"CVE-2022-30774","aliases":["INSYDE-SA-2022043"],"title":"Insyde InsydeH2O (PnpSmm parameter buffer, DMA TOCTOU): The plug-and-play SMI handler's parameters can be swapped by DMA between the check and the use. PnpSmm is the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PnpSmm parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The plug-and-play SMI handler's parameters can be swapped by DMA between the check and the use. PnpSmm is the driver that writes SMBIOS/platform-description data, so corrupting it lets an attacker both corrupt SMRAM and poison the hardware inventory the OS and your fleet-management tooling read back - a node can be made to misreport what it is.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.29, 5.3 / 05.36.25, 5.4 / 05.44.25, 5.5 / 05.52.25.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-30774","https://www.insyde.com/security-pledge/SA-2022043"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31243","cve":"CVE-2022-31243","aliases":["INSYDE-SA-2022044"],"title":"Insyde InsydeH2O (FvbServicesRuntimeDxe input buffer, DMA TOCTOU): Firmware Volume Block services are the abstraction SMM uses to read and write the SPI flash. A DMA race on…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FvbServicesRuntimeDxe input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Firmware Volume Block services are the abstraction SMM uses to read and write the SPI flash. A DMA race on this handler's input buffer means an attacker influences what gets written to the boot flash - the most direct path in this whole batch to firmware that survives OS reinstall, disk replacement and reimaging between tenants.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.21, 5.3 / 05.36.21, 5.4 / 05.44.21, 5.5 / 05.52.21. Verify SPI write protection (BIOS Lock Enable, protected range registers) is actually set - it blunts the primitive even before the flash lands. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31243","https://www.insyde.com/security-pledge/SA-2022044"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31602","cve":"CVE-2022-31602","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: An out-of-bounds write in IpSecDxe, exploitable against a preconditioned heap, reaches firmware code…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"An out-of-bounds write in IpSecDxe, exploitable against a preconditioned heap, reaches firmware code execution. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31602","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31603","cve":"CVE-2022-31603","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: Improper array index validation in IpSecDxe gives firmware-phase code execution against preconditioned global…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Improper array index validation in IpSecDxe gives firmware-phase code execution against preconditioned global data. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5367. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31603","https://github.com/NVIDIA/product-security/tree/main/2022/5367"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-129"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-31667","cve":"CVE-2022-31667","aliases":[],"title":"Harbor: Robot accounts in other projects can be updated","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Robot accounts in other projects can be updated","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31667"],"status":"curated"},{"id":"CVE-2022-32266","cve":"CVE-2022-32266","aliases":["INSYDE-SA-2022045"],"title":"Insyde InsydeH2O (PcdSmmDxe parameter buffer, DMA TOCTOU): A DMA race against the Platform Configuration Database SMI handler corrupts ACPI fields and adjacent memory.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (PcdSmmDxe parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"A DMA race against the Platform Configuration Database SMI handler corrupts ACPI fields and adjacent memory. ACPI tables are consumed by the OS after boot, so this driver is the one that reaches OS-visible platform description - an attacker can corrupt what the kernel believes about the hardware, not just SMRAM. Insyde notes exploitation needs detailed knowledge of the PCD contents on the specific platform, which raises the bar but does not close the race.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in the kernel releases named in the advisory (Insyde does not enumerate per-kernel versions for this one).  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32266","https://www.insyde.com/security-pledge/SA-2022045"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-32267","cve":"CVE-2022-32267","aliases":["INSYDE-SA-2022046"],"title":"Insyde InsydeH2O (SmmResourceCheckDxe input buffer, DMA TOCTOU): The sharpest irony in the batch: SmmResourceCheckDxe is the driver whose job is to validate SMM resource…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (SmmResourceCheckDxe input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The sharpest irony in the batch: SmmResourceCheckDxe is the driver whose job is to validate SMM resource access, and its own input buffer is racy. An attacker who wins this race corrupts SMRAM through the guard rather than around it, which means the check other handlers rely on can be made to approve what it should reject.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-32267","https://www.insyde.com/security-pledge/SA-2022046"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33906","cve":"CVE-2022-33906","aliases":["INSYDE-SA-2022048"],"title":"Insyde InsydeH2O (FwBlockServiceSmm input buffer, DMA TOCTOU): The firmware block service is the SMM-side writer for the boot flash. Racing its input buffer with DMA…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (FwBlockServiceSmm input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The firmware block service is the SMM-side writer for the boot flash. Racing its input buffer with DMA corrupts SMRAM and puts the attacker on the flash-write path - a persistent implant that no reimage between tenant leases will remove, sitting underneath Secure Boot and underneath whatever the node later attests.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23. Confirm SPI flash write protection is enforced as a stopgap. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33906","https://www.insyde.com/security-pledge/SA-2022048"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33907","cve":"CVE-2022-33907","aliases":["INSYDE-SA-2022049"],"title":"Insyde InsydeH2O (IdeBusDxe SMI input buffer, DMA TOCTOU): SMRAM corruption via a DMA race on the legacy IDE/ATA bus driver. Worth noting for fleet operators that this…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IdeBusDxe SMI input buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"SMRAM corruption via a DMA race on the legacy IDE/ATA bus driver. Worth noting for fleet operators that this driver is usually only live when CSM / legacy storage compatibility is enabled - so unlike most of this batch, there is a real chance the attack surface is simply not present on a modern UEFI-only server profile.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.25, 5.3 / 05.36.25, 5.4 / 05.44.25. Genuine config lever on this one: disable CSM / legacy storage support in BIOS on UEFI-only nodes, which removes the driver rather than just patching it. The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33907","https://www.insyde.com/security-pledge/SA-2022049"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33982","cve":"CVE-2022-33982","aliases":["INSYDE-SA-2022052"],"title":"Insyde InsydeH2O (Int15ServiceSmm parameter buffer, DMA TOCTOU): DMA race against the legacy INT15 services SMI handler corrupts SMRAM. This is legacy BIOS callback plumbing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (Int15ServiceSmm parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"DMA race against the legacy INT15 services SMI handler corrupts SMRAM. This is legacy BIOS callback plumbing that most operators do not know is still resident on a modern server image - it is, and it is reachable, which is the general lesson of this batch: the attack surface is the union of every driver the IBV compiled in, not the subset you actually use.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.2 / 05.27.23, 5.3 / 05.36.23, 5.4 / 05.44.23, 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33982","https://www.insyde.com/security-pledge/SA-2022052"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-33986","cve":"CVE-2022-33986","aliases":["INSYDE-SA-2022056"],"title":"Insyde InsydeH2O (VariableRuntimeDxe parameter buffer, DMA TOCTOU): The UEFI variable store is where the Secure Boot key databases live - PK, KEK, db and dbx. A DMA race on this…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (VariableRuntimeDxe parameter buffer, DMA TOCTOU)","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"The UEFI variable store is where the Secure Boot key databases live - PK, KEK, db and dbx. A DMA race on this handler's parameter buffer lets an attacker corrupt SMRAM through the driver that guards the keys deciding what firmware and bootloaders are allowed to run. Of everything in this batch, this is the driver whose compromise most directly invalidates the boot-integrity story you sell to tenants.","attack_vector":"An attacker able to drive DMA at host memory while the SMI handler is mid-flight - a malicious PCIe device, a peripheral running attacker-flashed firmware (NIC, GPU, NVMe), or a tenant with a passed-through device that is not behind a correctly configured IOMMU. Notably does NOT require host root, which is what separates this family from the ordinary SMM callout bugs.","remediation":"Firmware flash from the server OEM, not from Insyde - the fixed Insyde kernel has to be rebased by Dell/HPE/Lenovo/Supermicro and re-qualified before it reaches you, which for this batch ran months behind Insyde's own release. One reboot per node, so schedule it against a GPU drain. Fixed in kernel 5.4 / 05.44.23 and 5.5 / 05.52.23.  The compensating control that actually works here is the IOMMU, and Insyde says so in the advisory: enable VT-d/AMD-Vi with pre-boot DMA protection so the ACPI runtime buffer the handler reads is not reachable by an untrusted device. That is a BIOS setting, deployable fleet-wide without a flash, and it should be on already on any node that passes devices through to tenants. Patch the batch, not the CVE - Insyde filed one advisory per driver for the same defect, so fixing this one leaves every sibling handler reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-33986","https://www.insyde.com/security-pledge/SA-2022056"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42283","cve":"CVE-2022-42283","aliases":[],"title":"DGX-2 BMC: Buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Buffer overflow -> code exec on BMC","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42283","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-120"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-22745","cve":"CVE-2023-22745","aliases":[],"title":"tpm2-tss (Tss2_RC_Decode / Tss2_RC_SetHandler): An 8-bit layer number indexes an array with far fewer entries, so a TPM response code the library did not…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tss (Tss2_RC_Decode / Tss2_RC_SetHandler)","year":"2023","cvss_score":6.4,"severity":"medium","kev":false,"impact":"An 8-bit layer number indexes an array with far fewer entries, so a TPM response code the library did not expect reads or writes outside the buffer - and the path to arbitrary code execution runs through the userspace component that every attestation and key-sealing tool on the node depends on. The disclosed trigger is a man-in-the-middle on the TPM bus returning 0xFFFFFFFF, which ties this directly to the physical-interposer threat model.","attack_vector":"Local, privileged - or an attacker sitting on the TPM's LPC/SPI bus, which is a hardware-interposer attack that a colo tenant or remote-hands contractor can mount.","remediation":"Package update to tpm2-tss 4.0.1 / 3.2.2 or later and restart anything linked against it - no reboot, no firmware flash. One of the genuinely cheap fixes in this cluster, so there is no reason to carry it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22745","https://github.com/tpm2-software/tpm2-tss/security/advisories"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2023-5088","cve":"CVE-2023-5088","aliases":[],"title":"QEMU (IDE/ATAPI): Improper IDE controller reset lets a guest overwrite the host MBR of an attached device","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (IDE/ATAPI)","year":"2023","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Improper IDE controller reset lets a guest overwrite the host MBR of an attached device","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart/live-migration. Also stop exposing raw block devices as IDE to tenant VMs","references":["https://access.redhat.com/security/cve/CVE-2023-5088"],"status":"curated"},{"id":"CVE-2024-21823","cve":"CVE-2024-21823","aliases":[],"title":"Intel DSA/IAA (idxd): Hardware erratum: direct access to Intel DSA/IAA accelerators by an untrusted application allows privilege…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel DSA/IAA (idxd)","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Hardware erratum: direct access to Intel DSA/IAA accelerators by an untrusted application allows privilege escalation","attack_vector":"Tenant process granted direct accelerator access; tenant VM guest with passthrough","remediation":"Microcode + kernel driver update + reboot. Relevant wherever DSA/IAA is exposed to tenants alongside GPUs","references":["https://access.redhat.com/security/cve/CVE-2024-21823"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2024-22278","cve":"CVE-2024-22278","aliases":[],"title":"Harbor: Incorrect permission validation lets authenticated users modify Harbor configuration","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Incorrect permission validation lets authenticated users modify Harbor configuration","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-22278"],"status":"curated"},{"id":"CVE-2024-25620","cve":"CVE-2024-25620","aliases":[],"title":"Helm: Relative path in a chart name writes the chart outside the intended directory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Helm","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Relative path in a chart name writes the chart outside the intended directory","attack_vector":"Malicious chart","remediation":"Upgrade Helm on CI and operator machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25620"],"status":"curated"},{"id":"CVE-2024-8105","cve":"CVE-2024-8105","aliases":["PKfail"],"title":"UEFI Secure Boot Platform Key: ~791 firmware releases across Acer, Dell, Fujitsu, Gigabyte, HP, Intel, Lenovo, Supermicro shipped with AMI's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI Secure Boot Platform Key","year":"2024","cvss_score":6.4,"severity":"medium","kev":false,"impact":"~791 firmware releases across Acer, Dell, Fujitsu, Gigabyte, HP, Intel, Lenovo, Supermicro shipped with AMI's *test* Platform Key, whose private half is public on GitHub. Anyone can sign a bootloader that the platform trusts — complete Secure Boot bypass with a below-OS, reimage-surviving implant","attack_vector":"Local, or supply-chain","remediation":"Requires generating and enrolling a real per-vendor Platform Key, which is a BIOS-level key-enrollment operation, not a patch. Many affected models never received a fixed firmware, so for those the only remediation is hardware replacement or accepting that Secure Boot is decorative","references":["https://www.binarly.io/advisories/brly-2024-005"],"status":"curated","fleet":{"ubiquity":"very common - ~900 device models across Dell, HP, Lenovo, Gigabyte, Supermicro, Intel, Fujitsu, spanning 2012-2024","remediation_pain":"firmware-flash + key re-provisioning - each node needs a genuinely secret PK enrolled and the KEK/db chain re-signed; a leaked private key cannot be patched, only rotated","pain_class":"firmware-flash","why_fleet_wide":"The private Platform Key is public, so Secure Boot is decorative on affected nodes: anyone who can write the ESP can sign a bootkit the firmware trusts, below the OS, across the whole affected SKU population."}},{"id":"CVE-2025-0622","cve":"CVE-2025-0622","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (commands/gpg): Module unload leaves registered hooks behind, so GRUB later calls through freed function pointers. Notable…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (commands/gpg)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Module unload leaves registered hooks behind, so GRUB later calls through freed function pointers. Notable because the affected module is the one meant to verify signatures - the bug is in the verification machinery itself.","attack_vector":"Local, via GRUB command sequences or grub.cfg.","remediation":"grub2 package update + reboot. Part of the same February 2025 distro update as the rest of the batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0622","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0677","cve":"CVE-2025-0677","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (UFS symlink handling): Integer overflow on symlink handling in UFS gives a heap out-of-bounds write and a path to loading unsigned…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (UFS symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow on symlink handling in UFS gives a heap out-of-bounds write and a path to loading unsigned code with Secure Boot on.","attack_vector":"Attacker-supplied UFS filesystem image.","remediation":"grub2 package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0677","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0684","cve":"CVE-2025-0684","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (ReiserFS symlink handling): Same symlink integer-overflow pattern in the ReiserFS parser - heap out-of-bounds write, pre-boot execution","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (ReiserFS symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Same symlink integer-overflow pattern in the ReiserFS parser - heap out-of-bounds write, pre-boot execution.","attack_vector":"Attacker-supplied ReiserFS image.","remediation":"grub2 package update + reboot; or build without the module, since nothing in a modern GPU fleet boots ReiserFS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0684","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0685","cve":"CVE-2025-0685","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (JFS symlink handling): Symlink integer overflow in the JFS parser producing a heap out-of-bounds write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (JFS symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Symlink integer overflow in the JFS parser producing a heap out-of-bounds write.","attack_vector":"Attacker-supplied JFS image.","remediation":"grub2 package update + reboot; or drop the unused module from the build.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0685","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0686","cve":"CVE-2025-0686","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (romfs symlink handling): Symlink integer overflow in the romfs parser producing a heap out-of-bounds write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (romfs symlink handling)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Symlink integer overflow in the romfs parser producing a heap out-of-bounds write.","attack_vector":"Attacker-supplied romfs image.","remediation":"grub2 package update + reboot; or drop the module.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0686","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0689","cve":"CVE-2025-0689","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (UDF filesystem parser): Heap buffer overflow in grub_udf_read_block. UDF is the optical/ISO filesystem, which is precisely what BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (UDF filesystem parser)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Heap buffer overflow in grub_udf_read_block. UDF is the optical/ISO filesystem, which is precisely what BMC virtual media presents - so this one is reachable by anyone with BMC credentials, not only by a local tenant.","attack_vector":"Crafted UDF/ISO image, including one mounted remotely through iDRAC/iLO/XCC virtual media.","remediation":"grub2 package update + reboot. Disable BMC virtual media where you do not need it - that closes the most convenient remote path to this and several other parser bugs at once.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0689","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0690","cve":"CVE-2025-0690","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (read command): Integer overflow in the read command's accumulator writes out of bounds. Straightforward pre-boot memory…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (read command)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow in the read command's accumulator writes out of bounds. Straightforward pre-boot memory corruption from GRUB script.","attack_vector":"Local, via grub.cfg or the GRUB shell.","remediation":"grub2 package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0690","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-1125","cve":"CVE-2025-1125","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (HFS filesystem parser): Integer overflow computing internal buffer sizes from HFS metadata, leading to a heap out-of-bounds write","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HFS filesystem parser)","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Integer overflow computing internal buffer sizes from HFS metadata, leading to a heap out-of-bounds write.","attack_vector":"Attacker-supplied HFS volume, physical or virtual media.","remediation":"grub2 package update + reboot; or build GRUB without HFS support.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1125","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-57852","cve":"CVE-2025-57852","aliases":[],"title":"KServe ModelMesh: Group-writable `/etc/passwd` in the container image","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"KServe ModelMesh","year":"2025","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Group-writable `/etc/passwd` in the container image → privilege escalation inside the container","attack_vector":"Tenant with code execution in a ModelMesh pod","remediation":"Rebuild the container images; no host patch. Provider owns the images if it ships a managed KServe","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-57852"],"status":"curated"},{"id":"CVE-2026-13318","cve":"CVE-2026-13318","aliases":[],"title":"KubeVirt: SSRF in the virt-api port-forward handler via attacker-influenced VMI status IP","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"SSRF in the virt-api port-forward handler via attacker-influenced VMI status IP","attack_vector":"Cluster user with namespace access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-13318"],"status":"curated"},{"id":"CVE-2026-24205","cve":"CVE-2026-24205","aliases":[],"title":"NVIDIA TensorRT-LLM: MULTI-TENANT ISOLATION: concurrent requests race inside TensorRT-LLM and reach data tampering with a changed…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA TensorRT-LLM","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: concurrent requests race inside TensorRT-LLM and reach data tampering with a changed scope. On a shared LLM serving tier, a race between concurrent requests means one tenant's request can affect another's - the mechanism by which response bleed-through happens.","attack_vector":"Network, low privileges. Any client able to issue concurrent requests to the serving endpoint, which is every client.","remediation":"Upgrade TensorRT-LLM per bulletin 5805 and roll the serving deployment. Cost: rolling restart. Until patched, the compensating control is reducing concurrency or dedicating an engine per tenant, both of which cost throughput.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24205","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:N","cwe":["CWE-362"],"fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-24220","cve":"CVE-2026-24220","aliases":[],"title":"TensorRT-LLM: Code exec via insecure deserialization on model load","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Code exec via insecure deserialization on model load","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24220","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-502"]},{"id":"CVE-2026-24259","cve":"CVE-2026-24259","aliases":[],"title":"TensorRT-LLM: Missing authentication in configuration processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"Missing authentication in configuration processing","attack_vector":"Network client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24259","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:H/A:H","cwe":["CWE-306"]},{"id":"CVE-2026-44774","cve":"CVE-2026-44774","aliases":[],"title":"Traefik: A tenant with HTTPRoute creation rights exposes the REST provider handler, bypassing provider isolation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":6.4,"severity":"medium","kev":false,"impact":"A tenant with HTTPRoute creation rights exposes the REST provider handler, bypassing provider isolation","attack_vector":"Cluster user with namespace access","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44774"],"status":"curated"},{"id":"CVE-2019-9508","cve":"CVE-2019-9508","aliases":[],"title":"Vertiv Avocent UMG-4000 universal management gateway: An authenticated admin can plant a maliciously named file in the web application that executes JavaScript…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Vertiv Avocent UMG-4000 universal management gateway","year":"2019","cvss_score":6.3,"severity":"medium","kev":false,"impact":"An authenticated admin can plant a maliciously named file in the web application that executes JavaScript every time any user (including a higher-privileged one) browses to the page listing it — a stepping stone to hijacking another operator's session on the KVM gateway.","attack_vector":"Requires an authenticated administrator account to upload/name the malicious file; the payload then fires against any user who later views that page.","remediation":"Same fixed firmware/software build as the UMG-4000 command-injection issue (CVE-2019-9507) — apply both in the same maintenance window since they land in the same release. One flash per gateway.","references":["https://www.vertiv.com/en-us/support/software-download/it-management/avocent-universal-management-gateway-appliance--software-downloads/"],"status":"curated"},{"id":"CVE-2020-5969","cve":"CVE-2020-5969","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: time-of-check to time-of-use on a shared resource between guest and host plugin. A…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":6.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: time-of-check to time-of-use on a shared resource between guest and host plugin. A tenant that wins the race gets host-side information disclosure or crashes the shared GPU. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5969"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-8554","cve":"CVE-2020-8554","aliases":[],"title":"Kubernetes (kube-apiserver): Any user who can create a Service with externalIPs (or patch LB status) intercepts cluster traffic to that IP","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Any user who can create a Service with externalIPs (or patch LB status) intercepts cluster traffic to that IP; MITM","attack_vector":"Cluster user with namespace access","remediation":"No upstream code fix; deploy an admission policy denying externalIPs and status.loadBalancer patches for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8554"],"status":"curated"},{"id":"CVE-2020-8555","cve":"CVE-2020-8555","aliases":[],"title":"Kubernetes (kube-controller-manager): Half-blind SSRF from the controller manager into the cloud metadata service and internal network","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-controller-manager)","year":"2020","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Half-blind SSRF from the controller manager into the cloud metadata service and internal network","attack_vector":"Cluster user able to create storage objects","remediation":"Rolling control-plane upgrade; block link-local metadata from control-plane nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8555"],"status":"curated"},{"id":"CVE-2021-1061","cve":"CVE-2021-1061","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: the plugin keeps using a resource it validated after the guest has changed it - a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":6.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: the plugin keeps using a resource it validated after the guest has changed it - a double-fetch race that yields host information disclosure or a shared-GPU crash. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM able to race the host's validation window.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1061"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-21334","cve":"CVE-2021-21334","aliases":[],"title":"containerd: Environment variables from an unrelated image leak into a container, exposing another tenant's secrets","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Environment variables from an unrelated image leak into a container, exposing another tenant's secrets","attack_vector":"Any tenant workload on a shared node","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-21334"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-41091","cve":"CVE-2021-41091","aliases":[],"title":"Docker / moby: /var/lib/docker subdirectories world-traversable","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":6.3,"severity":"medium","kev":false,"impact":"/var/lib/docker subdirectories world-traversable; unprivileged host user reaches container filesystems","attack_vector":"Any local user on the node","remediation":"Upgrade Docker Engine and fix directory modes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41091"],"status":"curated"},{"id":"CVE-2022-21820","cve":"CVE-2022-21820","aliases":[],"title":"NVIDIA DCGM - nv-hostengine: A network-reachable caller drives nv-hostengine into an unhandled error condition, reaching limited code…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DCGM - nv-hostengine","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A network-reachable caller drives nv-hostengine into an unhandled error condition, reaching limited code execution and privilege escalation. DCGM runs as root on every GPU node and holds the fleet's telemetry, so it is a high-value target sitting on an open port.","attack_vector":"Network, with low privileges. nv-hostengine listens on TCP 5555 by default and many operators leave it bound beyond localhost so a central collector can scrape it - that binding is the exposure.","remediation":"Update DCGM per bulletin 5328 and restart nv-hostengine. Cost: restarting the host engine briefly interrupts telemetry but does not touch running GPU jobs - no drain needed. While you are there, bind nv-hostengine to localhost and scrape via a local exporter instead of exposing 5555.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21820","https://github.com/NVIDIA/product-security/tree/main/2022/5328"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-20"]},{"id":"CVE-2022-23823","cve":"CVE-2022-23823","aliases":["Hertzbleed (AMD)"],"title":"AMD processors - frequency scaling / power management: A remote or local attacker times operations and infers secret data from how DVFS frequency scaling responds…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - frequency scaling / power management","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A remote or local attacker times operations and infers secret data from how DVFS frequency scaling responds to the data being processed - turning a power side channel into a timing side channel that works over the network. On a GPU host node the exposure is the CPU-side crypto: TLS termination for your API, key material in the control plane, tenant secrets handled by the host. It does not read GPU memory.","attack_vector":"An authenticated attacker able to time operations on the target, including remotely for network-facing crypto. Co-tenancy is not required, which is what made Hertzbleed notable.","remediation":"AMD's guidance is not a microcode patch: the fix is constant-time or blinded implementations in the affected cryptographic software, and optionally disabling frequency boost - which costs you real performance on every workload on the node. Practically: update OpenSSL/libcrypto and any SIKE-like primitives, and treat disabling boost as a last resort. Effectively UNPATCHABLE at the silicon level.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-1038","https://www.hertzbleed.com/","https://nvd.nist.gov/vuln/detail/CVE-2022-23823"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-24436","cve":"CVE-2022-24436","aliases":["Hertzbleed (Intel)","INTEL-SA-00698"],"title":"Intel processors - power management throttling: The Intel half of Hertzbleed: observable behaviour in power-management throttling lets an authenticated user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors - power management throttling","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"The Intel half of Hertzbleed: observable behaviour in power-management throttling lets an authenticated user infer processed data over a network path. Same operator exposure as the AMD variant - host-side cryptography on your GPU nodes and control plane, not GPU memory.","attack_vector":"Authenticated user, remotely exploitable via timing of network-facing crypto operations. No co-tenancy required.","remediation":"Intel likewise did not ship a microcode fix. Mitigation is constant-time cryptographic software, or disabling Turbo Boost / SpeedStep which costs substantial throughput on a GPU host that is already CPU-bound in the data loader. Treat as UNPATCHABLE in hardware; patch the crypto libraries instead.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00698.html","https://www.hertzbleed.com/","https://nvd.nist.gov/vuln/detail/CVE-2022-24436"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-35888","cve":"CVE-2022-35888","aliases":["Hertzbleed (Ampere Altra)"],"title":"Ampere Altra / Altra Max processors: The Arm-server variant of Hertzbleed. Relevant because Ampere Altra is a common host CPU under GPU nodes in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Ampere Altra / Altra Max processors","year":"2022","cvss_score":6.3,"severity":"medium","kev":false,"impact":"The Arm-server variant of Hertzbleed. Relevant because Ampere Altra is a common host CPU under GPU nodes in Arm-based AI racks and in several neocloud fleets - operators who assumed the Hertzbleed story was x86-only still have it.","attack_vector":"Authenticated user able to time operations, including over the network against host crypto.","remediation":"Ampere published a security bulletin rather than a firmware fix. Mitigation is constant-time crypto and, at high cost, disabling frequency scaling. Treat as UNPATCHABLE in hardware.","references":["https://amperecomputing.com/products/security-bulletins/hertzbleed.html","https://nvd.nist.gov/vuln/detail/CVE-2022-35888"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-25163","cve":"CVE-2023-25163","aliases":[],"title":"Argo CD: Repository access credentials leaked in error messages surfaced in the UI and logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2023","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Repository access credentials leaked in error messages surfaced in the UI and logs","attack_vector":"Any Argo CD user who can trigger a sync error","remediation":"Rolling Argo CD upgrade; rotate repo credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25163"],"status":"curated"},{"id":"CVE-2023-26604","cve":"CVE-2023-26604","aliases":[],"title":"systemd: Privilege escalation via the systemctl `less` pager when sudo-granted systemctl is available","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"systemd","year":"2023","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Privilege escalation via the systemctl `less` pager when sudo-granted systemctl is available","attack_vector":"Local user with any sudo systemctl grant","remediation":"Package update; audit sudoers for systemctl grants. No reboot","references":["https://access.redhat.com/security/cve/CVE-2023-26604"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2023-34471","cve":"CVE-2023-34471","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC cryptography / HMAC): A step is missing when the BMC generates its HMAC, so the authentication tag it produces is weaker than…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC cryptography / HMAC)","year":"2023","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A step is missing when the BMC generates its HMAC, so the authentication tag it produces is weaker than intended and can be forged. The result is that an attacker can present traffic the BMC accepts as authentic - loss of authentication as well as confidentiality and integrity. Practically it degrades whatever assurance you thought you had that a management command came from your orchestration system rather than from something else on the wire.","attack_vector":"Adjacent network, requires an existing high-privilege position and user interaction, at high attack complexity. This is a chaining bug rather than a standalone break-in - it matters mostly as the thing that lets an attacker who already has partial management-plane access forge their way further.","remediation":"Firmware flash to SPx_12.2 / SPx_13.0 or later; fixed in early SPx branches, so the real work is verifying that the ODM build actually running on each node is from a fixed branch rather than trusting AMI's fix version. No config-only remediation. Compensate by treating the management network as untrusted transit: bastion-only access, mutual TLS with your own CA, and alerting on BMC sessions that do not originate from your management hosts.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34471"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-0085","cve":"CVE-2024-0085","aliases":[],"title":"vGPU Manager: Improper permission management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Improper permission management","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0085","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-266"]},{"id":"CVE-2024-0129","cve":"CVE-2024-0129","aliases":[],"title":"NVIDIA NeMo: SaveRestoreConnector extracts .tar archives unsafely, so a crafted archive writes files outside the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA NeMo","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"SaveRestoreConnector extracts .tar archives unsafely, so a crafted archive writes files outside the extraction directory and reaches code execution. In an AI datacenter this is the model-and-data supply chain problem: the code runs with whatever the training or inference job holds, which is usually a GPU, a service account, and mounted object storage credentials.","attack_vector":"Requires the job to load an attacker-influenced artifact - a checkpoint, .nemo file, config, tokenizer or dataset. Any pipeline that pulls from a public model hub, a customer bucket, or a tenant-supplied path is in scope.","remediation":"Bump the package to the fixed version in bulletin 5580 and rebuild every training/inference image that embeds it. Cost: image rebuild and job restart; no host driver or firmware change. The durable control is refusing to deserialize untrusted checkpoints at all - prefer safetensors-style formats and treat pickle-bearing artifacts as executable code.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0129","https://github.com/NVIDIA/product-security/tree/main/2024/5580"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:L","cwe":["CWE-22"]},{"id":"CVE-2024-10026","cve":"CVE-2024-10026","aliases":[],"title":"gVisor: Weak hashing and small seeds let a remote attacker derive a local IP and per-boot identifier","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Weak hashing and small seeds let a remote attacker derive a local IP and per-boot identifier; sandbox fingerprinting","attack_vector":"Unauthenticated network","remediation":"Upgrade runsc","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10026"],"status":"curated"},{"id":"CVE-2024-10603","cve":"CVE-2024-10603","aliases":[],"title":"gVisor: Predictable TCP/UDP source ports and header values enable off-path attacks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Predictable TCP/UDP source ports and header values enable off-path attacks","attack_vector":"Unauthenticated network","remediation":"Upgrade runsc","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10603"],"status":"curated"},{"id":"CVE-2024-36319","cve":"CVE-2024-36319","aliases":[],"title":"AMD Video Decoder Engine Firmware (VCN FW) - debug code left active: MULTI-TENANT ISOLATION: Debug code was shipped active in AMD's Video Core Next firmware, so a maliciously…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Video Decoder Engine Firmware (VCN FW) - debug code left active","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Debug code was shipped active in AMD's Video Core Next firmware, so a maliciously crafted command makes the VCN firmware read and write hardware registers. This is the shipped-debug-hooks failure applied to a GPU IP block: an attacker who can submit VCN commands - which any workload with the GPU device node can - gets arbitrary hardware register access, and hardware registers are how you reconfigure memory apertures, power state and access control on the device. On MI-series parts the VCN block is present whether or not your AI workload uses video decode, so 'we do not do video' is not a mitigation.","attack_vector":"Local, by submitting a crafted command to the video decode engine - reachable from any process holding /dev/dri/renderD*, i.e. an unprivileged tenant container.","remediation":"Fixed in AMD GPU firmware, which on Instinct parts is delivered as a firmware bundle through the ROCm/amdgpu driver package (the PSP loads the signed blobs at driver init) rather than through the server BIOS. Practically: update the AMD GPU driver/firmware package, then **drain the node and reboot** - the firmware is loaded once at driver init, so a reload of the module with no process holding /dev/kfd is the minimum, and a reboot is what you will actually schedule. Some fixes at this layer also require a **GPU VBIOS flash** via AMD's amdvbflash/amdfwtool, which is an offline, per-card operation with real bricking risk - check the AMD bulletin for whether a VBIOS update is called out before assuming a driver package covers it. If your workloads genuinely never touch video decode, consider whether the VCN block can be gated off in your deployment as a stopgap - but verify rather than assume, since the ROCm stack initialises IP blocks it does not use.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36319","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-38805","cve":"CVE-2024-38805","aliases":["GHSA-p7wp-52j7-6r5x"],"title":"EDK II NetworkPkg (IScsiDxe, iSCSI login response processing): A hostile iSCSI target answers the firmware initiator with a malformed login response and gets out-of-bounds…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IScsiDxe, iSCSI login response processing)","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"A hostile iSCSI target answers the firmware initiator with a malformed login response and gets out-of-bounds reads and writes in the pre-OS network stack. Realistic outcome per the upstream advisory is a hung or crashed boot rather than clean code execution, but for a fleet that boots from SAN this is an attacker holding nodes down from the storage side, and the write primitive is a corruption bug that has not been proven unexploitable so much as judged unlikely.","attack_vector":"Whoever controls or can impersonate the iSCSI target the node boots from - a compromised storage appliance, an attacker on the storage VLAN, or a rogue target answering discovery. Unauthenticated from the firmware's point of view, pre-OS.","remediation":"OEM BIOS update; flash + reboot per node. Effective config workaround exists and is cheap: disable the UEFI iSCSI initiator on nodes that do not boot from SAN (a BIOS setting, no flash), and where you do boot from iSCSI, enable mutual CHAP so a rogue target cannot complete the login, and keep the storage network on its own VLAN.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-38805","https://github.com/tianocore/edk2/security/advisories/GHSA-p7wp-52j7-6r5x"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-9594","cve":"CVE-2024-9594","aliases":[],"title":"Kubernetes Image Builder: Default credentials present during the build window for several providers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes Image Builder","year":"2024","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Default credentials present during the build window for several providers","attack_vector":"Attacker present during the image build","remediation":"Rebuild images; isolate the build network","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-23262","cve":"CVE-2025-23262","aliases":[],"title":"ConnectX: Access-control flaw in NIC firmware","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX","year":"2025","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Access-control flaw in NIC firmware","attack_vector":"Tenant with a VF","remediation":"Flash NIC firmware; node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23262","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:H/A:H","cwe":["CWE-863"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37752","cve":"CVE-2025-37752","aliases":[],"title":"Linux kernel (net/sched SFQ): Missing limit validation in sch_sfq - out-of-bounds write","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (net/sched SFQ)","year":"2025","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Missing limit validation in sch_sfq - out-of-bounds write","attack_vector":"Any tenant process in a container with CAP_NET_ADMIN in a userns","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2025-37752"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-15044","cve":"CVE-2026-15044","aliases":[],"title":"TrustyAI Service Operator / NeMo Guardrails: Unverified inter-service channels expose guardrail/orchestrator components to the cluster","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TrustyAI Service Operator / NeMo Guardrails","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Unverified inter-service channels expose guardrail/orchestrator components to the cluster","attack_vector":"Any workload with cluster network access","remediation":"Upgrade the operator; add mTLS/network policy between guardrail services","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=NeMo%20Guardrails"],"status":"curated"},{"id":"CVE-2026-24142","cve":"CVE-2026-24142","aliases":[],"title":"TensorRT-LLM: Code exec via insecure deserialization","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Code exec via insecure deserialization","attack_vector":"Malicious model artifact","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24142","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:C/C:L/I:L/A:L","cwe":["CWE-502"]},{"id":"CVE-2026-24226","cve":"CVE-2026-24226","aliases":[],"title":"TensorRT-LLM: Unsafe external file loading","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Unsafe external file loading -> code exec","attack_vector":"Malicious model reference","remediation":"Bump TensorRT-LLM; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24226","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:R/S:U/C:H/I:H/A:H","cwe":["CWE-829"]},{"id":"CVE-2026-49943","cve":"CVE-2026-49943","aliases":[],"title":"CZ.NIC BIRD Internet Routing Daemon (BGP AS_PATH mask matching): Stack-based buffer overflow in BIRD's AS_PATH mask matching: `as_path_match()` uses a fixed 2049-entry stack…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"CZ.NIC BIRD Internet Routing Daemon (BGP AS_PATH mask matching)","year":"2026","cvss_score":6.3,"severity":"medium","kev":false,"impact":"Stack-based buffer overflow in BIRD's AS_PATH mask matching: `as_path_match()` uses a fixed 2049-entry stack array while the parsed path can exceed it. BIRD is a common choice for route servers, for BGP-to-the-host designs, and inside open networking stacks — including some SONiC and Linux-router-based cluster underlays. A stack overflow in the BGP path-attribute parser is reachable from any peer, and in a route-reflector topology from beyond the direct peer.","attack_vector":"A BGP peer, or anything upstream of one whose AS_PATH propagates, sending a long AS_PATH that hits the mask-matching path. Requires a BGP filter using AS path masks.","remediation":"Upgrade BIRD past 2.19.0 and restart the daemon — package upgrade plus service restart, which briefly drops BGP sessions and reconverges. Interim: apply an inbound AS_PATH length limit on every eBGP session, a live config change and sound policy regardless.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-49943"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2019-3016","cve":"CVE-2019-3016","aliases":[],"title":"Linux KVM - PV TLB shootdown leaks memory between guest processes: MULTI-TENANT ISOLATION: In a KVM guest with paravirtualised TLB enabled, one process in the guest can read…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - PV TLB shootdown leaks memory between guest processes","year":"2019","cvss_score":6.2,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: In a KVM guest with paravirtualised TLB enabled, one process in the guest can read memory belonging to another process in the same guest. The isolation that breaks is inside the VM rather than between VMs - which matters for any operator whose customers run multi-user workloads inside a single VM, and for confidential guests where the tenant assumed process separation held.","attack_vector":"Local, from one process to another inside a KVM guest with PV TLB enabled. The host must be running Linux with KVM.","remediation":"Fixed in the Linux kernel. The fix belongs in the **guest** kernel, so update confidential and tenant VM images, not just hosts. Interim mitigation: disable PV TLB flush in the guest (the kvm.pv_tlb boot option / KVM_FEATURE_PV_TLB_FLUSH), which costs some scheduling efficiency and needs a guest reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-3016"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1093","cve":"CVE-2021-1093","aliases":[],"title":"NVIDIA GPU Display Driver, GPU firmware: An attacker-triggerable assert in GPU firmware aborts harder than it needs to, crashing the system.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, GPU firmware","year":"2021","cvss_score":6.2,"severity":"medium","kev":false,"impact":"An attacker-triggerable assert in GPU firmware aborts harder than it needs to, crashing the system. Firmware-level asserts are worth noting because the failure is below the driver: the recovery is a node reset, not a service restart. Debian and Gentoo shipped it as a security update.","attack_vector":"Any local user or GPU container able to drive the firmware into the asserting path.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1093"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1100","cve":"CVE-2021-1100","aliases":[],"title":"NVIDIA vGPU Manager kernel module (nvidia.ko, host): MULTI-TENANT ISOLATION: the host vGPU kernel module dereferences an unvalidated user-space pointer.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager kernel module (nvidia.ko, host)","year":"2021","cvss_score":6.2,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: the host vGPU kernel module dereferences an unvalidated user-space pointer. Guest-reachable crash of the hypervisor's GPU kernel module, taking down every tenant on the card. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1100"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-42284","cve":"CVE-2022-42284","aliases":[],"title":"DGX servers BMC: Sensitive data exposure from BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Sensitive data exposure from BMC","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+; rotate credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42284","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-312"]},{"id":"CVE-2022-46146","cve":"CVE-2022-46146","aliases":[],"title":"Prometheus (exporter-toolkit): Poisoning the built-in auth cache bypasses basic-auth on exporters","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Prometheus (exporter-toolkit)","year":"2022","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Poisoning the built-in auth cache bypasses basic-auth on exporters","attack_vector":"Local","remediation":"Data-plane: rebuild and roll every exporter - node_exporter/DCGM run on GPU nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-46146"],"status":"curated"},{"id":"CVE-2023-25153","cve":"CVE-2023-25153","aliases":[],"title":"containerd: Unbounded read on OCI image import causes containerd OOM","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2023","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Unbounded read on OCI image import causes containerd OOM; node DoS","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25153"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-8982","cve":"CVE-2024-8982","aliases":[],"title":"OpenLLM: Local file inclusion via the web application","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"OpenLLM","year":"2024","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Local file inclusion via the web application","attack_vector":"Network user of the OpenLLM UI","remediation":"Upgrade past 0.6.10","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-8982"],"status":"curated"},{"id":"CVE-2025-0426","cve":"CVE-2025-0426","aliases":[],"title":"Kubernetes (kubelet): Unauthenticated node DoS through the kubelet checkpoint API filling node disk","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2025","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Unauthenticated node DoS through the kubelet checkpoint API filling node disk","attack_vector":"Any pod on the cluster network reaching the kubelet port","remediation":"Rolling kubelet upgrade with node drain; disable the ContainerCheckpoint feature gate","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-33176","cve":"CVE-2025-33176","aliases":[],"title":"NVIDIA Run:ai: Improper restriction of communication channels lets an attacker on an adjacent network reach privilege…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Run:ai","year":"2025","cvss_score":6.2,"severity":"medium","kev":false,"impact":"Improper restriction of communication channels lets an attacker on an adjacent network reach privilege escalation, data tampering and information disclosure in the Run:ai scheduler. Run:ai is the component deciding which tenant's job lands on which GPU, so control over it is control over the fleet's allocation and quota model.","attack_vector":"Adjacent network, low privileges, user interaction, high complexity. A tenant workload or a compromised pod inside the cluster network.","remediation":"Upgrade Run:ai per bulletin 5719. Cost: a control-plane upgrade - the scheduler restarts, queued jobs pause, running jobs keep their GPUs. Plan it in a low-submission window rather than draining nodes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33176","https://github.com/NVIDIA/product-security/tree/main/2025/5719"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:L/UI:R/S:C/C:L/I:H/A:N","cwe":["CWE-923"]},{"id":"CVE-2026-24271","cve":"CVE-2026-24271","aliases":[],"title":"TensorRT-LLM: DoS via large tensor allocation","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DoS via large tensor allocation","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; add request limits","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24271","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-770"]},{"id":"CVE-2026-41187","cve":"CVE-2026-41187","aliases":[],"title":"Calico: DeleteCollection skips AuthorizeTierOperation, so a tenant can delete tiered NetworkPolicies they cannot…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DeleteCollection skips AuthorizeTierOperation, so a tenant can delete tiered NetworkPolicies they cannot delete individually","attack_vector":"Cluster user with namespace access","remediation":"Rolling Calico apiserver upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41187"],"status":"curated"},{"id":"CVE-2026-47470","cve":"CVE-2026-47470","aliases":[],"title":"TensorRT-LLM: DoS / memory corruption (insufficient tensor validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DoS / memory corruption (insufficient tensor validation)","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47470","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"]},{"id":"CVE-2026-47475","cve":"CVE-2026-47475","aliases":[],"title":"TensorRT-LLM: DoS via assertion failure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":6.2,"severity":"medium","kev":false,"impact":"DoS via assertion failure","attack_vector":"Any inference client","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47475","https://github.com/NVIDIA/product-security/tree/main/2026/5840"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-617"]},{"id":"CVE-2020-15157","cve":"CVE-2020-15157","aliases":[],"title":"containerd: \"ContainerDrip\": registry credentials leaked to an attacker-controlled URL referenced in an image manifest","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2020","cvss_score":6.1,"severity":"medium","kev":false,"impact":"\"ContainerDrip\": registry credentials leaked to an attacker-controlled URL referenced in an image manifest","attack_vector":"Malicious image pulled from an untrusted registry","remediation":"Upgrade containerd; rotate any registry pull credentials that may have leaked","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15157"],"status":"curated"},{"id":"CVE-2021-1094","cve":"CVE-2021-1094","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Out-of-bounds array access in the escape handler on Windows and Linux, giving disclosure or a crash from…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Out-of-bounds array access in the escape handler on Windows and Linux, giving disclosure or a crash from unprivileged local code.","attack_vector":"Any local user or GPU container with device access.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1094"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-22810","cve":"CVE-2021-22810","aliases":["SEVD-2021-313-03"],"title":"APC Network Management Card 2 (AP9630/AP9631/AP9635) in Smart-UPS, Symmetra and Galaxy 3500: Stored/reflected cross-site scripting in the NMC2 policy-file pages. On its own it is a browser bug","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC Network Management Card 2 (AP9630/AP9631/AP9635) in Smart-UPS, Symmetra and Galaxy 3500","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Stored/reflected cross-site scripting in the NMC2 policy-file pages. On its own it is a browser bug; in context it is a route to hijack a facility engineer's authenticated session on the card that controls UPS behaviour. An attacker with an NMC session can change shutdown policies, thresholds and outlet-group behaviour - which is a path to a power event, not just a defacement.","attack_vector":"Requires tricking an already-privileged NMC user into clicking a crafted URL. Realistic in a colo where facility staff routinely click links in tickets.","remediation":"Firmware update to NMC2 AOS v6.9.6 or later (SEVD-2021-313-03 covers the whole CVE-2021-22810 through -22815 batch, so treat it as one campaign). Non-disruptive flash. Enforce that NMC admin sessions are only opened from a dedicated management workstation.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2021-313-03"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-28509","cve":"CVE-2021-28509","aliases":[],"title":"Arista EOS (TerminAttr / OpenConfig telemetry transport): TENANT ISOLATION: the streaming-telemetry agent can leak MACsec keys over the telemetry transport. Whoever…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (TerminAttr / OpenConfig telemetry transport)","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: the streaming-telemetry agent can leak MACsec keys over the telemetry transport. Whoever consumes your telemetry stream — often a monitoring platform with far weaker access control than the switches themselves — ends up holding the keys that protect inter-site and inter-pod links. From there an attacker decrypts traffic for every tenant crossing those links.","attack_vector":"An attacker with access to the telemetry stream or to the collector storing it. That is usually a much softer target than the switch.","remediation":"EOS/TerminAttr upgrade plus agent restart. Then rotate every MACsec key that could have been exposed — a fabric-wide key rotation is disruptive and is the real cost here, not the upgrade. Treat telemetry collectors as secret-bearing systems going forward.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28509"],"status":"curated"},{"id":"CVE-2021-38961","cve":"CVE-2021-38961","aliases":["IBM X-Force 212049"],"title":"IBM OpenBMC OP910 web UI (phosphor-webui lineage): Stored/reflected script injection in the BMC web interface. The victim is your own operator: an admin opens…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC OP910 web UI (phosphor-webui lineage)","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Stored/reflected script injection in the BMC web interface. The victim is your own operator: an admin opens the BMC console and the injected script runs with their authenticated session, which is a session that can power-cycle nodes, mount virtual media and push firmware. On a GPU fleet the practical scenario is an attacker who has read-only or low-privilege access to one BMC planting the payload and waiting for an administrator to visit, converting a foothold into administrative control without ever cracking a password.","attack_vector":"Requires getting attacker-controlled content into a field the BMC web UI renders, plus an administrator subsequently loading that page. Network access to the BMC web interface.","remediation":"Fixed in later OP910 firmware - per-node system firmware update, maintenance window. Cheap compensating control: do not browse BMC web UIs from the same browser profile you use for anything else, and prefer Redfish API calls over the web UI for routine operations. Fleet-scale automation against Redfish rather than humans clicking through per-node web UIs removes the victim this bug needs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-38961","https://www.ibm.com/support/pages/node/6536720","https://exchange.xforce.ibmcloud.com/vulnerabilities/212049"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-46759","cve":"CVE-2021-46759","aliases":[],"title":"AMD TEE / ASP bootloader syscall input validation: Insufficient validation of syscall inputs in the AMD trusted execution environment lets an attacker who…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD TEE / ASP bootloader syscall input validation","year":"2021","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Insufficient validation of syscall inputs in the AMD trusted execution environment lets an attacker who controls a user application running under the ASP bootloader read back ASP bootloader memory - disclosing firmware internals and, more usefully to an attacker, the layout and secrets needed to build a reliable exploit against the secure processor.","attack_vector":"Local plus physical access, and control of a Uapp running under the bootloader. High bar; realistic for an attacker with hands on the hardware (supply chain, colocation insider, returned hardware).","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Physical-access requirement means datacenter physical controls and tamper-evident handling are a genuine compensating control here.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46759","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-21123","cve":"CVE-2022-21123","aliases":[],"title":"Intel CPU (MMIO Stale Data / SBDR): Incomplete cleanup of multi-core shared buffers - stale data read across domains","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (MMIO Stale Data / SBDR)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Incomplete cleanup of multi-core shared buffers - stale data read across domains","attack_vector":"Tenant VM guest; any tenant process in a container","remediation":"Microcode + kernel mitigation + reboot; consider disabling SMT for hard-isolation tenants","references":["https://access.redhat.com/security/cve/CVE-2022-21123"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-21813","cve":"CVE-2022-21813","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged local user gets limited write access to memory the driver treats as protected, which is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An unprivileged local user gets limited write access to memory the driver treats as protected, which is enough to crash the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5312. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21813","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-280"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-21814","cve":"CVE-2022-21814","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged local user gets limited write access to protected memory through the kernel driver package…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An unprivileged local user gets limited write access to protected memory through the kernel driver package, ending in a node-wide GPU denial of service. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5312. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21814","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-280"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-28186","cve":"CVE-2022-28186","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper input validation in the DxgkDdiEscape handler lets a local user crash the node or tamper with driver…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Improper input validation in the DxgkDdiEscape handler lets a local user crash the node or tamper with driver state. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28186","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31616","cve":"CVE-2022-31616","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An out-of-bounds read through DxgkDdiEscape yields a crash or kernel information disclosure. Only matters to…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An out-of-bounds read through DxgkDdiEscape yields a crash or kernel information disclosure. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5383. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31616","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-0186","cve":"CVE-2023-0186","aliases":[],"title":"GPU Display Driver: DoS (GPU firmware buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS (GPU firmware buffer overflow)","attack_vector":"Any tenant with a container","remediation":"Driver + GPU firmware bundle upgrade; reboot node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0186","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0187","cve":"CVE-2023-0187","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0187","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0199","cve":"CVE-2023-0199","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): An out-of-bounds write in the kernel mode handler gives a local user denial of service and data tampering…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"An out-of-bounds write in the kernel mode handler gives a local user denial of service and data tampering across the driver boundary. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5452. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0199","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-787"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-22655","cve":"CVE-2023-22655","aliases":[],"title":"Intel 3rd/4th Gen Xeon with SGX or TDX (protection mechanism failure): MULTI-TENANT ISOLATION: A protection mechanism in 3rd and 4th generation Xeon fails when SGX or TDX is in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel 3rd/4th Gen Xeon with SGX or TDX (protection mechanism failure)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A protection mechanism in 3rd and 4th generation Xeon fails when SGX or TDX is in use, letting a privileged local user escalate. Affects the exact Xeon generations most AI datacenter hosts were built on between 2021 and 2024.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update plus TCB recovery. Late-loadable microcode, reboot, then re-attest enclaves and trust domains.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22655","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00960.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2023-28642","cve":"CVE-2023-28642","aliases":[],"title":"runc: AppArmor bypass when /proc inside the container is symlinked with a specific mount config","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"AppArmor bypass when /proc inside the container is symlinked with a specific mount config","attack_vector":"Malicious image or tenant-controlled pod spec","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-28642"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-31012","cve":"CVE-2023-31012","aliases":[],"title":"DGX H100 BMC (REST): DoS / data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS / data tampering","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-20"]},{"id":"CVE-2023-31013","cve":"CVE-2023-31013","aliases":[],"title":"DGX H100 BMC (REST): DoS / data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS / data tampering","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-20"]},{"id":"CVE-2023-31020","cve":"CVE-2023-31020","aliases":[],"title":"GPU Display Driver (Windows): DoS / data tampering","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS / data tampering","attack_vector":"Local user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-284"]},{"id":"CVE-2023-40546","cve":"CVE-2023-40546","aliases":["shim 15.8 batch"],"title":"shim (mok.c mirror_one_esl): NULL pointer dereference while printing an error message stops the node from booting. On a fleet this is a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (mok.c mirror_one_esl)","year":"2023","cvss_score":6.1,"severity":"medium","kev":false,"impact":"NULL pointer dereference while printing an error message stops the node from booting. On a fleet this is a denial of service you cannot fix over SSH - it needs console or BMC access per affected node, which is expensive at scale.","attack_vector":"Requires attacker-influenced MOK/ESL data on the node.","remediation":"shim package update + reboot. Availability-only, so it can ride a normal maintenance window rather than an emergency one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40546","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0075","cve":"CVE-2024-0075","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0075","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0115","cve":"CVE-2024-0115","aliases":[],"title":"NVIDIA CV-CUDA: A long-running CV-CUDA Python process consumes resources without bound, ending in denial of service and data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CV-CUDA","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A long-running CV-CUDA Python process consumes resources without bound, ending in denial of service and data loss. On a shared inference node this means one tenant's preprocessing job can starve the box.","attack_vector":"Local, low privileges - a user able to submit work through the CV-CUDA Python API.","remediation":"Update CV-CUDA per bulletin 5560 and rebuild affected images. Cost: package update and job restart; no driver or firmware change. Consider cgroup memory limits on preprocessing containers as a standing control.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0115","https://github.com/NVIDIA/product-security/tree/main/2024/5560"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:H","cwe":["CWE-400"]},{"id":"CVE-2024-10086","cve":"CVE-2024-10086","aliases":[],"title":"HashiCorp Consul: Missing Content-Type header lets user input be reinterpreted","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Consul","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Missing Content-Type header lets user input be reinterpreted -> reflected XSS on the Consul UI","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; keep the Consul UI behind the ops VPN","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10086"],"status":"curated"},{"id":"CVE-2024-10649","cve":"CVE-2024-10649","aliases":[],"title":"Weights & Biases OpenUI: Unauthenticated endpoints allow file upload and download","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Weights & Biases OpenUI","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Unauthenticated endpoints allow file upload and download","attack_vector":"Unauthenticated network","remediation":"Upgrade; only affects the OpenUI project, not the core W&B SDK","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-10649"],"status":"curated"},{"id":"CVE-2024-25630","cve":"CVE-2024-25630","aliases":[],"title":"Cilium: WireGuard transparent encryption not applied to some pod traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"WireGuard transparent encryption not applied to some pod traffic","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25630"],"status":"curated"},{"id":"CVE-2024-25631","cve":"CVE-2024-25631","aliases":[],"title":"Cilium: With an external kvstore and WireGuard, pod-to-pod traffic is unencrypted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"With an external kvstore and WireGuard, pod-to-pod traffic is unencrypted","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-25631"],"status":"curated"},{"id":"CVE-2024-28249","cve":"CVE-2024-28249","aliases":[],"title":"Cilium: IPsec-eligible traffic matching L7 policy is sent unencrypted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"IPsec-eligible traffic matching L7 policy is sent unencrypted","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28249"],"status":"curated"},{"id":"CVE-2024-28250","cve":"CVE-2024-28250","aliases":[],"title":"Cilium: WireGuard-eligible traffic matching L7 policy is sent unencrypted between nodes","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"WireGuard-eligible traffic matching L7 policy is sent unencrypted between nodes","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-28250"],"status":"curated"},{"id":"CVE-2024-34923","cve":"CVE-2024-34923","aliases":[],"title":"Avocent DSR2030 / SVIP1020 KVM-over-IP appliance: A reflected XSS in the appliance's web interface lets an attacker who can get an operator to click a crafted…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Avocent DSR2030 / SVIP1020 KVM-over-IP appliance","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A reflected XSS in the appliance's web interface lets an attacker who can get an operator to click a crafted link run JavaScript in that operator's browser session — enough to steal their session cookie and act as them on the KVM appliance.","attack_vector":"Requires social engineering: the victim (an operator with legitimate access to the KVM appliance) has to click a link the attacker controls while authenticated to the device.","remediation":"Software upgrade — DSR2030 to firmware 03.07.01.23 or later, SVIP1020 to 01.07.00.00 or later. Standard firmware flash per unit; no serial/KVM downtime beyond the reboot itself.","references":["https://ka1ne1.github.io/avocent_xss.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-48869","cve":"CVE-2024-48869","aliases":[],"title":"Intel Xeon 6 E-core with TDX or SGX: MULTI-TENANT ISOLATION: Improper restriction of software interfaces to hardware features on Xeon 6 E-core…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Xeon 6 E-core with TDX or SGX","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper restriction of software interfaces to hardware features on Xeon 6 E-core parts when TDX or SGX is in use. Reaches both confidential-compute technologies on the same silicon, so a single platform update covers both boundaries.","attack_vector":"Local access on an affected Xeon 6 platform with TDX or SGX enabled.","remediation":"OEM platform firmware/BIOS update plus TCB recovery for whichever technology you use. Drain, reboot, re-attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48869","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01268.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-50302","cve":"CVE-2024-50302","aliases":[],"title":"Linux kernel (HID): Uninitialised HID report buffer leaks kernel memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (HID)","year":"2024","cvss_score":6.1,"severity":"medium","kev":true,"impact":"Uninitialised HID report buffer leaks kernel memory [KEV]","attack_vector":"Local user with USB device access","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2024-50302"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-5321","cve":"CVE-2024-5321","aliases":[],"title":"Kubernetes (kubelet): Incorrect permissions on Windows container log directories allow privilege escalation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2024","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Incorrect permissions on Windows container log directories allow privilege escalation","attack_vector":"Any tenant workload on a Windows node","remediation":"Rolling kubelet upgrade; Windows node drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-23528","cve":"CVE-2026-23528","aliases":[],"title":"Dask distributed (+ Jupyter proxy): Exposure when Dask, JupyterLab and jupyter-server-proxy are combined","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Dask distributed (+ Jupyter proxy)","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Exposure when Dask, JupyterLab and jupyter-server-proxy are combined","attack_vector":"Notebook user / network attacker","remediation":"Upgrade to 2026.1.0+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23528"],"status":"curated"},{"id":"CVE-2026-26963","cve":"CVE-2026-26963","aliases":[],"title":"Cilium: With native routing plus WireGuard node encryption, traffic from pods on other nodes is wrongly permitted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"With native routing plus WireGuard node encryption, traffic from pods on other nodes is wrongly permitted","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-26963"],"status":"curated"},{"id":"CVE-2026-29777","cve":"CVE-2026-29777","aliases":[],"title":"Traefik: A tenant with HTTPRoute write access injects backtick-delimited rule tokens into Traefik's router rule…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"A tenant with HTTPRoute write access injects backtick-delimited rule tokens into Traefik's router rule language; cross-tenant route hijack","attack_vector":"Cluster user with namespace access","remediation":"Rolling Traefik upgrade to 3.6.10+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-29777"],"status":"curated"},{"id":"CVE-2026-41568","cve":"CVE-2026-41568","aliases":[],"title":"Docker / moby: Companion `docker cp` mount-setup race","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2026","cvss_score":6.1,"severity":"medium","kev":false,"impact":"Companion `docker cp` mount-setup race","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine to 29.5.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41568"],"status":"curated"},{"id":"CVE-2017-5703","cve":"CVE-2017-5703","aliases":["INTEL-SA-00087"],"title":"SPI flash configuration (flash descriptor / protected range registers) across multiple Intel platforms: Misconfiguration of SPI flash protection lets a local attacker change how the SPI flash…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SPI flash configuration (flash descriptor / protected range registers) across multiple Intel platforms","year":"2017","cvss_score":6,"severity":"medium","kev":false,"impact":"Misconfiguration of SPI flash protection lets a local attacker change how the SPI flash behaves, up to and including bricking the node. In a GPU fleet the operator-facing outcome is a hard denial of service that no reimage fixes: the node will not POST and needs a physical flash recovery (external programmer or OEM RMA), so it is an unplanned rack visit and a node out of revenue for days. Where write protection is incomplete rather than merely unstable, the same weakness is the standard route to a persistent BIOS implant that survives every reimage and every tenant handoff.","attack_vector":"Local privileged code on the host writing to the SPI controller's configuration and protected-range registers. Reachable by any tenant with root on a bare-metal node.","remediation":"BIOS/platform firmware update from the OEM that sets the flash descriptor and protected-range/BIOS-lock registers correctly - HPE, Dell, Supermicro and Lenovo all shipped these; reboot and drain required. Independently of the patch, audit the flash-protection state on your actual fleet (BIOSWE/BLE/SMM_BWP/PRx and descriptor lock) with a tool like CHIPSEC as part of node acceptance and node reclaim, because these bits are set by the OEM's BIOS build and vary by SKU and by BIOS version even within one OEM.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5703","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00087.html","https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf03867en_us"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-43784","cve":"CVE-2021-43784","aliases":[],"title":"runc: Netlink bytemsg length integer overflow in libcontainer allows config injection / partial escape","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2021","cvss_score":6,"severity":"medium","kev":false,"impact":"Netlink bytemsg length integer overflow in libcontainer allows config injection / partial escape","attack_vector":"Any tenant workload with control over container config","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-43784"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-21233","cve":"CVE-2022-21233","aliases":[],"title":"Intel CPU (AEPIC Leak): Stale data read from the legacy xAPIC MMIO page - leaks SGX enclave and cross-domain data","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (AEPIC Leak)","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"Stale data read from the legacy xAPIC MMIO page - leaks SGX enclave and cross-domain data","attack_vector":"Local user; tenant VM guest","remediation":"Microcode + reboot. Kills naive SGX-based confidential-compute claims on affected parts","references":["https://access.redhat.com/security/cve/CVE-2022-21233"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-36382","cve":"CVE-2022-36382","aliases":[],"title":"Intel Ethernet E810 Series and Ethernet 700 Series firmware: Out-of-bounds write in firmware across both the E810 line and the older 700 Series (X710/XL710/XXV710)…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet E810 Series and Ethernet 700 Series firmware","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"Out-of-bounds write in firmware across both the E810 line and the older 700 Series (X710/XL710/XXV710), triggerable by a privileged host user. Notable because it spans two adapter generations — if you have a mixed fleet, the version target differs per family (E810 before 1.7.0.8, 700 Series before 9.101) and a single blanket firmware policy will miss half of it.","attack_vector":"Privileged local user on the host.","remediation":"Flash E810 to 1.7.0.8+ and 700-Series adapters to 9.101+. Cold power cycle on both. Because the two families need different images, build the NVM update into your provisioning pipeline keyed on device ID rather than doing it by hand.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36382"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-38090","cve":"CVE-2022-38090","aliases":[],"title":"Intel processors with SGX (shared resource isolation): MULTI-TENANT ISOLATION: Improper isolation of shared microarchitectural resources lets a privileged user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors with SGX (shared resource isolation)","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper isolation of shared microarchitectural resources lets a privileged user extract information from SGX enclaves. Another TCB-recovery event for anyone selling enclave-backed confidentiality.","attack_vector":"Privileged local access on the host.","remediation":"Microcode update and re-attestation. Late-loadable microcode plus reboot; no OEM BIOS strictly required for the microcode component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-38090","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00767.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-42285","cve":"CVE-2022-42285","aliases":[],"title":"NVIDIA DGX A100 - SBIOS / SMM firmware: A privileged user can disable SPI flash write protection during the Pre-EFI Initialization phase, which…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX A100 - SBIOS / SMM firmware","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"A privileged user can disable SPI flash write protection during the Pre-EFI Initialization phase, which removes the platform's own guard against firmware overwrite. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5435. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42285","https://github.com/NVIDIA/product-security/tree/main/2022/5435"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-1231"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42286","cve":"CVE-2022-42286","aliases":[],"title":"DGX-2 SBIOS: OOB write in BIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 SBIOS","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"OOB write in BIOS","attack_vector":"Local operator","remediation":"Flash SBIOS out-of-band; node power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42286","https://github.com/NVIDIA/product-security/tree/main/2023/5449"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-42287","cve":"CVE-2022-42287","aliases":[],"title":"DGX-2 BMC: Path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX-2 BMC","year":"2022","cvss_score":6,"severity":"medium","kev":false,"impact":"Path traversal","attack_vector":"Network-adjacent authenticated","remediation":"Flash DGX-2 BMC firmware","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42287","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:N","cwe":["CWE-22"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20588","cve":"CVE-2023-20588","aliases":[],"title":"AMD CPU (DIV0): Division-by-zero leaves stale quotient data readable across contexts - confidentiality loss on Zen 1","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (DIV0)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Division-by-zero leaves stale quotient data readable across contexts - confidentiality loss on Zen 1","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation + reboot; low perf cost","references":["https://access.redhat.com/security/cve/CVE-2023-20588"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-25509","cve":"CVE-2023-25509","aliases":[],"title":"NVIDIA DGX-1 - SBIOS / SMM firmware: A flaw in the Bds phase reaches firmware code execution and privilege escalation. This is firmware-level…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX-1 - SBIOS / SMM firmware","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"A flaw in the Bds phase reaches firmware code execution and privilege escalation. This is firmware-level persistence: it survives OS reinstall, image re-flash and tenant handoff, and it is invisible to anything running above it. On a bare-metal GPU rental business it is the difference between wiping a node between tenants and not actually being able to.","attack_vector":"Local and already privileged - host root, or code that has reached the platform firmware/SMM path. It is not a first foothold; it is what turns a one-time root compromise into something you cannot remediate by reimaging.","remediation":"Flash the fixed SBIOS from bulletin 5458. Cost: not live-patchable. Full node drain, host power cycle, and on DGX the SBIOS ships inside a firmware bundle alongside BMC and CPLD components, so budget 30-60 minutes of node downtime plus a post-flash health check. Firmware rollback protection means you cannot cleanly revert - stage on one node before the fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25509","https://github.com/NVIDIA/product-security/tree/main/2023/5458"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-30456","cve":"CVE-2023-30456","aliases":[],"title":"KVM (nested VMX): Missing CR0/CR4 consistency checks in nVMX - L2 guest can break nested-virt assumptions / crash host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM (nested VMX)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Missing CR0/CR4 consistency checks in nVMX - L2 guest can break nested-virt assumptions / crash host","attack_vector":"Tenant VM guest running nested virtualisation","remediation":"Kernel patch + reboot. Cheaper interim control: disable nested virtualisation for tenant VMs, which most GPU tenants do not need","references":["https://access.redhat.com/security/cve/CVE-2023-30456"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-31026","cve":"CVE-2023-31026","aliases":[],"title":"vGPU software (Virtual GPU Manager): Host-side DoS (null deref in vGPU Manager)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU software (Virtual GPU Manager)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Host-side DoS (null deref in vGPU Manager)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager on hypervisor; migrate guest VMs then reboot host","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-31355","cve":"CVE-2023-31355","aliases":["decommissioned guest memory disclosure","UMC seed reuse"],"title":"AMD SEV-SNP firmware, guest teardown / UMC key seed handling: TENANT HANDOFF FAILURE. A malicious hypervisor can overwrite a guest's UMC seed such that memory belonging to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP firmware, guest teardown / UMC key seed handling","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"TENANT HANDOFF FAILURE. A malicious hypervisor can overwrite a guest's UMC seed such that memory belonging to an already-decommissioned confidential guest becomes readable. In a rented-GPU business the slot a customer just released is immediately resold; this says the previous tenant's plaintext - checkpoints, weights, prompts, keys still resident in DRAM - can be recovered after their VM is gone. It breaks the one property you cannot buy back with an apology, and it does so on a path your own automation exercises thousands of times a day.","attack_vector":"Malicious or compromised hypervisor / host root, acting after a confidential guest terminates. No access to the victim tenant needed at all - only control of the host they used to be on.","remediation":"Same package as CVE-2024-21980 (AMD-SB-3011): hot-loadable SEV firmware 1.37.14 hex (Milan) / 1.37.24 hex (Genoa) with no reboot, or Platform Initialization firmware MilanPI 1.0.0.D / GenoaPI 1.0.0.C via OEM BIOS with a reboot. Prioritize this one over the rest of the batch. Until it is fixed, do not treat guest teardown as a memory-sanitization boundary - force an explicit scrub or a full host reboot between confidential tenants. The fix bumps TCB[SNP], so re-baseline attestation policies.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3011.html","https://nvd.nist.gov/vuln/detail/CVE-2023-31355"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-34342","cve":"CVE-2023-34342","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (IPMI handler): Arbitrary file upload and download through the BMC's IPMI handler. Download gives the attacker the BMC's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (IPMI handler)","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"Arbitrary file upload and download through the BMC's IPMI handler. Download gives the attacker the BMC's stored secrets and configuration; upload gives them a way to drop a payload onto the controller's filesystem and, depending on where it lands, get it executed - which is how a credentialed foothold becomes a persistent BMC implant. Availability damage is also on the table: writing over the wrong file bricks the controller.","attack_vector":"Local access to the BMC with high privileges per AMI's vector - i.e. an attacker who already holds a BMC admin credential or has landed on the controller. Its role in a real chain is post-exploitation persistence, not initial access.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. Config-only reduction: disable IPMI-over-LAN so the handler is not reachable from the network at all and drive management through Redfish, accepting that this breaks ipmitool-based provisioning and monitoring tooling. Also worth doing regardless: alert on any BMC firmware or filesystem change, because this class of bug is invisible from the host OS.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34342"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-47165","cve":"CVE-2023-47165","aliases":[],"title":"Intel Data Center GPU Max Series 1100 / 1550: An improper conditions check lets a privileged local user take the accelerator out of service. Low ceiling…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center GPU Max Series 1100 / 1550","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"An improper conditions check lets a privileged local user take the accelerator out of service. Low ceiling - denial of service only - but it is one of the very few CVEs that exist against Intel's Ponte Vecchio datacenter GPUs at all, which is itself worth knowing if you are evaluating them.","attack_vector":"Local, privileged. Host root on the node holding the GPU.","remediation":"Apply the Intel firmware/driver update for the Max Series. Cost: a GPU firmware update requires a drain and reboot; the driver alone needs a module reload.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2023-47165"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-47855","cve":"CVE-2023-47855","aliases":[],"title":"Intel TDX module: MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2023","cvss_score":6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: a second input-validation gap in the same module version range.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-47855","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01036.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-20399","cve":"CVE-2024-20399","aliases":[],"title":"Cisco NX-OS CLI: **[KEV]** Command injection giving root on the switch's underlying OS from an admin CLI session","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS CLI","year":"2024","cvss_score":6,"severity":"medium","kev":true,"impact":"**[KEV]** Command injection giving root on the switch's underlying OS from an admin CLI session; exploited in the wild by the Velvet Ant group, who used it to install persistent malware on the switch","attack_vector":"Network, authenticated administrator","remediation":"NX-OS upgrade with fabric failover — one leaf at a time, relying on the fabric's redundancy; a spine upgrade on a rail-optimised GPU fabric costs measurable job throughput","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-20399"],"status":"curated"},{"id":"CVE-2024-21850","cve":"CVE-2024-21850","aliases":[],"title":"Intel TDX SEAM loader (Seamldr): MULTI-TENANT ISOLATION: Sensitive information is not cleared before a resource is reused in the SEAM loader…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX SEAM loader (Seamldr)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Sensitive information is not cleared before a resource is reused in the SEAM loader - the component that loads and measures the TDX module itself. Anything wrong at the SEAM loader layer is below the TDX module in the trust stack, so it undermines the measurement every TD attestation ultimately chains to.","attack_vector":"Privileged host user.","remediation":"Update the TDX SEAM loader (Seamldr) to 1.5.02.00 or later alongside the TDX module. Loaded at boot, so drain all trust domains and reboot the node. Re-attest afterwards; the SEAM loader version feeds the attestation chain. No OEM BIOS dependency for the loader itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21850","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01076.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-21961","cve":"CVE-2024-21961","aliases":["PCIe link guest-to-host DoS"],"title":"AMD PCIe link handling (memory buffer bounds): A guest VM can drive the PCIe link into an out-of-bounds condition and deny service to the entire host. On an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PCIe link handling (memory buffer bounds)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"A guest VM can drive the PCIe link into an out-of-bounds condition and deny service to the entire host. On an AI node the PCIe fabric is the load-bearing structure - GPUs, NVMe scratch, RDMA NICs all hang off it - so a link-level fault does not degrade one tenant, it takes the box down and kills every job on it. This is a guest-to-host availability break reachable from a normal VM, which is a materially different risk class from the ring 0 firmware issues elsewhere in this set.","attack_vector":"Attacker with access to a guest virtual machine - an ordinary paying tenant. Network-adjacent attack vector per AMD's scoring, low privilege required.","remediation":"Firmware update per AMD-SB-4013 from the OEM; BIOS flash and reboot. In the meantime the practical control is blast-radius management rather than prevention: do not co-locate high-value long-running training jobs with untrusted short-lived tenants on the same PCIe complex, and make sure checkpointing intervals assume the node can vanish. Note this bulletin targets client and embedded platform audits, so confirm applicability to your specific EPYC or Instinct SKUs with your vendor before planning a fleet-wide flash.","references":["https://www.amd.com/en/resources/product-security/bulletin/AMD-SB-4013.html","https://nvd.nist.gov/vuln/detail/CVE-2024-21961"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21978","cve":"CVE-2024-21978","aliases":[],"title":"AMD SEV-SNP firmware - input validation: MULTI-TENANT ISOLATION: Improper input validation in SEV-SNP lets a malicious hypervisor read or overwrite…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP firmware - input validation","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper input validation in SEV-SNP lets a malicious hypervisor read or overwrite guest memory. That is the whole point of SEV-SNP defeated in one line: the host, which SNP exists to exclude, gets both read and write access to the confidential guest's pages.","attack_vector":"Malicious or compromised hypervisor. The guest need do nothing.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21978","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-24595","cve":"CVE-2024-24595","aliases":[],"title":"ClearML: Passwords stored in plaintext in MongoDB","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"ClearML","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Passwords stored in plaintext in MongoDB","attack_vector":"Anyone who compromises the ClearML server","remediation":"Upgrade; rotate all credentials after any exposure","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-24595"],"status":"curated"},{"id":"CVE-2024-36346","cve":"CVE-2024-36346","aliases":[],"title":"AMD Power Management Firmware (PMFW) - guest VM input validation causing GPU reset: MULTI-TENANT ISOLATION: Improper input validation in AMD's GPU Power Management Firmware lets a **guest VM**…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Power Management Firmware (PMFW) - guest VM input validation causing GPU reset","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper input validation in AMD's GPU Power Management Firmware lets a **guest VM** send arbitrary input data that forces a GPU reset. On a virtualised GPU host this is a tenant taking the accelerator out from under everyone sharing it: a GPU reset kills in-flight work on the whole device, so one tenant's malformed PMFW message destroys other tenants' training progress since their last checkpoint. Cheap to trigger, expensive to absorb, and it does not require the attacker to escape their VM at all.","attack_vector":"From inside a guest VM with GPU access - SR-IOV virtual function or passthrough. No host privilege and no escape needed; the guest simply talks to the power management firmware through the interface it is legitimately given.","remediation":"Fixed in AMD GPU firmware, which on Instinct parts is delivered as a firmware bundle through the ROCm/amdgpu driver package (the PSP loads the signed blobs at driver init) rather than through the server BIOS. Practically: update the AMD GPU driver/firmware package, then **drain the node and reboot** - the firmware is loaded once at driver init, so a reload of the module with no process holding /dev/kfd is the minimum, and a reboot is what you will actually schedule. Some fixes at this layer also require a **GPU VBIOS flash** via AMD's amdvbflash/amdfwtool, which is an offline, per-card operation with real bricking risk - check the AMD bulletin for whether a VBIOS update is called out before assuming a driver package covers it. Prioritise this on any GPU virtualisation deployment with untrusted tenants - the attacker prerequisite is just 'has a GPU assigned'. Interim mitigation is thin: you cannot easily filter PMFW messages from a VF, so the practical stopgap is not co-tenanting untrusted guests on a shared physical GPU until the firmware is updated.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36346","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-39283","cve":"CVE-2024-39283","aliases":[],"title":"Intel TDX module: MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: incomplete filtering of special elements, reachable by an authenticated user.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39283","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01010.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-45779","cve":"CVE-2024-45779","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (BFS filesystem parser): Integer overflow producing a heap out-of-bounds read in the BeFS parser - leaks bootloader memory, useful for…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (BFS filesystem parser)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Integer overflow producing a heap out-of-bounds read in the BeFS parser - leaks bootloader memory, useful for defeating any layout randomisation before pairing with a write primitive.","attack_vector":"Attacker-supplied BFS image.","remediation":"grub2 package update + reboot; or strip unused filesystem modules from the GRUB build.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45779","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50264","cve":"CVE-2024-50264","aliases":[],"title":"Linux kernel (vsock/virtio): Dangling pointer in vsk","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (vsock/virtio)","year":"2024","cvss_score":6,"severity":"medium","kev":false,"impact":"Dangling pointer in vsk->trans during AF_VSOCK connect - use-after-free, local root; vsock is the host/guest channel on virtualised nodes","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Livepatchable; otherwise drain + reboot. Blacklist vsock modules on nodes that do not use them","references":["https://access.redhat.com/security/cve/CVE-2024-50264"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-0033","cve":"CVE-2025-0033","aliases":[],"title":"AMD SEV-SNP - RMP write access during SNP initialization: MULTI-TENANT ISOLATION: There is a window during SEV-SNP initialization in which an admin-privileged attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - RMP write access during SNP initialization","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: There is a window during SEV-SNP initialization in which an admin-privileged attacker can write to the reverse-map table itself. Corrupting the RMP at init time means the ownership map that governs every subsequent guest page assignment starts out wrong - the attacker can arrange for guest pages to be host-writable from the moment the platform comes up, and no later check catches it because the check is the RMP.","attack_vector":"Local, admin-privileged, and specifically during SNP platform initialization - so an attacker who controls the host boot sequence or the SNP init path.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Since the exposure is at SNP init, the practical control is boot integrity: measured boot on the host, and refusing to admit confidential tenants onto nodes whose boot chain you cannot attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0033","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-20067","cve":"CVE-2025-20067","aliases":[],"title":"Intel CSME / SPS firmware (timing side channel): An observable timing discrepancy in CSME/SPS firmware allows a privileged local user to infer information…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CSME / SPS firmware (timing side channel)","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"An observable timing discrepancy in CSME/SPS firmware allows a privileged local user to infer information from the management engine. Low direct impact; useful as a step toward key or state recovery from a component that is supposed to be opaque to the host.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20067","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01280.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-24296","cve":"CVE-2025-24296","aliases":[],"title":"Intel E810 Ethernet controller firmware: Improper input validation in E810 firmware lets a privileged local user deny service on the adapter. On a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Improper input validation in E810 firmware lets a privileged local user deny service on the adapter. On a node whose NIC carries collective traffic, denying the NIC denies the job.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target E810 firmware 4.6 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24296","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01257.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-24851","cve":"CVE-2025-24851","aliases":["INTEL-SA-01171"],"title":"Intel Ethernet Controller E810 (100GbE) firmware: Uncaught exception in 100GbE E810 firmware, reachable from privileged host software, causing denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 (100GbE) firmware","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Uncaught exception in 100GbE E810 firmware, reachable from privileged host software, causing denial of service. Ships in the same advisory as the out-of-bounds write, so a single NVM update fixes both — worth listing separately so operators tracking by CVE do not think they are done after one.","attack_vector":"Privileged local software on the host.","remediation":"Same NVM update as CVE-2025-27243 — E810 firmware cvl fw 1.7.8.x or later, cold power cycle. One flash, two CVEs.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01171.html","https://nvd.nist.gov/vuln/detail/CVE-2025-24851"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-27243","cve":"CVE-2025-27243","aliases":["INTEL-SA-01171"],"title":"Intel Ethernet Controller E810 firmware: Out-of-bounds write inside E810 firmware, reachable from a privileged Ring-0 software adversary on the host…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Ethernet Controller E810 firmware","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Out-of-bounds write inside E810 firmware, reachable from a privileged Ring-0 software adversary on the host, causing denial of service. The important framing for an operator is that this is a **write** primitive into NIC firmware from the host: on a bare-metal rental where the tenant has kernel privilege, the boundary between 'a tenant had root on the node' and 'the NIC's firmware state was modified' is exactly what this class of bug erodes. The published impact is DoS, but the primitive is the concern.","attack_vector":"Privileged local software on the host (Ring 0 / bare-metal OS). Any tenant with root on a rented bare-metal node qualifies.","remediation":"Flash E810 firmware to cvl fw 1.7.8.x or later; cold power cycle. Beyond the patch: if you rent bare metal, reflash NIC firmware from a known-good image at tenant handoff and verify the version afterwards, because a patched-but-unverified NIC is not a clean NIC.","references":["https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01171.html","https://nvd.nist.gov/vuln/detail/CVE-2025-27243"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-37149","cve":"CVE-2025-37149","aliases":["HPESBHF04952"],"title":"HPE ProLiant RL300 Gen11 (UEFI firmware, out-of-bounds read): Out-of-bounds reads in the UEFI firmware of the ProLiant RL300 Gen11 - HPE's Arm-based (Ampere) ProLiant. The…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE ProLiant RL300 Gen11 (UEFI firmware, out-of-bounds read)","year":"2025","cvss_score":6,"severity":"medium","kev":false,"impact":"Out-of-bounds reads in the UEFI firmware of the ProLiant RL300 Gen11 - HPE's Arm-based (Ampere) ProLiant. The CVSS vector marks scope as changed with high confidentiality impact, meaning the leak crosses a trust boundary out of the firmware context. Firmware-level memory disclosure is how an attacker recovers the addresses and secrets that make a subsequent firmware write reliable, so treat it as an enabler for a persistence attack rather than as a standalone data-loss event. Relevant to operators running mixed-architecture racks where Arm nodes handle inference or control-plane duty alongside x86 GPU boxes.","attack_vector":"Local to the host with high privileges - a root/administrator account on the operating system. Not remotely reachable and not exposed on the management VLAN.","remediation":"UEFI firmware update on the affected RL300 Gen11 nodes. As with any system firmware, it applies on the next reboot, so it costs a maintenance window per node rather than a live out-of-band flash. Small affected footprint means the campaign should be quick to scope - identify RL300 Gen11 nodes specifically, since the rest of the ProLiant Gen11 line is not in scope. No config-only mitigation.","references":["https://support.hpe.com/hpesc/public/docDisplay?docId=hpesbhf04952en_us&docLocale=en_US","https://nvd.nist.gov/vuln/detail/CVE-2025-37149"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-1386","cve":"CVE-2026-1386","aliases":[],"title":"Firecracker: Symlink following in the jailer lets a local host user with write access to pre-created jailer directories…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"Symlink following in the jailer lets a local host user with write access to pre-created jailer directories overwrite arbitrary files","attack_vector":"A local host user","remediation":"Upgrade Firecracker; tighten jailer directory ownership","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-1386"],"status":"curated"},{"id":"CVE-2026-15792","cve":"CVE-2026-15792","aliases":[],"title":"BuildKit: Malicious client or frontend panics the BuildKit daemon","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"Malicious client or frontend panics the BuildKit daemon","attack_vector":"Anyone with build API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15792"],"status":"curated"},{"id":"CVE-2026-41184","cve":"CVE-2026-41184","aliases":[],"title":"Calico: install-cni logs the rendered CNI config including the substituted service-account token","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"install-cni logs the rendered CNI config including the substituted service-account token","attack_vector":"Anyone with pod-log read access","remediation":"Rolling Calico upgrade; rotate the CNI service-account token; scrub logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41184"],"status":"curated"},{"id":"CVE-2026-41185","cve":"CVE-2026-41185","aliases":[],"title":"Calico: Azure IPAM helper logs the mutated CNI config including credentials","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"Azure IPAM helper logs the mutated CNI config including credentials","attack_vector":"Anyone with pod-log read access","remediation":"Rolling Calico upgrade; rotate leaked credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41185"],"status":"curated"},{"id":"CVE-2026-41186","cve":"CVE-2026-41186","aliases":[],"title":"Calico: kube-controllers and Goldmane bind an unauthenticated pprof listener to 0.0.0.0","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2026","cvss_score":6,"severity":"medium","kev":false,"impact":"kube-controllers and Goldmane bind an unauthenticated pprof listener to 0.0.0.0","attack_vector":"Any pod on the cluster network","remediation":"Rolling Calico upgrade; keep the debug server disabled (it is off by default)","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41186"],"status":"curated"},{"id":"CVE-2017-15361","cve":"CVE-2017-15361","aliases":["ROCA","Return of Coppersmith's Attack"],"title":"Infineon TPM firmware (RSA key generation): RSA keys generated inside affected Infineon TPMs are factorable from the public key alone - no access to the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Infineon TPM firmware (RSA key generation)","year":"2017","cvss_score":5.9,"severity":"medium","kev":false,"impact":"RSA keys generated inside affected Infineon TPMs are factorable from the public key alone - no access to the machine required. Every key the TPM ever produced is retroactively compromised: attestation identity keys, sealed storage keys, machine certificates, and any SSH or code-signing key an operator generated in the TPM believing it was hardware-protected. For a fleet, that means the attestation evidence you have been collecting is forgeable by anyone who saw a public key.","attack_vector":"No access to the hardware at all. The attacker needs only a public key that the TPM generated - which by definition has been published to whatever service consumed it.","remediation":"Two-part and expensive. First a TPM firmware update from the platform OEM, usually shipped inside a BIOS package, so it is a per-node flash plus reboot. Then - and this is the part that gets skipped - every key generated by the vulnerable TPM must be regenerated and re-enrolled, and the old ones revoked. Sealed data must be unsealed before the update or it becomes unrecoverable. Inventory which nodes carry Infineon TPMs before planning; on old hardware the OEM may never have shipped the fix, in which case move the trust anchor off the TPM.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-15361","https://www.infineon.com/cms/en/product/promopages/tpm-update/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-12130","cve":"CVE-2018-12130","aliases":[],"title":"Intel CPU (MDS / ZombieLoad): Microarchitectural Fill Buffer Data Sampling - cross-domain leak from fill buffers, including across SMT…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (MDS / ZombieLoad)","year":"2018","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Microarchitectural Fill Buffer Data Sampling - cross-domain leak from fill buffers, including across SMT siblings","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel VERW mitigation + reboot; SMT disable for full protection, with the corresponding throughput loss","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12130"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2019-16863","cve":"CVE-2019-16863","aliases":["TPM-FAIL"],"title":"STMicroelectronics ST33 TPM (ECDSA timing): Discrete TPM leaks ECDSA nonce data through timing, allowing private key recovery from observed signatures.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"STMicroelectronics ST33 TPM (ECDSA timing)","year":"2019","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Discrete TPM leaks ECDSA nonce data through timing, allowing private key recovery from observed signatures. The discrete-chip half of TPM-FAIL - which matters because the usual answer to the fTPM flaw is 'use a real TPM', and this shows the real TPM had the same class of problem.","attack_vector":"Local attacker able to request signatures from the TPM, with accurate timing measurement.","remediation":"TPM firmware update from ST, distributed through the platform OEM's BIOS package - per-node flash plus reboot, and vendor availability was patchy. Rotate any long-lived key the TPM produced. Where no update exists, treat that TPM's keys as software-grade rather than hardware-protected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16863","https://tpm.fail/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-16843","cve":"CVE-2020-16843","aliases":[],"title":"Firecracker: Network stack freezes under heavy ingress","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Firecracker","year":"2020","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Network stack freezes under heavy ingress; microVM DoS","attack_vector":"Unauthenticated network","remediation":"Upgrade Firecracker","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-16843"],"status":"curated"},{"id":"CVE-2020-26569","cve":"CVE-2020-26569","aliases":[],"title":"Arista EOS (EVPN VXLAN MAC/IP binding): TENANT ISOLATION: malformed packets create incorrect MAC-to-IP bindings in an EVPN VXLAN fabric, and packets…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (EVPN VXLAN MAC/IP binding)","year":"2020","cvss_score":5.9,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: malformed packets create incorrect MAC-to-IP bindings in an EVPN VXLAN fabric, and packets get forwarded across VLAN boundaries as a result. An attacker who can source crafted frames from inside one tenant's overlay can poison the fabric's bindings and cause traffic to cross into the wrong VLAN. The advisory notes traffic is discarded on the receiving VLAN, which limits it to a leak-and-drop rather than a clean interception — but it is still a control-plane-driven breach of the segmentation model.","attack_vector":"A host inside an EVPN VXLAN tenant network able to emit specific malformed packets.","remediation":"EOS upgrade plus reload across the VTEP layer. No live workaround. If you run EVPN multi-tenancy, also audit the MAC/IP binding table for entries that do not correspond to a real workload.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-26569"],"status":"curated"},{"id":"CVE-2020-8553","cve":"CVE-2020-8553","aliases":[],"title":"ingress-nginx: A tenant can overwrite another ingress's basic-auth password file","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2020","cvss_score":5.9,"severity":"medium","kev":false,"impact":"A tenant can overwrite another ingress's basic-auth password file","attack_vector":"Cluster user able to create namespaces and Ingress objects","remediation":"Rolling controller upgrade, no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8553"],"status":"curated"},{"id":"CVE-2021-20199","cve":"CVE-2021-20199","aliases":[],"title":"Podman: Rootless containers see all traffic as coming from 127.0.0.1, defeating localhost-trust checks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2021","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Rootless containers see all traffic as coming from 127.0.0.1, defeating localhost-trust checks","attack_vector":"Unauthenticated network","remediation":"Upgrade Podman; never trust source 127.0.0.1 in containerised apps","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-20199"],"status":"curated"},{"id":"CVE-2022-24769","cve":"CVE-2022-24769","aliases":[],"title":"Docker / moby: Containers started with non-empty inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2022","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Containers started with non-empty inheritable capabilities; unexpected privilege retention on setuid binaries","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine / moby; restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24769"],"status":"curated"},{"id":"CVE-2022-29162","cve":"CVE-2022-29162","aliases":[],"title":"runc: `runc exec --cap` created processes with non-empty inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2022","cvss_score":5.9,"severity":"medium","kev":false,"impact":"`runc exec --cap` created processes with non-empty inheritable capabilities; unexpected privilege retention","attack_vector":"Any tenant workload","remediation":"Replace runc binary; restart affected containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-29162"],"status":"curated"},{"id":"CVE-2023-20902","cve":"CVE-2023-20902","aliases":[],"title":"Harbor: Timing condition allows creating and stopping jobs and retrieving job info","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2023","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Timing condition allows creating and stopping jobs and retrieving job info","attack_vector":"Unauthenticated network access to Harbor","remediation":"Upgrade Harbor","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20902"],"status":"curated"},{"id":"CVE-2023-48795","cve":"CVE-2023-48795","aliases":[],"title":"OpenSSH (transport): Terrapin: prefix-truncation attack on the SSH Binary Packet Protocol - silently downgrades channel integrity","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (transport)","year":"2023","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Terrapin: prefix-truncation attack on the SSH Binary Packet Protocol - silently downgrades channel integrity","attack_vector":"Unauthenticated network (MITM position)","remediation":"Package update on both ends + sshd restart; no reboot","references":["https://access.redhat.com/security/cve/CVE-2023-48795"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2023-5502","cve":"CVE-2023-5502","aliases":[],"title":"Arista EOS (802.1X on access/trunk ports): TENANT ISOLATION: with 802.1X configured on access or trunk ports and routing enabled on the access VLAN, a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (802.1X on access/trunk ports)","year":"2023","cvss_score":5.9,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: with 802.1X configured on access or trunk ports and routing enabled on the access VLAN, a malicious supplicant can skip 802.1X authentication entirely. Port-based admission control is what stops an unauthorized machine being plugged into a rack and joining the fabric — this makes it optional. Companion issue CVE-2024-6858 does the same thing in multi-auth mode via a device in the fallback VLAN.","attack_vector":"A device physically connected to a switch port that has 802.1X configured. Colocation, shared cages, and contractor rack-and-stack are the realistic scenarios.","remediation":"EOS upgrade plus reload. Do not rely on 802.1X alone as the tenant admission boundary; combine it with per-port VLAN pinning and MAC allowlisting (live config) so a bypassed supplicant still lands nowhere useful.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-5502","https://nvd.nist.gov/vuln/detail/CVE-2024-6858"],"status":"curated"},{"id":"CVE-2024-29018","cve":"CVE-2024-29018","aliases":[],"title":"Docker / moby: DNS requests from an internal network can be forwarded to external resolvers, leaking data out of an isolated…","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":5.9,"severity":"medium","kev":false,"impact":"DNS requests from an internal network can be forwarded to external resolvers, leaking data out of an isolated network","attack_vector":"Any tenant workload on an \"internal\" Docker network","remediation":"Upgrade moby","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29018"],"status":"curated"},{"id":"CVE-2024-9042","cve":"CVE-2024-9042","aliases":[],"title":"Kubernetes (kubelet): Command injection on Windows nodes via the nodes/*/logs/query API","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2024","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Command injection on Windows nodes via the nodes/*/logs/query API","attack_vector":"Cluster user with node log-query rights","remediation":"Rolling kubelet upgrade; Windows node drain; restrict nodes/log RBAC","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"ubiquity":"Niche in GPU clouds - Windows GPU nodes are rare, but the `nodes/*/logs/query` path is a reminder that kubelet log endpoints reach the host shell","remediation_pain":"`daemon-restart` - kubelet upgrade to v1.32.1 / v1.31.5 / v1.30.9 / v1.29.13, node-by-node","pain_class":"daemon-restart","why_fleet_wide":"Anyone with `nodes/*/logs` read rights injects into PowerShell and executes as SYSTEM on the node; low ubiquity in GPU fleets keeps this off the emergency list"}},{"id":"CVE-2025-23333","cve":"CVE-2025-23333","aliases":[],"title":"NVIDIA Triton Inference Server: MULTI-TENANT ISOLATION: manipulating the Python backend's shared memory region produces an out-of-bounds read…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: manipulating the Python backend's shared memory region produces an out-of-bounds read and leaks data. Where one Triton instance serves several models or tenants, that shared memory holds other requests' tensors. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23333","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"]},{"id":"CVE-2025-23334","cve":"CVE-2025-23334","aliases":[],"title":"NVIDIA Triton Inference Server: MULTI-TENANT ISOLATION: a crafted request causes an out-of-bounds read in the Python backend, disclosing…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: a crafted request causes an out-of-bounds read in the Python backend, disclosing memory that can include co-resident request data. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23334","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"]},{"id":"CVE-2025-26466","cve":"CVE-2025-26466","aliases":[],"title":"OpenSSH (sshd): Pre-auth memory/CPU amplification - denial of service against sshd","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"OpenSSH (sshd)","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Pre-auth memory/CPU amplification - denial of service against sshd","attack_vector":"Unauthenticated network","remediation":"Package update + sshd restart; rate-limit with PerSourcePenalties","references":["https://access.redhat.com/security/cve/CVE-2025-26466"],"status":"curated"},{"id":"CVE-2025-29948","cve":"CVE-2025-29948","aliases":[],"title":"AMD SEV firmware - RMP protection bypass: MULTI-TENANT ISOLATION: An access-control failure in SEV firmware lets a malicious hypervisor bypass…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - RMP protection bypass","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An access-control failure in SEV firmware lets a malicious hypervisor bypass reverse-map table protections and break SEV-SNP guest memory integrity. The RMP is the single structure standing between a hostile host and a confidential guest's pages; a firmware-level bypass of it means the isolation you are selling is not enforced.","attack_vector":"Malicious hypervisor - host-privileged.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29948","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-29952","cve":"CVE-2025-29952","aliases":[],"title":"AMD SEV firmware - improper initialization corrupting RMP-covered memory: MULTI-TENANT ISOLATION: An initialization defect in SEV firmware lets an admin-privileged attacker corrupt…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - improper initialization corrupting RMP-covered memory","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An initialization defect in SEV firmware lets an admin-privileged attacker corrupt memory covered by the RMP, costing confidential guest integrity. Same family as the other RMP issues in this batch and shipped in the same AMD advisory wave - the recurring theme is that SNP's protections are only as good as the firmware that sets them up.","attack_vector":"Local, admin-privileged.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29952","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33242","cve":"CVE-2025-33242","aliases":[],"title":"HGX / DGX B300: Authentication bypass in the hardware management interface","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"HGX / DGX B300","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Authentication bypass in the hardware management interface","attack_vector":"Network-adjacent attacker on the mgmt path","remediation":"Flash HGX/DGX management firmware out-of-band","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33242","https://github.com/NVIDIA/product-security/tree/main/2026/5768"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:N/I:H/A:H","cwe":["CWE-1234"]},{"id":"CVE-2025-54510","cve":"CVE-2025-54510","aliases":[],"title":"AMD Secure Processor firmware - MMIO routing lock (Zen 5): MULTI-TENANT ISOLATION: A missing lock check in ASP firmware on some Zen 5 parts lets a locally authenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor firmware - MMIO routing lock (Zen 5)","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A missing lock check in ASP firmware on some Zen 5 parts lets a locally authenticated administrator alter MMIO routing after the point where it should have been frozen. Redirecting MMIO means pointing a device's window somewhere it should not go - which is how a host administrator reaches into a confidential guest's memory despite SEV-SNP.","attack_vector":"Local, administrative privilege on the host. Affects Zen 5 (Turin / EPYC 9005) generation parts.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54510","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-61971","cve":"CVE-2025-61971","aliases":[],"title":"AMD NBIO register lock bits - MMIO routing configuration: MULTI-TENANT ISOLATION: The sibling of the SMN issue: unprotected NBIO lock bits let a local administrator…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD NBIO register lock bits - MMIO routing configuration","year":"2025","cvss_score":5.9,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The sibling of the SMN issue: unprotected NBIO lock bits let a local administrator rewrite MMIO routing configuration and break SEV-SNP guest integrity. The host operator can point address windows at confidential guest memory that the RMP was supposed to fence off.","attack_vector":"Local, host administrator privilege.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Same OEM BIOS package as the SMN lock-bit issue; do not patch one and leave the other.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-61971","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24231","cve":"CVE-2026-24231","aliases":[],"title":"NemoClaw: SSRF","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NemoClaw","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"SSRF -> internal service access from a tenant-facing service","attack_vector":"Tenant supplying a crafted URL","remediation":"Upgrade the service; add egress network policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24231","https://github.com/NVIDIA/product-security/tree/main/2026/5837"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:C/C:H/I:N/A:N","cwe":["CWE-918"]},{"id":"CVE-2026-24266","cve":"CVE-2026-24266","aliases":[],"title":"Triton Inference Server: Use-after-free in request processing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Use-after-free in request processing","attack_vector":"Any inference client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24266","https://github.com/NVIDIA/product-security/tree/main/2026/5848"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-416"]},{"id":"CVE-2026-27482","cve":"CVE-2026-27482","aliases":[],"title":"Ray (dashboard DELETE endpoints): Browser-origin protection covers POST/PUT but not DELETE","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Ray (dashboard DELETE endpoints)","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"Browser-origin protection covers POST/PUT but not DELETE; key DELETE endpoints unauthenticated","attack_vector":"Malicious page visited by an operator with dashboard access","remediation":"Upgrade past 2.53.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-27482"],"status":"curated"},{"id":"CVE-2026-56742","cve":"CVE-2026-56742","aliases":[],"title":"Cilium: A namespaced HTTPRoute can mirror another tenant's HTTP traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":5.9,"severity":"medium","kev":false,"impact":"A namespaced HTTPRoute can mirror another tenant's HTTP traffic; cross-tenant data exfiltration","attack_vector":"Cluster user with namespace access and Gateway API rights","remediation":"Rolling Cilium upgrade; restrict HTTPRoute mirroring","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-56742"],"status":"curated"},{"id":"CVE-2020-15115","cve":"CVE-2020-15115","aliases":[],"title":"etcd: No password length validation permits one-character etcd passwords","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"etcd","year":"2020","cvss_score":5.8,"severity":"medium","kev":false,"impact":"No password length validation permits one-character etcd passwords","attack_vector":"Unauthenticated network brute force","remediation":"Rolling etcd upgrade; move to certificate auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15115"],"status":"curated"},{"id":"CVE-2021-25736","cve":"CVE-2021-25736","aliases":[],"title":"Kubernetes (kube-proxy): Windows kube-proxy forwards LoadBalancer traffic to local processes on the same port","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-proxy)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Windows kube-proxy forwards LoadBalancer traffic to local processes on the same port","attack_vector":"Unauthenticated network via the LB","remediation":"Rolling kube-proxy upgrade on Windows nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25736"],"status":"curated"},{"id":"CVE-2021-28511","cve":"CVE-2021-28511","aliases":[],"title":"Arista EOS (security ACL vs NAT rule interaction): TENANT ISOLATION: a security ACL drop rule is bypassed when a NAT ACL permit rule matches the same packet.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (security ACL vs NAT rule interaction)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: a security ACL drop rule is bypassed when a NAT ACL permit rule matches the same packet. Traffic you explicitly denied is forwarded. Same class of problem as the VXLAN ACL bug — the enforcement does not match the config, so your segmentation audit passes while the boundary is open.","attack_vector":"Any source whose traffic matches both a NAT permit and a security deny. Requires NAT to be configured on the device, which is common on the cluster's egress or storage-gateway leaves.","remediation":"EOS upgrade plus reload. Interim: avoid overlapping NAT and security ACL match spaces on the same device, and verify enforcement with actual traffic tests rather than reading the config. Live config change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28511"],"status":"curated"},{"id":"CVE-2021-4228","cve":"CVE-2021-4228","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 (BMC default TLS certificate): The BMC ships with a hard-coded default TLS certificate, so HTTPS to the management interface can be…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 (BMC default TLS certificate)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"The BMC ships with a hard-coded default TLS certificate, so HTTPS to the management interface can be transparently intercepted by anyone holding the extracted key - which is everyone, since it is identical across all devices built from that firmware. The operator loses exactly what they thought they were buying with HTTPS: BMC admin passwords, Redfish tokens and KVM traffic are readable to an attacker who can sit in the path. Because the same certificate is on every node, a single extraction compromises the whole management plane.","attack_vector":"Requires a man-in-the-middle position on the management network - a compromised jump host, a rogue device on the management VLAN, or control of a switch or DHCP server on that segment. No credentials needed. Disclosed by Nozomi Labs against a Lanner IAC-AST2500A platform; AMI's own advisory confirms the affected code is part of MegaRAC SPx, so the exposure is not limited to that one vendor's box.","remediation":"Firmware flash to SPx_12-update-3.00 or later (AMI states SPx_13 is not affected), but flashing alone does not fix a node whose certificate is already installed. The operative fix is config-only and can be done today across the fleet without a reboot: generate a unique certificate per BMC from your own internal CA and push it over Redfish, then have your management tooling actually pin or verify it rather than skipping certificate validation - which is the default in most homegrown BMC scrapers and is the reason this bug stays exploitable.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-4228/","https://nvd.nist.gov/vuln/detail/CVE-2021-4228"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-46279","cve":"CVE-2021-46279","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC web session management): Session fixation combined with sessions that never properly expire. An attacker can plant a session…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC web session management)","year":"2021","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Session fixation combined with sessions that never properly expire. An attacker can plant a session identifier, wait for an administrator to authenticate with it, and inherit a live admin session on the BMC - power control, console, virtual media, firmware update. The never-expiring half means stolen sessions stay valid long after the admin walked away, so a token lifted from a browser, a proxy log or a shared jump host keeps working for as long as the attacker wants it.","attack_vector":"Network access to the BMC web interface plus getting an administrator to interact with an attacker-supplied session - a link, a shared workstation, or a proxy on the management path. High complexity, no credentials required.","remediation":"Firmware flash to SPx_12-update-7.00 / SPx_13-update-5.00 or later, out-of-band per node, ODM-gated. Config-only mitigations that help immediately: restrict BMC web access to a bastion, do not let admins browse anything else from that host, and force logout rather than closing the tab - the sessions this bug leaves behind are the ones that never got explicitly ended.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-46279/","https://nvd.nist.gov/vuln/detail/CVE-2021-46279"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-2538","cve":"CVE-2023-2538","aliases":[],"title":"Tyan S5552 BMC web interface, firmware version 3.00: An unauthenticated attacker downloads the BMC's TLS private key. With it they can decrypt captured management…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tyan S5552 BMC web interface, firmware version 3.00","year":"2023","cvss_score":5.8,"severity":"medium","kev":false,"impact":"An unauthenticated attacker downloads the BMC's TLS private key. With it they can decrypt captured management traffic and impersonate the BMC to your own tooling - meaning your provisioning system, monitoring collector and operators can be fed a controller that is not the controller, and will hand over BMC credentials to it. Because vendors frequently ship the same certificate across a production run, a key pulled from one Tyan S5552 may authenticate an impersonated BMC across every node of that model in the fleet. That turns a medium-scored file disclosure into a fleet-wide management-plane credential harvest. The private key for the TLS certificate the BMC presents is retrievable by forced browsing, with no authentication.","attack_vector":"Unauthenticated HTTP access to the BMC web interface - anything routable to the out-of-band management VLAN. No credential and no host foothold required.","remediation":"Firmware update from Tyan, and this is where an operator hits a wall: Tyan has been folded into MiTAC Computing, www.tyan.com no longer presents a valid TLS certificate for its own hostname, and mitaccomputing.com returns 403 to automated clients - so there is no reachable vendor PSIRT to obtain a fixed image from. The advisory that exists is third-party, from Nozomi Networks. Regardless of firmware state, replace the BMC's TLS certificate with one you generated and control, and rotate any BMC credentials that were transmitted to that BMC over a session an attacker could have decrypted. Certificate replacement is a config-only change and should be done on every BMC in the fleet as standard practice, not just Tyan ones.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2538","https://www.nozominetworks.com/labs/vulnerability-advisories-cve-2023-2538/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-52529","cve":"CVE-2024-52529","aliases":[],"title":"Cilium: L3 port-range plus L7 allow combination results in over-permissive policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":5.8,"severity":"medium","kev":false,"impact":"L3 port-range plus L7 allow combination results in over-permissive policy","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-52529"],"status":"curated"},{"id":"CVE-2024-53197","cve":"CVE-2024-53197","aliases":[],"title":"Linux kernel (ALSA usb-audio): Out-of-bounds access for Extigy/Mbox devices","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ALSA usb-audio)","year":"2024","cvss_score":5.8,"severity":"medium","kev":true,"impact":"Out-of-bounds access for Extigy/Mbox devices; part of a real-world Android forensic-unlock chain [KEV]","attack_vector":"Local user with USB device access","remediation":"Livepatchable; otherwise drain + reboot. Physical-access class - matters mainly for colo cages with weak physical controls","references":["https://access.redhat.com/security/cve/CVE-2024-53197"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-6095","cve":"CVE-2024-6095","aliases":[],"title":"LocalAI (`/models/apply`): SSRF and partial local file inclusion","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"LocalAI (`/models/apply`)","year":"2024","cvss_score":5.8,"severity":"medium","kev":false,"impact":"SSRF and partial local file inclusion","attack_vector":"Unauthenticated network","remediation":"Upgrade past 2.15.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6095"],"status":"curated"},{"id":"CVE-2024-6437","cve":"CVE-2024-6437","aliases":[],"title":"Arista EOS (PBR / BGP Flowspec / interface traffic policy): TENANT ISOLATION: IPv4 packets carrying IP options can bypass policy-based routing, BGP Flowspec and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (PBR / BGP Flowspec / interface traffic policy)","year":"2024","cvss_score":5.8,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: IPv4 packets carrying IP options can bypass policy-based routing, BGP Flowspec and interface traffic policy redirection. If you use PBR or Flowspec to steer a tenant's traffic through an inspection or scrubbing path, an attacker sets an IP option and goes around it. Flowspec-based DDoS mitigation on the cluster edge fails the same way.","attack_vector":"Any sender able to emit IPv4 packets with IP options toward an interface with the affected redirection configured.","remediation":"EOS upgrade plus reload. Interim: drop IPv4 packets with IP options at the edge with an ACL — a live config change, and reasonable policy in a datacenter fabric where IP options have no legitimate use.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-6437"],"status":"curated"},{"id":"CVE-2025-13281","cve":"CVE-2025-13281","aliases":[],"title":"Kubernetes (kube-controller-manager): Half-blind SSRF via the Portworx in-tree volume plugin in kube-controller-manager","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-controller-manager)","year":"2025","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Half-blind SSRF via the Portworx in-tree volume plugin in kube-controller-manager","attack_vector":"Cluster user able to create a Portworx volume","remediation":"Rolling control-plane upgrade; remove the in-tree Portworx plugin","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-33043","cve":"CVE-2025-33043","aliases":["AMI-SA-2025005","Binarly rediscovery"],"title":"AMI AptioV UEFI BIOS: Improper input validation in the BIOS with an integrity impact and a changed scope. The score understates why…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS","year":"2025","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Improper input validation in the BIOS with an integrity impact and a changed scope. The score understates why this one matters: AMI states in the advisory that the issue was identified, fixed and disclosed under NDA back in 2018, and that Binarly rediscovered it still unmitigated in multiple systems in the field seven years later. For a fleet operator that is a direct statement about the supply chain - a fix existing at AMI is not the same as a fix reaching your hardware, and the gap can be measured in years. Assume the same is true of every other AptioV CVE in this cluster on any SKU you have not explicitly verified.","attack_vector":"Local access with high privileges and user interaction, at high attack complexity. Requires an attacker who already has administrative control of the host and can induce the right operation - so it is a persistence and privilege-depth bug rather than an entry point.","remediation":"BIOS update to AptioV_5.011 or later - firmware flash plus reboot per node. The specific action this CVE demands is different from the others: do not trust version numbers, verify. Pull the actual firmware image off a representative node per SKU and confirm the fix is present, because the whole point of this advisory is that vendors shipped systems that never got the 2018 fix. Prioritise older and white-box SKUs, and any hardware acquired second-hand or through a broker.","references":["https://go.ami.com/hubfs/Security%20Advisories/2025/AMI-SA-2025005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2025-33043"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24201","cve":"CVE-2026-24201","aliases":[],"title":"vGPU Manager: Host impact via GPU command-buffer overflow","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2026","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Host impact via GPU command-buffer overflow","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate guest VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24201","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-787"]},{"id":"CVE-2026-44210","cve":"CVE-2026-44210","aliases":[],"title":"Kata Containers: Default configuration allows pod creators more than intended","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kata Containers","year":"2026","cvss_score":5.8,"severity":"medium","kev":false,"impact":"Default configuration allows pod creators more than intended","attack_vector":"Cluster user with namespace access","remediation":"Upgrade Kata to 3.31.0+ and harden the default config","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-44210"],"status":"curated"},{"id":"CVE-2026-7473","cve":"CVE-2026-7473","aliases":[],"title":"Arista EOS (tunnel decapsulation): **[KEV]** With VXLAN, decap-groups or GRE configured, the switch incorrectly decapsulates and forwards…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (tunnel decapsulation)","year":"2026","cvss_score":5.8,"severity":"medium","kev":true,"impact":"**[KEV]** With VXLAN, decap-groups or GRE configured, the switch incorrectly decapsulates and forwards unexpected tunnelled packets — a tenant-isolation break on a VXLAN-segmented GPU fabric, exploited in the wild","attack_vector":"Network, fabric-local","remediation":"EOS upgrade with fabric failover; on a multi-tenant VXLAN fabric this is a segmentation failure, so it also demands a review of whether cross-tenant traffic actually occurred","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-7473"],"status":"curated"},{"id":"CVE-2020-15707","cve":"CVE-2020-15707","aliases":["BootHole family"],"title":"GRUB2 (initrd size handling): Integer overflows in the initrd command's size arithmetic corrupt GRUB's heap. Same end state as the rest of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (initrd size handling)","year":"2020","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Integer overflows in the initrd command's size arithmetic corrupt GRUB's heap. Same end state as the rest of the family - unsigned code running pre-kernel with Secure Boot still claiming to be enforcing.","attack_vector":"Requires the attacker to control the initrd list, i.e. write access to boot configuration on the node.","remediation":"grub2 package update + reboot. No config-only mitigation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15707","https://ubuntu.com/security/CVE-2020-15707"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-3696","cve":"CVE-2021-3696","aliases":[],"title":"GRUB2 (PNG grayscale reader): Out-of-bounds write on the grayscale PNG path. Same shape as the colour path but a separate code path that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (PNG grayscale reader)","year":"2021","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Out-of-bounds write on the grayscale PNG path. Same shape as the colour path but a separate code path that the first patch did not cover - relevant if you patched early and stopped.","attack_vector":"Attacker-supplied boot splash image.","remediation":"grub2 package update + reboot. Confirm your package version covers this one specifically, not just the headline PNG CVE.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3696","https://access.redhat.com/security/cve/CVE-2021-3696"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-23471","cve":"CVE-2022-23471","aliases":[],"title":"containerd: Goroutine leak in the CRI stream server terminal-resize path exhausts host memory","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2022","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Goroutine leak in the CRI stream server terminal-resize path exhausts host memory; node DoS","attack_vector":"Any tenant workload that opens exec/attach sessions","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23471"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-25534","cve":"CVE-2023-25534","aliases":[],"title":"DGX H100 BMC (IPMI): Info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (IPMI)","year":"2023","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Info disclosure","attack_vector":"Network-adjacent IPMI client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-20"]},{"id":"CVE-2023-32690","cve":"CVE-2023-32690","aliases":["libspdm CTExponent"],"title":"DMTF libspdm - SPDM Requester timeout handling: A libspdm Requester stores the Responder's CTExponent without validating it, so a malicious or faulty…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DMTF libspdm - SPDM Requester timeout handling","year":"2023","cvss_score":5.7,"severity":"medium","kev":false,"impact":"A libspdm Requester stores the Responder's CTExponent without validating it, so a malicious or faulty responder can force an enormous computed timeout and hang the requester. In an attestation flow this is a denial of service against the thing that decides whether a device is trustworthy - and a hung attestation is often failed open by the surrounding orchestration.","attack_vector":"Adjacent, unauthenticated with user interaction. A device on the link that answers CAPABILITIES dishonestly.","remediation":"Update to libspdm 2.3.3 / 3.0 or later, again through your device vendor's firmware. Separately, check what your orchestration does when attestation times out rather than fails - failing open on timeout is the more damaging half of this.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-56h8-4gv5-jf2c","https://nvd.nist.gov/vuln/detail/CVE-2023-32690"],"status":"curated"},{"id":"CVE-2023-34472","cve":"CVE-2023-34472","aliases":["AMI-SA-2023006","Nozomi Labs BMC audit"],"title":"AMI MegaRAC SPx (BMC web interface, HTTP header handling): CRLF sequences are not neutralised in HTTP headers, so an attacker can split responses and inject headers of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (BMC web interface, HTTP header handling)","year":"2023","cvss_score":5.7,"severity":"medium","kev":false,"impact":"CRLF sequences are not neutralised in HTTP headers, so an attacker can split responses and inject headers of their choosing. Against a BMC web UI the payoff is session and cache manipulation against an administrator's browser - poisoning what the admin sees, planting cookies, or setting up a follow-on credential capture. It is an integrity bug that is useful as a stepping stone toward hijacking an admin's BMC session rather than a direct takeover.","attack_vector":"Adjacent network with a low-privilege BMC account. Requires an administrator to subsequently interact with the BMC web interface for the payoff, so it depends on your ops team actually using the web UI - which most do for KVM and console access.","remediation":"Firmware flash to SPx_12.5 / SPx_13.3 or later, out-of-band per node, ODM-gated. Low priority relative to the rest of this cluster, so fold it into the same flash campaign rather than scheduling separately. Config-only reduction: reach BMC web UIs only from a hardened jump host with a dedicated browser profile, so an admin session cannot be crossed with anything else.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023006.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34472"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21981","cve":"CVE-2024-21981","aliases":[],"title":"AMD Secure Processor - cryptographic key usage control: MULTI-TENANT ISOLATION: Once an attacker has arbitrary code execution inside the ASP, weak key-usage controls…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - cryptographic key usage control","year":"2024","cvss_score":5.7,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Once an attacker has arbitrary code execution inside the ASP, weak key-usage controls let them extract the ASP's cryptographic keys rather than merely using them. Extracted platform keys are portable: they can be used off-box to forge attestation material or decrypt data long after you have re-imaged the node, so this turns a contained firmware compromise into a lasting one.","attack_vector":"Local, and requires having already achieved code execution in the ASP - it is a privilege-amplifier chained behind one of the other ASP bugs, not a standalone entry point.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. If you have reason to believe a node's ASP was compromised, patching does not undo key extraction - the platform keys must be considered burned and the node's attestation identity retired.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21981","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23272","cve":"CVE-2025-23272","aliases":[],"title":"CUDA Toolkit: Info disclosure / DoS (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Info disclosure / DoS (buffer over-read)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23272","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:H","cwe":["CWE-125"]},{"id":"CVE-2025-33191","cve":"CVE-2025-33191","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An invalid memory read in OSROOT firmware crashes the platform. These sit in the GB10 root-of-trust chain, so…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"An invalid memory read in OSROOT firmware crashes the platform. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33191","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-20"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33192","cve":"CVE-2025-33192","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: An arbitrary memory read in SROOT firmware gives a denial of service and exposes firmware memory layout.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"An arbitrary memory read in SROOT firmware gives a denial of service and exposes firmware memory layout. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33192","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-690"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33193","cve":"CVE-2025-33193","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: SROOT firmware validates integrity improperly, leaking information that the root of trust was supposed to…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"SROOT firmware validates integrity improperly, leaking information that the root of trust was supposed to protect. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33193","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-354"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33194","cve":"CVE-2025-33194","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: Improper input processing in SROOT firmware yields information disclosure or a crash. These sit in the GB10…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":5.7,"severity":"medium","kev":false,"impact":"Improper input processing in SROOT firmware yields information disclosure or a crash. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33194","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:L/I:N/A:L","cwe":["CWE-180"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24215","cve":"CVE-2026-24215","aliases":[],"title":"Triton Inference Server: DoS via connection flooding","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":5.7,"severity":"medium","kev":false,"impact":"DoS via connection flooding","attack_vector":"Network-adjacent unauthenticated","remediation":"Upgrade Triton; add rate limiting","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24215","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-400"]},{"id":"CVE-2017-5715","cve":"CVE-2017-5715","aliases":["Spectre v2","Branch Target Injection"],"title":"Intel processors (indirect branch prediction): MULTI-TENANT ISOLATION: Spectre v2: an attacker trains the indirect branch predictor so that a victim context…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (indirect branch prediction)","year":"2017","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Spectre v2: an attacker trains the indirect branch predictor so that a victim context - another process, another VM, or the kernel - speculatively executes an attacker-chosen gadget and leaks its memory through a cache side channel. On a shared GPU host this is the canonical cross-VM and container-to-host read primitive, and it is still the reason retpoline, IBPB and eIBRS exist in every kernel you run.","attack_vector":"Local code execution anywhere on the host - any container, any VM. No privilege needed.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5715","https://security-center.intel.com/advisory.aspx?intelid=INTEL-SA-00088&languageid=en-fr","https://spectreattack.com/"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2017-5753","cve":"CVE-2017-5753","aliases":["Spectre v1","Bounds Check Bypass"],"title":"Intel processors (bounds check bypass): Spectre v1: speculative execution past a bounds check lets an attacker read memory the check was supposed to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (bounds check bypass)","year":"2017","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Spectre v1: speculative execution past a bounds check lets an attacker read memory the check was supposed to protect, within the same address space. The practical exposure on an AI node is inside anything that JITs or interprets untrusted input - eBPF, a Python runtime, a model-serving framework's custom-op path.","attack_vector":"Local code execution, including code inside a sandbox or interpreter that is meant to be confined.","remediation":"Software mitigation in the kernel and in individual programs (array index masking, speculation barriers) rather than microcode. Take kernel updates, keep runtimes current, and assume any interpreter you expose to untrusted input needs its own hardening. Reboot for the kernel component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5753","https://spectreattack.com/","https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/advisory-guidance/bounds-check-bypass.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2017-5754","cve":"CVE-2017-5754","aliases":["Meltdown","Rogue Data Cache Load"],"title":"Intel processors (rogue data cache load): MULTI-TENANT ISOLATION: Meltdown: unprivileged code reads kernel memory - and on affected parts, memory…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (rogue data cache load)","year":"2017","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Meltdown: unprivileged code reads kernel memory - and on affected parts, memory belonging to other tenants mapped into the kernel's direct map - by exploiting out-of-order execution past a permission check. The mitigation, kernel page-table isolation, is the reason every syscall on affected hardware got measurably slower.","attack_vector":"Any local unprivileged code on an affected processor.","remediation":"Kernel page-table isolation (KPTI/PTI), shipped in the kernel; take the kernel update and reboot. Newer silicon fixes it in hardware. No microcode or BIOS component for the mitigation itself. Expect a syscall-heavy throughput regression on affected parts - that is the fix working.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5754","https://meltdownattack.com/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-12126","cve":"CVE-2018-12126","aliases":["MSBDS","Fallout"],"title":"Intel processors (microarchitectural data sampling): MULTI-TENANT ISOLATION: One of the MDS family: store buffers retain data from other contexts and can be…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (microarchitectural data sampling)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: One of the MDS family: store buffers retain data from other contexts and can be sampled speculatively. The operator-relevant consequence is that data crosses between hyperthread siblings, between VMs, and out of enclaves without any architectural access - so on a node with SMT enabled and untrusted co-tenants, tenant isolation is not holding.","attack_vector":"Local code on the same physical core - with SMT enabled that includes a co-tenant on the sibling thread, which is the configuration most density-optimised fleets run.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The software half is buffer clearing on context switch (VERW), already in current kernels. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12126","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00233.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-12127","cve":"CVE-2018-12127","aliases":["MLPDS","RIDL"],"title":"Intel processors (microarchitectural data sampling): MULTI-TENANT ISOLATION: One of the MDS family: load ports retain data from other contexts and can be sampled…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (microarchitectural data sampling)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: One of the MDS family: load ports retain data from other contexts and can be sampled speculatively. The operator-relevant consequence is that data crosses between hyperthread siblings, between VMs, and out of enclaves without any architectural access - so on a node with SMT enabled and untrusted co-tenants, tenant isolation is not holding.","attack_vector":"Local code on the same physical core - with SMT enabled that includes a co-tenant on the sibling thread, which is the configuration most density-optimised fleets run.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The software half is buffer clearing on context switch (VERW), already in current kernels. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-12127","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00233.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-3620","cve":"CVE-2018-3620","aliases":["Foreshadow-OS","L1TF"],"title":"Intel processors (L1 terminal fault, OS/SMM): The OS-level variant of L1 terminal fault: a local user can speculatively read data present in the L1 cache…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (L1 terminal fault, OS/SMM)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"The OS-level variant of L1 terminal fault: a local user can speculatively read data present in the L1 cache belonging to the kernel or to system-management memory, by faulting on a carefully crafted page-table entry. Sits alongside the hypervisor-level L1TF variant as the reason page-table entry inversion exists in every modern kernel.","attack_vector":"Any local unprivileged code on an affected processor.","remediation":"Microcode update plus the kernel's PTE-inversion mitigation, then reboot. Microcode is late-loadable at boot without an OEM BIOS release. Verify with the kernel's l1tf sysfs vulnerability file after reboot rather than assuming.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3620","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00161.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-3640","cve":"CVE-2018-3640","aliases":["Spectre v3a","RSRE","Rogue System Register Read"],"title":"Intel processors (rogue system register read): Spectre v3a: speculative reads of system registers leak system parameters - MSR contents and similar - to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (rogue system register read)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Spectre v3a: speculative reads of system registers leak system parameters - MSR contents and similar - to unprivileged local code. On its own it exposes configuration rather than data, but that configuration is what an attacker needs to aim the more serious attacks.","attack_vector":"Local unprivileged code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3640","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00115.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-3646","cve":"CVE-2018-3646","aliases":[],"title":"Intel CPU (L1TF / Foreshadow-NG): L1 Terminal Fault: a guest reads any data present in the L1 data cache, including other guests' and the…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (L1TF / Foreshadow-NG)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"L1 Terminal Fault: a guest reads any data present in the L1 data cache, including other guests' and the hypervisor's memory","attack_vector":"Tenant VM guest","remediation":"Microcode + hypervisor L1D-flush mitigation + reboot; full mitigation requires core scheduling or disabling SMT, which costs roughly half the CPU throughput on a hyperthreaded host","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3646"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-3665","cve":"CVE-2018-3665","aliases":["LazyFP"],"title":"Intel processors (lazy FP state restore): LazyFP: when the OS restores FPU/vector state lazily, one process can speculatively read another process's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (lazy FP state restore)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"LazyFP: when the OS restores FPU/vector state lazily, one process can speculatively read another process's floating-point and vector register contents. On an AI node the vector registers hold tensor data and, in crypto-adjacent code, key material.","attack_vector":"Local code on a host whose OS uses lazy FPU state restore.","remediation":"Fixed in the OS by switching to eager FPU restore - take the kernel update and reboot. No microcode or BIOS component. Modern kernels already default to eager restore; this matters mostly on frozen legacy images.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3665","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00145.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-3693","cve":"CVE-2018-3693","aliases":["Spectre 1.1","Bounds Check Bypass Store"],"title":"Intel processors (bounds check bypass store): Spectre 1.1: speculative stores can overflow a bounds-checked buffer, letting an attacker write speculatively…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (bounds check bypass store)","year":"2018","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Spectre 1.1: speculative stores can overflow a bounds-checked buffer, letting an attacker write speculatively into structures that steer later speculation. Extends Spectre v1 from a read primitive to a write-and-redirect one inside a sandbox.","attack_vector":"Local code execution, particularly inside interpreters and JITs handling untrusted input.","remediation":"Software mitigation in compilers, kernels and runtimes rather than microcode. Take kernel and runtime updates; reboot for the kernel component.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3693","https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/advisory-guidance/bounds-check-bypass-store.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2019-11091","cve":"CVE-2019-11091","aliases":["MDSUM"],"title":"Intel processors (microarchitectural data sampling): MULTI-TENANT ISOLATION: One of the MDS family: uncacheable-memory accesses leave sampleable residue in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (microarchitectural data sampling)","year":"2019","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: One of the MDS family: uncacheable-memory accesses leave sampleable residue in microarchitectural buffers. The operator-relevant consequence is that data crosses between hyperthread siblings, between VMs, and out of enclaves without any architectural access - so on a node with SMT enabled and untrusted co-tenants, tenant isolation is not holding.","attack_vector":"Local code on the same physical core - with SMT enabled that includes a co-tenant on the sibling thread, which is the configuration most density-optimised fleets run.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The software half is buffer clearing on context switch (VERW), already in current kernels. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11091","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00233.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2019-1125","cve":"CVE-2019-1125","aliases":["SWAPGS","Spectre v1 SWAPGS variant","Spectre-SWAPGS"],"title":"Intel x86-64 CPUs (Ivy Bridge onward); Windows and Linux kernel entry paths: The kernel's syscall/interrupt entry path uses SWAPGS to switch the GS base between user and kernel values","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel x86-64 CPUs (Ivy Bridge onward); Windows and Linux kernel entry paths","year":"2019","cvss_score":5.6,"severity":"medium","kev":false,"impact":"The kernel's syscall/interrupt entry path uses SWAPGS to switch the GS base between user and kernel values; the CPU speculates past the conditional that decides whether to swap, so a user process can make the kernel speculatively dereference a user-controlled GS base and leak kernel memory through a cache channel. It bypasses the standard Spectre v1 mitigations because it lives in the entry code that runs before them. Operator exposure is a local process reading kernel memory - which on a container host means other tenants' data in the page cache and the kernel's credential structures.","attack_vector":"Unprivileged local code on the host or inside a guest, exercised through ordinary syscalls and interrupts. No SMT or core-sharing requirement.","remediation":"Kernel/hypervisor patch only - no microcode, no BIOS flash. The fix adds LFENCE serialization in the SWAPGS entry paths. Reboot into the patched kernel (drain the node), and that is the whole job. Runtime cost is a serializing instruction on kernel entry - measurable on syscall-microbenchmarks, effectively invisible on GPU workloads. Patched in Linux since 5.2 and backported everywhere, so on a current fleet this is closed; the audit item is confirming no host runs a pre-August-2019 kernel and that 'mitigations=off' is not set anywhere.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-1125","https://docs.kernel.org/admin-guide/hw-vuln/spectre.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-0550","cve":"CVE-2020-0550","aliases":["Snoop-assisted L1D Sampling","SnoopAssist"],"title":"Intel processors (snoop-assisted L1D sampling): Data can be leaked out of L1D during snoop transactions, crossing privilege boundaries. Lower practical yield…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (snoop-assisted L1D sampling)","year":"2020","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Data can be leaked out of L1D during snoop transactions, crossing privilege boundaries. Lower practical yield than the fill-buffer attacks but it targets the same shared L1 that makes SMT co-tenancy risky.","attack_vector":"Local code on the same physical core as the victim.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0550","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00330.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2020-0551","cve":"CVE-2020-0551","aliases":["LVI","Load Value Injection"],"title":"Intel processors / SGX (load value injection): MULTI-TENANT ISOLATION: The inverse of Meltdown: instead of leaking data out of the enclave, the attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors / SGX (load value injection)","year":"2020","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The inverse of Meltdown: instead of leaking data out of the enclave, the attacker injects a value into a faulting load inside the victim enclave and steers its transient execution into attacker-chosen gadgets. That gives enclave-secret extraction from outside the enclave, again breaking the SGX guarantee against a privileged host.","attack_vector":"Local privileged code on the host targeting a victim enclave on the same machine.","remediation":"SGX SDK/PSW update that inserts LFENCE serialisation in enclave code, plus microcode. The software mitigation requires recompiling enclaves with the patched SDK - a code change for whoever ships the enclave, not something the operator can apply unilaterally - and it carries a heavy performance cost. Microcode is late-loadable at boot; enclave recompilation is not. Re-attestation required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0551","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00334.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2021-26401","cve":"CVE-2021-26401","aliases":["LFENCE/JMP insufficiency"],"title":"AMD processors - LFENCE/JMP mitigation for Spectre v2 (CVE-2017-5715): MULTI-TENANT ISOLATION: The LFENCE/JMP sequence AMD originally recommended as the cheap Spectre-v2 mitigation…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - LFENCE/JMP mitigation for Spectre v2 (CVE-2017-5715)","year":"2021","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The LFENCE/JMP sequence AMD originally recommended as the cheap Spectre-v2 mitigation (mitigation V2-2) turns out not to sufficiently block branch target injection on some AMD CPUs. Anyone who took AMD's early guidance and chose LFENCE/JMP over retpoline because it was faster has been running with a Spectre-v2 mitigation that does not actually hold - a cross-VM and cross-process speculative disclosure channel that people believe is already closed. The dangerous part is the false sense of coverage, not the novelty of the attack.","attack_vector":"Local, cross-privilege and cross-guest speculative execution. Reachable from any tenant workload on affected hardware.","remediation":"Switch from LFENCE/JMP to retpoline or hardware IBRS/IBPB. On Linux, verify what is actually active by reading /sys/devices/system/cpu/vulnerabilities/spectre_v2 on your fleet - do not assume, check, because the string tells you exactly which mitigation the kernel selected. Changing it needs a kernel update and/or boot parameter change plus a reboot; on some platforms full IBRS also needs microcode from an SBIOS update. Retpoline and IBRS both cost performance relative to LFENCE/JMP, which is why the weaker option got chosen in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26401","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-46778","cve":"CVE-2021-46778","aliases":["SQUIP"],"title":"AMD Zen 1 / Zen 2 / Zen 3 - execution unit scheduler queue contention (SMT): MULTI-TENANT ISOLATION: AMD's split scheduler design gives each execution unit its own queue, and contention…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zen 1 / Zen 2 / Zen 3 - execution unit scheduler queue contention (SMT)","year":"2021","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: AMD's split scheduler design gives each execution unit its own queue, and contention on those queues is observable from the sibling SMT thread. An attacker running on one hardware thread measures scheduler pressure and reconstructs what the co-resident thread is computing - the SQUIP researchers recovered a full RSA-4096 private key from a victim on the sibling thread. This is a genuine cross-tenant confidentiality break with no memory access involved at all: nothing in your container, VM or cgroup boundary sees it happening, because no boundary is crossed in software.","attack_vector":"Local, unprivileged, requires SMT enabled and the attacker scheduled on the sibling thread of the victim's physical core. On a bin-packed cluster that co-residency happens by default - your scheduler arranges it for you. Affects Zen 1, Zen 2 and Zen 3.","remediation":"AMD's guidance is that software should use constant-time / secret-independent control flow rather than a microcode fix, so **treat this as effectively unpatchable in hardware**. The operator-side controls are the real answer: disable SMT on nodes that mix tenants, or enforce whole-core (not thread) allocation so a physical core is never shared across trust boundaries. Kubernetes operators can get this with the CPU manager's full-pcpus-only policy. Disabling SMT needs a reboot; core-pinning policy needs a kubelet restart and a drain. The zero-cost mitigation available today is scheduling policy rather than patching: these attacks need the attacker and victim co-resident on sibling SMT threads, so either disable SMT (costing roughly 10-25% throughput on most inference and training workloads) or enforce core isolation so no two tenants ever share a physical core. On a GPU fleet the CPU is rarely the bottleneck, which makes disabling SMT a cheaper trade than it looks on paper.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46778","https://stefangast.eu/papers/squip.pdf","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-23825","cve":"CVE-2022-23825","aliases":[],"title":"AMD CPU (Branch Type Confusion): Non-Retbleed branch type confusion - speculative cross-domain leak","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Branch Type Confusion)","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Non-Retbleed branch type confusion - speculative cross-domain leak","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation + reboot; standing perf cost","references":["https://access.redhat.com/security/cve/CVE-2022-23825"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-23960","cve":"CVE-2022-23960","aliases":["Spectre-BHB","Branch History Injection","BHI","TFV-9","CVE-2022-25368 (Ampere variant)"],"title":"Arm Cortex-A and Neoverse cores (Neoverse N1/N2/V1 among them); Trusted Firmware-A; also tracked by Ampere as AMP-SB-0001 for Altra, Altra Max and AmpereOne, and by CVE-2023-3006 for AmpereOne…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Cortex-A and Neoverse cores (Neoverse N1/N2/V1 among them); Trusted Firmware-A; also tracked by Ampere as…","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Branch history is shared across privilege boundaries, so an attacker in one context can poison the Branch History Buffer and make a victim context mispredict into a gadget that leaves the secret in the cache. Practically: a guest reads hypervisor memory, or one container reads another's, at side-channel bandwidth. Slow, but it does not crash anything and leaves no trace, which is exactly the profile that matters when the secret is a customer's model weights or an API key sitting in host memory on a shared Arm node.","attack_vector":"Unprivileged code in any guest or container on an affected Arm core that shares a physical core with the victim. Worse when SMT or aggressive core-sharing is enabled; a dedicated-core scheduling policy raises the bar considerably.","remediation":"Layered and none of it is optional. EL3 firmware exposes the mitigation via SMCCC (updated TF-A from the OEM, flash + reboot); the guest and host kernels must run the arm64 Spectre-BHB workarounds (loop or clearbhb sequence per core type). Newer cores get FEAT_CLEARBHB in hardware; older Neoverse N1 parts pay for a software loop on every exception entry, which is a real syscall-path cost. Ampere published AMP-SB-0001 for Altra/Altra Max/AmpereOne. The durable operator-level control is to stop co-tenanting untrusted workloads on one physical core.","references":["https://developer.arm.com/Arm%20Security%20Center/Speculative%20Processor%20Vulnerability","https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-9.html","https://amperecomputing.com/products/security-bulletins/impact-of-spectre-bhb-on-ampere.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-29900","cve":"CVE-2022-29900","aliases":[],"title":"AMD CPU (Retbleed): Retbleed: arbitrary speculative code execution via return instructions - cross-domain secret disclosure","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Retbleed)","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Retbleed: arbitrary speculative code execution via return instructions - cross-domain secret disclosure","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation (retbleed=) + reboot. Standing performance cost, historically severe on Zen 1/2 - measure before enabling fleet-wide","references":["https://access.redhat.com/security/cve/CVE-2022-29900"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-29901","cve":"CVE-2022-29901","aliases":[],"title":"Intel CPU (Retbleed): Retbleed on Intel - speculative execution of return instructions leaks across privilege boundaries","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Retbleed)","year":"2022","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Retbleed on Intel - speculative execution of return instructions leaks across privilege boundaries","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation + reboot; standing perf cost (IBRS on Skylake-era parts is expensive)","references":["https://access.redhat.com/security/cve/CVE-2022-29901"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-20569","cve":"CVE-2023-20569","aliases":[],"title":"AMD CPU (Inception / SRSO): Inception: Speculative Return Stack Overflow - attacker-controlled speculative disclosure across privilege…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (Inception / SRSO)","year":"2023","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Inception: Speculative Return Stack Overflow - attacker-controlled speculative disclosure across privilege domains on Zen 1-4","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation (`spec_rstack_overflow=`) + reboot. Standing perf cost - the safe-RET mitigation is measurable on syscall-heavy workloads","references":["https://access.redhat.com/security/cve/CVE-2023-20569"],"status":"curated","fleet":{"ubiquity":"very common - spans Zen 1 through Zen 4, i.e. essentially every AMD host CPU generation in service","remediation_pain":"microcode+reboot **and** a kernel update (SRSO safe-return / IBPB-on-entry); on Zen 1/2 the mitigation additionally requires **disabling SMT**, which permanently cuts logical core count on those nodes","pain_class":"microcode + reboot","why_fleet_wide":"Kernel-memory disclosure via speculative return redirection across the whole AMD lineup; the SMT-disable mitigation is a capacity hit the operator has to absorb fleet-wide, not a one-time reboot."}},{"id":"CVE-2023-25775","cve":"CVE-2023-25775","aliases":[],"title":"Intel irdma driver (Ethernet Controller RDMA for Linux): MULTI-TENANT ISOLATION: Improper access control in the Intel RDMA driver lets an unauthenticated user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel irdma driver (Ethernet Controller RDMA for Linux)","year":"2023","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper access control in the Intel RDMA driver lets an unauthenticated user escalate privilege. RDMA is the transport for collective operations across an AI cluster, and the driver maps queue pairs directly into user processes - so an access-control failure here is one workload reaching another's RDMA resources on the same host.","attack_vector":"Unauthenticated, which on an RDMA fabric means anything that can present traffic to the verbs interface. Treat the RDMA fabric as an authentication boundary that is not actually authenticating.","remediation":"Update the Intel irdma driver to 1.9.30 or later. Module reload drops RDMA connections and kills in-flight collectives - drain the node. No firmware component for this one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25775","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00794.html"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-3301","cve":"CVE-2023-3301","aliases":[],"title":"QEMU (net): Triggerable assertion via a race on NIC hot-unplug - guest can abort the host QEMU process","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU (net)","year":"2023","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Triggerable assertion via a race on NIC hot-unplug - guest can abort the host QEMU process","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart/live-migration","references":["https://access.redhat.com/security/cve/CVE-2023-3301"],"status":"curated"},{"id":"CVE-2024-28956","cve":"CVE-2024-28956","aliases":["XSA-469"],"title":"Xen / Intel CPU (ITS): Indirect Target Selection - speculative execution leak across privilege domains on Intel parts","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen / Intel CPU (ITS)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Indirect Target Selection - speculative execution leak across privilege domains on Intel parts","attack_vector":"Tenant VM guest; any tenant process in a container","remediation":"Microcode + hypervisor/kernel mitigation + reboot; standing perf cost","references":["https://xenbits.xen.org/xsa/advisory-469.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2024-33607","cve":"CVE-2024-33607","aliases":[],"title":"Intel TDX module: MULTI-TENANT ISOLATION: An out-of-bounds read in the TDX module reachable by an authenticated user, leaking…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds read in the TDX module reachable by an authenticated user, leaking information across the TDX boundary. Read primitives inside the module are the ones to watch: they are the shape that leaks other trust domains' state.","attack_vector":"An authenticated user on the host.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-33607","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01192.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-36350","cve":"CVE-2024-36350","aliases":[],"title":"AMD CPU (TSA-L1): Transient Scheduler Attack - store-to-load forwarding leak across contexts on Zen 3/4","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (TSA-L1)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Transient Scheduler Attack - store-to-load forwarding leak across contexts on Zen 3/4","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation (VERW on transitions) + reboot; standing perf cost. Also tracked as Xen XSA-471","references":["https://access.redhat.com/security/cve/CVE-2024-36350"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2024-36357","cve":"CVE-2024-36357","aliases":[],"title":"AMD CPU (TSA-SQ): Transient Scheduler Attack via the store queue - cross-context information leak","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD CPU (TSA-SQ)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Transient Scheduler Attack via the store queue - cross-context information leak","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + kernel mitigation + reboot; standing perf cost; consider disabling SMT for isolation-sensitive tenants","references":["https://access.redhat.com/security/cve/CVE-2024-36357"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2024-43420","cve":"CVE-2024-43420","aliases":[],"title":"Intel Atom processors (shared predictor transient execution): Shared microarchitectural predictor state influences transient execution on Atom parts, allowing local…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Atom processors (shared predictor transient execution)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Shared microarchitectural predictor state influences transient execution on Atom parts, allowing local information disclosure. Same advisory wave as the branch privilege injection work; relevant to Atom-based appliances inside the datacenter footprint.","attack_vector":"Local authenticated code on an affected Atom platform.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43420","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01247.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2024-45332","cve":"CVE-2024-45332","aliases":["Branch Privilege Injection","BPI","Branch Predictor Race Conditions"],"title":"Intel processors (indirect branch predictor race): MULTI-TENANT ISOLATION: Branch Privilege Injection: a race in how the indirect branch predictor associates…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (indirect branch predictor race)","year":"2024","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Branch Privilege Injection: a race in how the indirect branch predictor associates predictions with privilege level lets unprivileged code get its predictions applied in kernel context, reading kernel memory even on parts with hardware Spectre-v2 mitigations. The researchers demonstrated reading /etc/shadow on a fully patched machine - it defeats the mitigations operators had been told were sufficient.","attack_vector":"Local unprivileged code on the node - any container or VM.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. Confirm the specific microcode revision Intel names for your stepping; this one is not fully closed by kernel changes alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45332","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01247.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2025-20623","cve":"CVE-2025-20623","aliases":[],"title":"Intel Core processors, 10th generation (shared predictor state): Shared predictor state influencing transient execution on 10th-generation Core parts, leaking information to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Core processors, 10th generation (shared predictor state)","year":"2025","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Shared predictor state influencing transient execution on 10th-generation Core parts, leaking information to a local authenticated user. Matters where Core-class hardware runs management, build or edge-inference roles alongside the Xeon fleet.","attack_vector":"Local authenticated code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20623","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01247.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2025-24495","cve":"CVE-2025-24495","aliases":["Training Solo"],"title":"Intel Core Ultra processors (branch prediction unit initialisation): MULTI-TENANT ISOLATION: Part of the Training Solo family: incorrect initialisation of the branch prediction…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Core Ultra processors (branch prediction unit initialisation)","year":"2025","cvss_score":5.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Part of the Training Solo family: incorrect initialisation of the branch prediction unit lets an attacker self-train a predictor within the victim's own domain, so the leak works without the cross-domain training that existing mitigations assume. That is the significance - it sidesteps domain-isolation mitigations rather than defeating them head-on, and it reopens guest-to-host and user-to-kernel leakage on parts believed fixed.","attack_vector":"Local unprivileged code on an affected processor.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. Training Solo also needs kernel-side changes for the eBPF and indirect-branch paths; take both.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-24495","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01322.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2026-15788","cve":"CVE-2026-15788","aliases":[],"title":"BuildKit: NTFS junctions inside the cache root escape the cache mount on Windows container workers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":5.6,"severity":"medium","kev":false,"impact":"NTFS junctions inside the cache root escape the cache mount on Windows container workers","attack_vector":"Untrusted build author on a WCOW builder","remediation":"Upgrade BuildKit; avoid shared WCOW builders","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15788"],"status":"curated"},{"id":"CVE-2026-24198","cve":"CVE-2026-24198","aliases":[],"title":"GPU Display Driver: Improper access control in driver operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":5.6,"severity":"medium","kev":false,"impact":"Improper access control in driver operations","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24198","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-200"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-50195","cve":"CVE-2026-50195","aliases":[],"title":"containerd: CRI checkpoint import does not validate image references in checkpoint metadata","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":5.6,"severity":"medium","kev":false,"impact":"CRI checkpoint import does not validate image references in checkpoint metadata","attack_vector":"Malicious checkpoint image","remediation":"Rolling containerd upgrade with node drain; disable checkpoint import","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-50195"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2017-7262","cve":"CVE-2017-7262","aliases":[],"title":"AMD Ryzen with AGESA microcode - FMA3 instruction sequence hang: A long series of FMA3 instructions hangs the system. FMA3 is the fused multiply-add path that every dense…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Ryzen with AGESA microcode - FMA3 instruction sequence hang","year":"2017","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A long series of FMA3 instructions hangs the system. FMA3 is the fused multiply-add path that every dense linear-algebra kernel hammers, so on an ML host this is not an exotic instruction sequence - it is the normal workload. An unprivileged tenant running a numerical benchmark can hang the node, deliberately or by accident.","attack_vector":"Local, unprivileged. Reachable from any container running numerical code.","remediation":"Fixed by an AGESA microcode update from 2017 onward, delivered as an OEM BIOS package. Client Ryzen silicon rather than EPYC, so on a server fleet this is mostly a non-issue - verify your SKUs before spending a window on it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-7262"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-3639","cve":"CVE-2018-3639","aliases":["Spectre v4","SSB","Speculative Store Bypass"],"title":"Intel processors (speculative store bypass): Spectre v4: a load speculatively executes before an older store to the same address is resolved, so it reads…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (speculative store bypass)","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Spectre v4: a load speculatively executes before an older store to the same address is resolved, so it reads stale data. The exposure that matters is inside language runtimes and sandboxes where the attacker supplies the code being JITed - which describes most of a model-serving stack.","attack_vector":"Local code execution, including inside sandboxes and JIT runtimes.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The mitigation (SSBD) is opt-in per process on Linux because it costs measurable performance; decide deliberately whether your JIT-hosting workloads get it via prctl or whether you enable it system-wide.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3639","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00115.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-3689","cve":"CVE-2018-3689","aliases":[],"title":"Intel SGX Platform Software for Linux (AESM daemon): A local attacker can disable the AESM daemon, which is the component that services remote attestation. Kill…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX Platform Software for Linux (AESM daemon)","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A local attacker can disable the AESM daemon, which is the component that services remote attestation. Kill it and enclaves on that node can no longer prove what they are - so attestation-gated workloads stop being admitted.","attack_vector":"Local attacker on the node, no special privilege required.","remediation":"Upgrade SGX PSW for Linux to 2.1.102 or later, and monitor AESM liveness as a first-class signal on confidential-compute nodes. Userspace daemon update and restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3689","https://access.redhat.com/security/cve/CVE-2018-3689"],"status":"curated"},{"id":"CVE-2018-6252","cve":"CVE-2018-6252","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The escape handler exposes functionality that should never have shipped to production callers","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The escape handler exposes functionality that should never have shipped to production callers; an unprivileged process can reach it and crash the display driver.","attack_vector":"Any local user on the host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-6252"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6253","cve":"CVE-2018-6253","aliases":[],"title":"NVIDIA GPU Display Driver, DirectX and OpenGL user-mode drivers: A crafted pixel shader drives the user-mode driver into infinite recursion and kills the rendering process or…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, DirectX and OpenGL user-mode drivers","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A crafted pixel shader drives the user-mode driver into infinite recursion and kills the rendering process or the session. On a shared render farm or VDI host, one tenant's shader takes out their own session and can pin a core doing it.","attack_vector":"Anyone who can submit shaders - remote session users, tenant VMs, or local processes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://usn.ubuntu.com/3662-1/","https://www.talosintelligence.com/vulnerability_reports/TALOS-2018-0522","https://nvd.nist.gov/vuln/detail/CVE-2018-6253"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2018-6260","cve":"CVE-2018-6260","aliases":[],"title":"NVIDIA GPU Display Driver, GPU hardware performance counters: MULTI-TENANT ISOLATION: GPU performance counters are readable by any local user and leak enough about another…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver, GPU hardware performance counters","year":"2018","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: GPU performance counters are readable by any local user and leak enough about another process's GPU activity to reconstruct data it is processing - published work recovered neural-network structure and input images this way. On a shared GPU node, one tenant's job can profile a co-tenant's inference or training work through the counters. NVIDIA's answer was not a code fix but an access-control change, and the restriction was not on by default for a long time in several driver branches, so a large installed base ran exposed for years. Treat any node where profiling is unrestricted as offering no side-channel isolation between GPU contexts.","attack_vector":"Any local user or container that can open the GPU device and run a profiling-capable API (CUPTI, nvprof, Nsight) against the same physical GPU as the victim.","remediation":"Update to a driver branch that ships the profiling restriction, then verify it is actually enforced - this is the step operators skip. On Linux check the nvidia module parameter NVreg_RestrictProfilingToAdminUsers and set it to 1 in /etc/modprobe.d, then reload the module (node drain) or reboot; on Windows apply the driver update and confirm the developer-mode registry setting is not re-enabling profiling. On a multi-tenant node also stop mapping profiling-capable device nodes into tenant containers. Note this is a mitigation, not a fix: MPS and time-sliced sharing still leave the counters shared at the hardware level, so the only hard isolation is one tenant per physical GPU or MIG.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://usn.ubuntu.com/3904-1/","https://nvd.nist.gov/vuln/detail/CVE-2018-6260"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-0157","cve":"CVE-2019-0157","aliases":[],"title":"Intel SGX driver for Linux: Insufficient input validation in the out-of-tree SGX Linux driver lets a local authenticated user deny…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX driver for Linux","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Insufficient input validation in the out-of-tree SGX Linux driver lets a local authenticated user deny service. Relevant on hosts still running the legacy Intel SGX DKMS driver rather than the in-kernel driver.","attack_vector":"Local authenticated user with access to the SGX device node.","remediation":"Move to the in-kernel SGX driver where the kernel supports it, otherwise update the Intel SGX DKMS driver. Driver reload or reboot; no firmware or microcode.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00235.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-11090","cve":"CVE-2019-11090","aliases":["TPM-FAIL"],"title":"Intel PTT / fTPM (ECDSA and ECSchnorr timing): The firmware TPM's signing operation leaks nonce information through timing, letting an attacker recover the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel PTT / fTPM (ECDSA and ECSchnorr timing)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The firmware TPM's signing operation leaks nonce information through timing, letting an attacker recover the private key after observing a few hundred signatures. Intel PTT is the TPM on a large share of server boards where nobody fitted a discrete chip - so this is the default configuration, not an edge case. Recovered attestation keys let an attacker sign forged quotes and make a compromised node present as measured-clean.","attack_vector":"Local unprivileged user who can request signatures, or a network attacker where the TPM key backs a network-facing service such as a VPN or TLS client certificate. No physical access needed, which is what separated this from earlier TPM side channels.","remediation":"Intel CSME/PTT firmware update, delivered as a BIOS/ME package from the server OEM - per-node flash and reboot. Regenerate and re-enroll every key the fTPM signed with, because the firmware fix does not un-leak an already-extracted key. If attestation is load-bearing for your product, this is the argument for a discrete TPM over the chipset one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11090","https://tpm.fail/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-19077","cve":"CVE-2019-19077","aliases":[],"title":"Linux bnxt_re RoCE driver (bnxt_re_create_srq memory leak): A tenant can exhaust host memory by repeatedly triggering shared-receive-queue creation failures through the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (bnxt_re_create_srq memory leak)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A tenant can exhaust host memory by repeatedly triggering shared-receive-queue creation failures through the RDMA verbs interface. This is a straightforward noisy-neighbour denial of service available to any container or VM granted RDMA access, and it needs no privilege beyond opening verbs.","attack_vector":"Any local process with access to the RDMA verbs device — in practice, any tenant container given RDMA.","remediation":"Kernel upgrade plus host reboot. Independently: apply memory cgroup limits to RDMA-capable workloads and restrict verbs device access to workloads that actually need it — both container-runtime config changes, and both good practice regardless of this CVE.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-19077"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5671","cve":"CVE-2019-5671","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Resource leak in the escape handler - a local process can exhaust kernel resources and take the display…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Resource leak in the escape handler - a local process can exhaust kernel resources and take the display driver, and with it the node's GPUs, out of service.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["http://support.lenovo.com/us/en/solutions/LEN-26250","https://nvd.nist.gov/vuln/detail/CVE-2019-5671"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5677","cve":"CVE-2019-5677","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Out-of-bounds read in the DeviceIoControl handler - reads past the target buffer and crashes the driver","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the DeviceIoControl handler - reads past the target buffer and crashes the driver.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5677"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5686","cve":"CVE-2019-5686","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): The escape handler relies on invariants that are not guaranteed, so a local caller can drive the kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The escape handler relies on invariants that are not guaranteed, so a local caller can drive the kernel driver into an unsupported state and crash the node.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://support.lenovo.com/us/en/product_security/LEN-28096","https://nvd.nist.gov/vuln/detail/CVE-2019-5686"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5693","cve":"CVE-2019-5693","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Uninitialised pointer use in nvlddmkm.sys. Local denial of service","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Uninitialised pointer use in nvlddmkm.sys. Local denial of service; uninitialised-memory bugs of this shape are often better than their rating suggests.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5693"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-5696","cve":"CVE-2019-5696","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: a guest VM hands the vGPU Manager a wrongly sized buffer and drives the host into an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2019","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: a guest VM hands the vGPU Manager a wrongly sized buffer and drives the host into an out-of-bounds GPU access. One tenant VM can take down the vGPU host, which means every other tenant sharing that physical GPU goes with it.","attack_vector":"Any unprivileged user inside any guest VM assigned a vGPU on the host.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5696"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-0543","cve":"CVE-2020-0543","aliases":[],"title":"Intel CPU (SRBDS / CrossTalk): Special Register Buffer Data Sampling - leaks RDRAND/RDSEED output across cores, i.e. across tenants on the…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (SRBDS / CrossTalk)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Special Register Buffer Data Sampling - leaks RDRAND/RDSEED output across cores, i.e. across tenants on the same socket","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Microcode + reboot; the mitigation serialises RDRAND/RDSEED with a large measured cost on RNG-heavy workloads","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0543"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2020-0548","cve":"CVE-2020-0548","aliases":["Vector Register Sampling","VRS"],"title":"Intel processors (vector register sampling): Stale values left in vector registers can be sampled by other contexts, leaking data across the process and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (vector register sampling)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Stale values left in vector registers can be sampled by other contexts, leaking data across the process and enclave boundary. Lower yield than the fill-buffer attacks but it targets the vector registers - which is where floating-point tensor data lives on an AI node.","attack_vector":"Local code on the same physical core as the victim.","remediation":"Microcode update plus the OS buffer-clearing mitigation, then reboot. Microcode is late-loadable at boot; no OEM BIOS release strictly needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0548","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00329.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2020-0549","cve":"CVE-2020-0549","aliases":["CacheOut","L1DES","SGAxe"],"title":"Intel processors (L1D eviction sampling) / SGX attestation keys: MULTI-TENANT ISOLATION: Stale data can be sampled out of L1D fill buffers during cache-line eviction, leaking…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (L1D eviction sampling) / SGX attestation keys","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Stale data can be sampled out of L1D fill buffers during cache-line eviction, leaking across privilege and enclave boundaries. The consequence operators care about is the SGAxe result built on it: recovery of the machine's SGX attestation keys, which lets an attacker produce quotes that pass Intel's attestation service for a machine whose enclaves are entirely under their control. Once that happens, remote attestation stops proving anything about that platform.","attack_vector":"Local code on the same core as the victim; with SMT enabled, a sibling-thread co-tenant.","remediation":"Microcode update, which is late-loadable at boot without an OEM BIOS release, plus a TCB recovery and re-attestation. Also disable SMT or enforce core scheduling on nodes serving untrusted tenants. Any attestation key material provisioned before the microcode update must be treated as compromised - the fix does not revoke it for you.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-0549","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00329.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2020-12966","cve":"CVE-2020-12966","aliases":[],"title":"AMD EPYC SEV-ES / SEV-SNP - information disclosure: MULTI-TENANT ISOLATION: An information-disclosure flaw in SEV-ES and SEV-SNP on EPYC lets a locally…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD EPYC SEV-ES / SEV-SNP - information disclosure","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An information-disclosure flaw in SEV-ES and SEV-SNP on EPYC lets a locally authenticated attacker recover data that the encrypted-state protections were meant to keep opaque. For an operator this is a confidentiality gap in the feature you are charging for when you sell confidential VMs on EPYC.","attack_vector":"Local, authenticated attacker on the host.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-12966","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-24502","cve":"CVE-2020-24502","aliases":[],"title":"Intel E810 adapter driver for Linux (< 1.0.4): Early E810 Linux driver flaw (improper input validation) reachable by an authenticated local user. Present on…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 adapter driver for Linux (< 1.0.4)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Early E810 Linux driver flaw (improper input validation) reachable by an authenticated local user. Present on nodes still running the original E810 driver, which happens when the out-of-tree driver was pinned at deployment and never revisited.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24502","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00462.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-24503","cve":"CVE-2020-24503","aliases":[],"title":"Intel E810 adapter driver for Linux (< 1.0.4): Early E810 Linux driver flaw (insufficient access control leading to information disclosure) reachable by an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 adapter driver for Linux (< 1.0.4)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Early E810 Linux driver flaw (insufficient access control leading to information disclosure) reachable by an authenticated local user. Present on nodes still running the original E810 driver, which happens when the out-of-tree driver was pinned at deployment and never revisited.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24503","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00462.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-24504","cve":"CVE-2020-24504","aliases":[],"title":"Intel E810 adapter driver for Linux (< 1.0.4): Early E810 Linux driver flaw (uncontrolled resource consumption) reachable by an authenticated local user.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 adapter driver for Linux (< 1.0.4)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Early E810 Linux driver flaw (uncontrolled resource consumption) reachable by an authenticated local user. Present on nodes still running the original E810 driver, which happens when the out-of-tree driver was pinned at deployment and never revisited.","attack_vector":"Authenticated local user on the node.","remediation":"Fixed in the Intel out-of-tree ice driver (or the equivalent in-kernel version). Updating the driver requires unloading and reloading the module, which drops every link on that NIC - on a node whose RDMA fabric carries collective traffic, that is a job-killing event, so drain first. If you take it via a distro kernel update instead, it is a reboot. No firmware flash for the driver-side fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24504","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00462.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-25596","cve":"CVE-2020-25596","aliases":[],"title":"Xen - x86 PV guest denial of service via SYSENTER: SYSENTER leaves state sanitisation to software, and on AMD hardware Xen's PV guest handling got it wrong…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen - x86 PV guest denial of service via SYSENTER","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"SYSENTER leaves state sanitisation to software, and on AMD hardware Xen's PV guest handling got it wrong, letting a PV guest kernel deny service to itself and destabilise the host path. Legacy PV territory, included for completeness of the Xen-on-AMD picture.","attack_vector":"From inside an x86 PV guest.","remediation":"Fixed in Xen (XSA-339). Hypervisor update plus host reboot. The durable answer is to stop running PV guests - HVM/PVH is the supported path and carries less of this legacy surface.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-25596","https://xenbits.xen.org/xsa/advisory-339.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5959","cve":"CVE-2020-5959","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated guest-supplied index in the vGPU plugin. A tenant VM crashes the host…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated guest-supplied index in the vGPU plugin. A tenant VM crashes the host plugin and takes out the GPU for every co-tenant.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5959"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5960","cve":"CVE-2020-5960","aliases":[],"title":"NVIDIA vGPU Manager kernel module (nvidia.ko, host): MULTI-TENANT ISOLATION: NULL dereference in the host-side nvidia.ko under the vGPU Manager. A guest can panic…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager kernel module (nvidia.ko, host)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: NULL dereference in the host-side nvidia.ko under the vGPU Manager. A guest can panic the hypervisor host's kernel module, killing every VM sharing that GPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5960"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5961","cve":"CVE-2020-5961","aliases":[],"title":"NVIDIA vGPU guest graphics driver: Bad cleanup on a failure path in the guest driver crashes the tenant's own VM. Self-inflicted blast radius…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU guest graphics driver","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Bad cleanup on a failure path in the guest driver crashes the tenant's own VM. Self-inflicted blast radius only, but it is a reliable way for a tenant to lose their own workload and blame the platform.","attack_vector":"A user inside the guest VM (or any workload that triggers the failure path).","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5961"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5965","cve":"CVE-2020-5965","aliases":[],"title":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll): A crafted shader causes an out-of-bounds access in the DX11 user-mode driver and kills the rendering process.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver, DirectX 11 user-mode driver (nvwgf2um.dll)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A crafted shader causes an out-of-bounds access in the DX11 user-mode driver and kills the rendering process. Anywhere tenants supply shaders - VDI, cloud gaming, remote rendering - a tenant can knock over their own and neighbouring sessions.","attack_vector":"Anyone who can submit a shader to the host, including from inside a guest VM or a remote session.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5965"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5986","cve":"CVE-2020-5986","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: guest-supplied size not validated in the vGPU plugin","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: guest-supplied size not validated in the vGPU plugin; a tenant tampers with host state or crashes the shared GPU. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5986"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5989","cve":"CVE-2020-5989","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: NULL dereference in the vGPU plugin reachable from a guest - one tenant crashes the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: NULL dereference in the vGPU plugin reachable from a guest - one tenant crashes the plugin and the shared GPU. vGPU 8.x before 8.5, 10.x before 10.4, and 11.0.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5989"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-8557","cve":"CVE-2020-8557","aliases":[],"title":"Kubernetes (kubelet): Pod writes to its own /etc/hosts unaccounted for in eviction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Pod writes to its own /etc/hosts unaccounted for in eviction; fills node disk","attack_vector":"Any tenant workload","remediation":"Rolling kubelet upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8557"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2020-8698","cve":"CVE-2020-8698","aliases":["Fast Store Forwarding Predictor","FSFP"],"title":"Intel processors (fast store forwarding predictor): Improper isolation of a shared microarchitectural resource lets a local authenticated user infer data from…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (fast store forwarding predictor)","year":"2020","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper isolation of a shared microarchitectural resource lets a local authenticated user infer data from another context. Part of the steady drip of shared-predictor leaks that each cost another microcode revision and another small performance tax.","attack_vector":"Local authenticated code on the host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8698","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00381"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2021-0145","cve":"CVE-2021-0145","aliases":[],"title":"Intel processors (fast store forwarding predictor initialisation): Improper initialisation of a shared predictor resource allows a local authenticated user to infer data across…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (fast store forwarding predictor initialisation)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper initialisation of a shared predictor resource allows a local authenticated user to infer data across contexts. Same family as the earlier fast-store-forwarding issue.","attack_vector":"Local authenticated code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0145","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00561.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2021-1053","cve":"CVE-2021-1053","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): Improper validation of a user-supplied pointer in the kernel-mode layer. An unprivileged local caller or GPU…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper validation of a user-supplied pointer in the kernel-mode layer. An unprivileged local caller or GPU container crashes the driver and takes every GPU job on the node with it.","attack_vector":"Any local user or GPU container with access to the NVIDIA device nodes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1053"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1054","cve":"CVE-2021-1054","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Missing authorization check in the escape handler - an unprivileged caller performs an action the driver…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing authorization check in the escape handler - an unprivileged caller performs an action the driver should have refused, resulting in denial of service.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1054"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1066","cve":"CVE-2021-1066","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: unvalidated guest input causes unbounded resource consumption on the host, so one…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: unvalidated guest input causes unbounded resource consumption on the host, so one tenant starves the vGPU host and denies service to co-tenants. vGPU 8.x before 8.6, 11.0 before 11.3.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1066"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1078","cve":"CVE-2021-1078","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in nvlddmkm.sys leading to a system crash - unprivileged local caller kills the whole node…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL dereference in nvlddmkm.sys leading to a system crash - unprivileged local caller kills the whole node and every job on it.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1078"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1087","cve":"CVE-2021-1087","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: information retrievable by a guest that defeats ASLR on the host side. On its own it…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: information retrievable by a guest that defeats ASLR on the host side. On its own it does nothing; combined with any of the memory-corruption bugs in the same bulletin it is what makes a guest-to-host escape reliable instead of a coin flip. vGPU 12.x before 12.2, 11.x before 11.4, 8.x before 8.7.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1087"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1095","cve":"CVE-2021-1095","aliases":[],"title":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko): All control calls with embedded parameters dereference an untrusted pointer - not one handler but the whole…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver (Windows nvlddmkm.sys + Linux nvidia.ko)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"All control calls with embedded parameters dereference an untrusted pointer - not one handler but the whole family of them. Unprivileged local caller crashes the driver on Windows or Linux.","attack_vector":"Any local user or GPU container with access to the NVIDIA device nodes.","remediation":"Install the fixed GPU Display Driver branch on both Windows and Linux nodes. The kernel component (nvlddmkm.sys / nvidia.ko) cannot be hot-swapped under load, so this is a node drain and reboot per host; restart the container runtime afterwards so mounted driver libraries match the kernel module. No VBIOS or BMC flash.","references":["https://lists.debian.org/debian-lts-announce/2022/01/msg00013.html","https://security.gentoo.org/glsa/202310-02","https://nvd.nist.gov/vuln/detail/CVE-2021-1095"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1096","cve":"CVE-2021-1096","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): NULL dereference in the escape handler causing a system crash. Any local process with GPU access can reboot…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL dereference in the escape handler causing a system crash. Any local process with GPU access can reboot the node.","attack_vector":"Any local user with GPU device access on a Windows host.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1096"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1101","cve":"CVE-2021-1101","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: NULL dereference in the vGPU plugin, guest-reachable","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: NULL dereference in the vGPU plugin, guest-reachable; one tenant crashes the shared GPU. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1101"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1102","cve":"CVE-2021-1102","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: guest input drives the vGPU plugin into a floating-point exception and takes the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: guest input drives the vGPU plugin into a floating-point exception and takes the shared GPU down. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1102"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1116","cve":"CVE-2021-1116","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): A NULL pointer created in user-mode code is dereferenced in the kernel, crashing the system. Unprivileged…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer created in user-mode code is dereferenced in the kernel, crashing the system. Unprivileged local user takes the node down.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1116"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1121","cve":"CVE-2021-1121","aliases":[],"title":"NVIDIA vGPU Manager kernel module (nvidia.ko, host): MULTI-TENANT ISOLATION: one vGPU can starve the other vGPUs hosted on the same physical GPU of resources.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager kernel module (nvidia.ko, host)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: one vGPU can starve the other vGPUs hosted on the same physical GPU of resources. This is the noisy-neighbour problem as a security bug: a tenant deliberately degrades every co-tenant on the card, and nothing in the vGPU scheduler stops them.","attack_vector":"Any user inside a guest VM sharing a physical GPU with other tenants.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1121"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1122","cve":"CVE-2021-1122","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: guest-reachable NULL dereference in the vGPU plugin causing denial of service across…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: guest-reachable NULL dereference in the vGPU plugin causing denial of service across the shared GPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1122"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1123","cve":"CVE-2021-1123","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: a guest can deadlock the vGPU plugin. A hung plugin does not crash-and-restart…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: a guest can deadlock the vGPU plugin. A hung plugin does not crash-and-restart cleanly - it wedges the GPU for every tenant on it until the host is reset.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1123"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-26320","cve":"CVE-2021-26320","aliases":[],"title":"AMD SEV firmware - ASK validation in SEND_START: Insufficient validation of the AMD SEV Signing Key in the SEND_START command lets a locally authenticated…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - ASK validation in SEND_START","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Insufficient validation of the AMD SEV Signing Key in the SEND_START command lets a locally authenticated attacker wedge the PSP. SEND_START is part of the guest migration/export flow, so a tenant-triggered migration path can take the secure processor out - and with the PSP down, every confidential guest on the node loses its attestation and key services.","attack_vector":"Local, authenticated. Exercised through the SEV guest export path.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26320","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26333","cve":"CVE-2021-26333","aliases":[],"title":"AMD PSP chipset driver - permissive device DACL: The PSP chipset driver's discretionary access control list lets low-privileged users open a handle to the PSP…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD PSP chipset driver - permissive device DACL","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The PSP chipset driver's discretionary access control list lets low-privileged users open a handle to the PSP device. Once open, an unprivileged process can query the driver and read back information it should not have - and, more importantly, it has a handle on the secure processor interface that every other PSP bug then becomes reachable through. Weak device permissions are what turn 'requires privilege' into 'requires a shell'.","attack_vector":"Local, unprivileged - which is the notable part. Worth explicitly checking on any node where you hand semi-trusted workloads local execution.","remediation":"Fixed by updating the AMD PSP chipset driver package and reloading it or rebooting. Driver-speed rather than BIOS-speed, so this is one you can actually close quickly. Audit the permissions on your PSP/ccp device node as part of the same pass - a device that unprivileged users can open is a standing invitation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26333","https://www.amd.com/en/resources/product-security.html"],"status":"curated"},{"id":"CVE-2021-26339","cve":"CVE-2021-26339","aliases":[],"title":"AMD CPU core logic - core hang triggered from an unprivileged VM: MULTI-TENANT ISOLATION: Specific code executed from an unprivileged VM can hang an AMD CPU core outright. A…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU core logic - core hang triggered from an unprivileged VM","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Specific code executed from an unprivileged VM can hang an AMD CPU core outright. A guest wedges a physical core on the host, denying it to every other workload scheduled there - and on a GPU node whose CPU cores feed data to accelerators, losing cores starves the GPUs. Availability attack from a tenant against the host, requiring no privilege.","attack_vector":"From inside an unprivileged guest VM. No escalation needed - the guest simply executes a particular sequence.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26339","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26349","cve":"CVE-2021-26349","aliases":[],"title":"AMD SEV-SNP migration agent (report ID assignment): MULTI-TENANT ISOLATION: An imported SEV-SNP guest is not assigned a fresh report ID, so the guest can be…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP migration agent (report ID assignment)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An imported SEV-SNP guest is not assigned a fresh report ID, so the guest can be tricked into trusting a dishonest Migration Agent. Live migration is the moment a confidential VM is most exposed - it has to hand its state to something - and this lets a malicious host present an MA the guest will accept. The guest then migrates its secrets into an attacker-controlled destination believing it is talking to a legitimate peer.","attack_vector":"Requires a malicious or compromised hypervisor driving guest migration. Only exercised if you actually use SEV-SNP live migration.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. If you do not offer live migration for confidential VMs, the path is not exercised and this can sit in the normal patch queue. If you do, it is a top-of-queue item, and worth reviewing whether migration should be disabled for confidential tenants until the fleet is patched.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26349","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26361","cve":"CVE-2021-26361","aliases":[],"title":"AMD AGESA Boot Loader (ABL) / ASP stage-2 bootloader: MULTI-TENANT ISOLATION: A malicious or compromised User Application or AGESA Boot Loader can exfiltrate…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD AGESA Boot Loader (ABL) / ASP stage-2 bootloader","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A malicious or compromised User Application or AGESA Boot Loader can exfiltrate arbitrary memory from the ASP stage-2 bootloader. What leaks is firmware memory at the deepest pre-boot stage - the material an attacker needs to build a reliable secure-processor exploit, and potentially key state handled during early boot.","attack_vector":"Local, requires control of a UApp or the ABL itself, so firmware-level access rather than OS-level.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26361","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-26393","cve":"CVE-2021-26393","aliases":[],"title":"AMD Secure Processor TEE - memory cleanup between trusted applications: MULTI-TENANT ISOLATION: The ASP's trusted execution environment fails to scrub memory between uses, so an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor TEE - memory cleanup between trusted applications","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The ASP's trusted execution environment fails to scrub memory between uses, so an attacker able to get a validly signed trusted application loaded can read or poison residue left by a previous TA. On a platform where the ASP handles fTPM state and SEV key material, that residue is exactly the material you least want leaking sideways.","attack_vector":"Local and privileged: the attacker must be able to produce and load a validly signed trusted application, which normally means a signing-key or supply-chain compromise rather than plain root.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26393","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-27562","cve":"CVE-2021-27562","aliases":[],"title":"Arm Trusted Firmware-M: Non-secure world can halt the system, overwrite secure data, or leak secure data via the NSPE handler —…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-M","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Non-secure world can halt the system, overwrite secure data, or leak secure data via the NSPE handler — relevant to BMC SoCs and DPUs built on Arm TrustZone","attack_vector":"Local","remediation":"Firmware update of the affected Arm-based management controller; on a BMC this is again an ODM-gated rebase","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27562"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-28689","cve":"CVE-2021-28689","aliases":[],"title":"Xen on x86 - speculative vulnerabilities with bare 32-bit PV guests: MULTI-TENANT ISOLATION: Bare (non-shim) 32-bit PV guests run in ring 1, an arrangement that leaves them…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on x86 - speculative vulnerabilities with bare 32-bit PV guests","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Bare (non-shim) 32-bit PV guests run in ring 1, an arrangement that leaves them exposed to speculative-execution attacks against the hypervisor and other guests. Listed here because the mitigation story on AMD hardware differs from Intel's and gets overlooked - if you still run 32-bit PV guests anywhere, this is a standing cross-guest speculative exposure.","attack_vector":"From inside a 32-bit PV guest.","remediation":"Fixed in Xen (XSA-370) by running 32-bit PV guests under the PV shim rather than bare. Update Xen, switch affected guests to shim mode, and reboot. The durable answer is to stop running 32-bit PV guests at all - on a modern AI fleet there is no reason to have any.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-28689","https://xenbits.xen.org/xsa/advisory-370.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-33096","cve":"CVE-2021-33096","aliases":["INTEL-SA-00571","CVE-2021-33061"],"title":"Intel 82599 Ethernet Controllers and Adapters - network-on-chip shared-resource isolation: Improper isolation of shared resources inside the controller's network-on-chip lets an authenticated user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel 82599 Ethernet Controllers and Adapters - network-on-chip shared-resource isolation","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper isolation of shared resources inside the controller's network-on-chip lets an authenticated user cause denial of service. This is the multi-tenant failure mode that matters for SR-IOV: one function on the adapter - a virtual function assigned to a tenant VM - can exhaust or wedge a resource shared with the physical function and the other VFs. A tenant who is only supposed to control their own VF takes down networking for every other tenant sharing that adapter, and for the host. Intel documented no firmware fix for the 82599 here, which makes it a design limit of the part rather than a bug you close.","attack_vector":"An authenticated local user on any function of the adapter - concretely, a tenant VM that has been assigned an SR-IOV virtual function on a shared 82599.","remediation":"Intel's guidance for the 82599 is mitigation, not a firmware fix: do not share a single adapter's virtual functions across mutually untrusted tenants. Practically that means either dedicating a physical adapter per tenant, moving multi-tenant workloads off 82599-class parts onto controllers with stronger VF isolation, or accepting the shared-fate risk and rate-limiting at the switch. If your multi-tenant story depends on SR-IOV VF isolation, this CVE is the reason to test that assumption on your actual silicon rather than reading it off a datasheet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33096","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00571.html","https://security.netapp.com/advisory/ntap-20220210-0010/"],"status":"curated"},{"id":"CVE-2021-33135","cve":"CVE-2021-33135","aliases":[],"title":"Intel SGX Linux kernel driver: Uncontrolled resource consumption in the in-kernel SGX driver lets a local authenticated user exhaust EPC or…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX Linux kernel driver","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Uncontrolled resource consumption in the in-kernel SGX driver lets a local authenticated user exhaust EPC or driver resources and deny SGX to everyone else on the node.","attack_vector":"Local authenticated user with SGX device access - on a confidential-compute node, any tenant.","remediation":"Kernel update and reboot. Kernel-only, no firmware.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33135","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00603.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-4453","cve":"CVE-2021-4453","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: fix a potential gpu_metrics_table memory leak","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-4453","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47042","cve":"CVE-2021-47042","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Free local data after use","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47042","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47253","cve":"CVE-2021-47253","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix potential memory leak in DMUB hw_init","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47253","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47362","cve":"CVE-2021-47362","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: Update intermediate power state for SI","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47362","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47410","cve":"CVE-2021-47410","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: fix svm_migrate_fini warning","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47410","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47420","cve":"CVE-2021-47420","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: fix a potential ttm->sg memory leak","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47420","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47431","cve":"CVE-2021-47431","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix gart.bo pin_count leak","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47431","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47550","cve":"CVE-2021-47550","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu): A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/amdgpu: fix potential memleak","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47550","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-47658","cve":"CVE-2021-47658","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2021","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: fix a potential gpu_metrics_table memory leak","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47658","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-0171","cve":"CVE-2022-0171","aliases":[],"title":"Linux KVM SEV API - host kernel crash from unprivileged guest creation: A non-root host user-level application can crash the host kernel simply by creating a confidential guest…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM SEV API - host kernel crash from unprivileged guest creation","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A non-root host user-level application can crash the host kernel simply by creating a confidential guest through the KVM SEV API. Denial of service against the whole node from an unprivileged local process - on a GPU host that means every training job on the box dies because somebody with shell access called an ioctl.","attack_vector":"Local, **unprivileged** - the notable part. Only needs access to /dev/kvm, which on many hosts is more widely granted than people assume.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. Also audit who has /dev/kvm on your GPU hosts; if nothing on the node runs VMs, the device should not be world-accessible.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0171"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-0563","cve":"CVE-2022-0563","aliases":[],"title":"util-linux (chfn/chsh): Partial disclosure of arbitrary files via libreadline in setuid chfn/chsh","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"util-linux (chfn/chsh)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Partial disclosure of arbitrary files via libreadline in setuid chfn/chsh","attack_vector":"Local user","remediation":"Package update; no reboot","references":["https://access.redhat.com/security/cve/CVE-2022-0563"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-20235","cve":"CVE-2022-20235","aliases":["PowerVR information page"],"title":"Imagination PowerVR GPU driver - cache subsystem information page: The driver's cache-subsystem information page, intended to be writable only by the driver, was mapped…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Imagination PowerVR GPU driver - cache subsystem information page","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The driver's cache-subsystem information page, intended to be writable only by the driver, was mapped writable to userspace before DDK 1.18. A tenant process could alter driver cache state. Included as another instance of the GPU-driver-maps-too-much class that recurs across vendors.","attack_vector":"Unprivileged local application holding the GPU device node.","remediation":"Update to Imagination DDK 1.18 or later. Not a datacenter part; treat as vendor-evaluation intelligence rather than a fleet action.","references":["https://source.android.com/security/bulletin/2022-08-01","https://nvd.nist.gov/vuln/detail/CVE-2022-20235"],"status":"curated"},{"id":"CVE-2022-21125","cve":"CVE-2022-21125","aliases":["SBDS","MMIO Stale Data"],"title":"Intel processors (shared buffers data sampling): MULTI-TENANT ISOLATION: Incomplete cleanup of microarchitectural fill buffers lets a local user sample data…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (shared buffers data sampling)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Incomplete cleanup of microarchitectural fill buffers lets a local user sample data left behind by other contexts. Part of the MMIO stale-data cluster, whose distinguishing feature is that a guest can pull data across the VM boundary through device MMIO accesses - relevant on any node passing devices through to tenants, which describes every GPU node.","attack_vector":"Local authenticated code, including inside a guest with a passed-through device.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The hypervisor half of the mitigation matters here: make sure your VMM's stale-data mitigations are enabled, not just the host microcode. On nodes that host untrusted co-tenants, also disable SMT or enforce core scheduling; that costs real throughput and is a capacity-planning decision, not a free toggle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21125","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00615.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-2153","cve":"CVE-2022-2153","aliases":[],"title":"KVM: NULL pointer dereference in kvm_irq_delivery_to_apic_fast() - guest crashes the host","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"KVM","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"NULL pointer dereference in kvm_irq_delivery_to_apic_fast() - guest crashes the host","attack_vector":"Tenant VM guest","remediation":"Kernel patch; KVM-module scope usually forces drain + reboot rather than livepatch","references":["https://access.redhat.com/security/cve/CVE-2022-2153"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-21815","cve":"CVE-2022-21815","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): A null pointer dereference created from user mode inside the kernel driver bluescreens the node. Only matters…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null pointer dereference created from user mode inside the kernel driver bluescreens the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5312. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21815","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-21816","cve":"CVE-2022-21816","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A user inside a guest VM triggers a GPU interrupt storm that lands on the hypervisor…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A user inside a guest VM triggers a GPU interrupt storm that lands on the hypervisor host, wedging the physical GPU. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. One tenant can take down every other tenant sharing that GPU with no memory-corruption skill required - it is a pure availability attack that any guest user can run.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5312. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21816","https://github.com/NVIDIA/product-security/tree/main/2022/5312"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-284"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-26373","cve":"CVE-2022-26373","aliases":["PBRSB","Post-barrier Return Stack Buffer"],"title":"Intel processors (post-barrier return stack buffer): MULTI-TENANT ISOLATION: PBRSB: return predictions made after an IBPB barrier can still use pre-barrier state…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (post-barrier return stack buffer)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: PBRSB: return predictions made after an IBPB barrier can still use pre-barrier state, so the barrier that hypervisors rely on to separate guests does not fully separate them. The specific worry is a guest reading host memory on a machine where the operator believed IBPB closed that door.","attack_vector":"Local code in a guest or unprivileged context on an affected host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect. The mitigation also requires a hypervisor/kernel change that stuffs the RSB after VM exit - patch both, and confirm through the spectre_v2 sysfs entry.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-26373","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00706.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-28187","cve":"CVE-2022-28187","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): The kernel mode layer fails to release a resource after its lifetime ends","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The kernel mode layer fails to release a resource after its lifetime ends; a local user leaks it until the node falls over. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28187","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-772"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-28188","cve":"CVE-2022-28188","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper input validation in the DxgkDdiEscape handler ends in a node crash. Only matters to you if you run…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper input validation in the DxgkDdiEscape handler ends in a node crash. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28188","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-28189","cve":"CVE-2022-28189","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): A null-pointer dereference in the DxgkDdiEscape handler bluescreens the node. Only matters to you if you run…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the DxgkDdiEscape handler bluescreens the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28189","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-28190","cve":"CVE-2022-28190","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper input validation in the DxgkDdiEscape handler gives a local user a denial of service. Only matters…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper input validation in the DxgkDdiEscape handler gives a local user a denial of service. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5353. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28190","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-28191","cve":"CVE-2022-28191","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: An unprivileged guest user drives uncontrolled resource consumption in the host vGPU…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An unprivileged guest user drives uncontrolled resource consumption in the host vGPU Manager until the host driver stops serving. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. This is the classic vGPU noisy-neighbour-as-attack: no exploit engineering, just resource exhaustion crossing from guest to host.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5353. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28191","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-28224","cve":"CVE-2022-28224","aliases":[],"title":"Calico: Route hijacking via the floating IP feature","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Calico","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Route hijacking via the floating IP feature","attack_vector":"Privileged cluster user","remediation":"Rolling Calico upgrade; disable floating IPs for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28224"],"status":"curated"},{"id":"CVE-2022-31030","cve":"CVE-2022-31030","aliases":[],"title":"containerd: Unbounded memory consumption in containerd daemon via repeated ExecSync","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unbounded memory consumption in containerd daemon via repeated ExecSync; node DoS","attack_vector":"Any tenant workload / anyone with kube API exec rights","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31030"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31615","cve":"CVE-2022-31615","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A basic local user triggers a null-pointer dereference in the kernel mode layer and panics the node.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A basic local user triggers a null-pointer dereference in the kernel mode layer and panics the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5383. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31615","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-31618","cve":"CVE-2022-31618","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A null-pointer dereference in the vGPU plugin lets a guest crash the host GPU stack.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A null-pointer dereference in the vGPU plugin lets a guest crash the host GPU stack. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5383. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-31618","https://github.com/NVIDIA/product-security/tree/main/2022/5383"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34675","cve":"CVE-2022-34675","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: The Virtual GPU Manager ignores a return value and dereferences null, giving a guest…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The Virtual GPU Manager ignores a return value and dereferences null, giving a guest a host-side denial of service. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5415. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34675","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34677","cve":"CVE-2022-34677","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user forces an integer truncation in the kernel handler, producing a crash or silent data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unprivileged user forces an integer truncation in the kernel handler, producing a crash or silent data corruption. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34677","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34679","cve":"CVE-2022-34679","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unhandled return value leads to a null-pointer dereference in the kernel handler, crashing the node from…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unhandled return value leads to a null-pointer dereference in the kernel handler, crashing the node from an unprivileged account. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34679","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34680","cve":"CVE-2022-34680","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Integer truncation leads to an out-of-bounds read in the kernel handler and a node crash. Everything with a…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Integer truncation leads to an out-of-bounds read in the kernel handler and a node crash. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34680","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-197"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34681","cve":"CVE-2022-34681","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Improper validation of a display-related data structure in the kernel handler crashes the node. Only matters…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Improper validation of a display-related data structure in the kernel handler crashes the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34681","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34682","cve":"CVE-2022-34682","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An unprivileged user triggers a null-pointer dereference in the kernel mode layer and panics the node.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An unprivileged user triggers a null-pointer dereference in the kernel mode layer and panics the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34682","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34683","cve":"CVE-2022-34683","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): A null-pointer dereference in the DxgkDdiEscape handler crashes the node. Only matters to you if you run…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the DxgkDdiEscape handler crashes the node. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34683","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-36056","cve":"CVE-2022-36056","aliases":[],"title":"cosign / sigstore: Multiple verify-blob flaws cause successful verification of unsigned or wrongly-signed artifacts","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Multiple verify-blob flaws cause successful verification of unsigned or wrongly-signed artifacts","attack_vector":"Malicious artifact","remediation":"Upgrade cosign; re-verify","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36056"],"status":"curated"},{"id":"CVE-2022-42266","cve":"CVE-2022-42266","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): The DxgkDdiEscape handler exposes information to a caller that should not have it - limited kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The DxgkDdiEscape handler exposes information to a caller that should not have it - limited kernel information disclosure to any local user. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5415. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42266","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42703","cve":"CVE-2022-42703","aliases":[],"title":"Linux kernel (mm anon_vma): Use-after-free from leaf anon_vma double reuse - memory corruption / privesc primitive","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (mm anon_vma)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Use-after-free from leaf anon_vma double reuse - memory corruption / privesc primitive","attack_vector":"Local user / tenant process","remediation":"Livepatchable; otherwise drain + reboot","references":["https://access.redhat.com/security/cve/CVE-2022-42703"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-43309","cve":"CVE-2022-43309","aliases":[],"title":"Supermicro X11SSL-CF hardware revision 1.01, BMC firmware v1.63: A local low-privilege actor gains write access to something they should not be able to modify. What makes…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro X11SSL-CF hardware revision 1.01, BMC firmware v1.63","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A local low-privilege actor gains write access to something they should not be able to modify. What makes this worth tracking despite the modest score is where it sits: the VRM advisory track covers voltage regulator firmware, and write access to power-delivery firmware on a server is a physical-damage and availability primitive, not just an integrity one. On dense GPU nodes where the VRMs are already running near their limits, an attacker able to tamper with regulator configuration has a plausible path to hardware damage or node-level denial of service that no software remediation reverses. Insecure permissions, disclosed under Supermicro's VRM (voltage regulator module) advisory track rather than the BMC track.","attack_vector":"Local access to the node with low privilege - a user account on the host, not necessarily root. The permissions problem is on the node itself rather than across the management network.","remediation":"Firmware update per Supermicro's January 2023 VRM advisory. VRM firmware is updated separately from BIOS and BMC on Supermicro platforms, which is the operational trap here: an operator who believes they have a fully patched node because BIOS and BMC are current may still be running vulnerable regulator firmware. Add VRM firmware to whatever inventory you use to track BIOS and BMC versions, because most fleet tooling does not enumerate it by default.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-43309","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2022/43xxx/CVE-2022-43309.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-4543","cve":"CVE-2022-4543","aliases":[],"title":"Linux kernel (KASLR): EntryBleed: prefetch side channel defeats KASLR even with KPTI - enabling primitive for every other kernel…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (KASLR)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"EntryBleed: prefetch side channel defeats KASLR even with KPTI - enabling primitive for every other kernel exploit","attack_vector":"Any tenant process in a container","remediation":"Kernel patch, drain + reboot. Not independently exploitable, but it removes the main mitigation everything else relies on - treat as a severity multiplier on the whole kernel-privesc set","references":["https://access.redhat.com/security/cve/CVE-2022-4543"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-48766","cve":"CVE-2022-48766","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Wrap dcn301_calculate_wm_and_dlg for FPU.","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48766","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-48849","cve":"CVE-2022-48849","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A race condition or locking defect in the amdgpu firmware, ACPI and IP-block initialisation. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu firmware, ACPI and IP-block initialisation. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: bypass tiling flag check in virtual display case (v2)","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48849","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-48853","cve":"CVE-2022-48853","aliases":[],"title":"Linux swiotlb - info leak with DMA_FROM_DEVICE bounce buffers: MULTI-TENANT ISOLATION: The software IO TLB leaks information through bounce buffers on DMA_FROM_DEVICE…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux swiotlb - info leak with DMA_FROM_DEVICE bounce buffers","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The software IO TLB leaks information through bounce buffers on DMA_FROM_DEVICE transfers - stale buffer contents are exposed rather than being overwritten by the device. swiotlb is the bounce-buffer layer that SEV and SEV-SNP guests are forced to use for all DMA, because a confidential guest cannot let a device write directly into encrypted memory. So this leak sits precisely on the path every confidential VM's I/O takes, and what leaks is whatever the previous user of that bounce buffer left behind.","attack_vector":"Local, through DMA operations that use bounce buffers - which is all device I/O in an SEV/SNP guest, and any DMA above the device's addressing limit on a normal host.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot; no firmware step. Prioritise on confidential-computing hosts and inside confidential guest images, since SEV guests route all I/O through swiotlb by design.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48853"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49055","cve":"CVE-2022-49055","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Check for potential null return of kmalloc_array()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49055","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49069","cve":"CVE-2022-49069","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix by adding FPU protection for dcn30_internal_validate_bw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49069","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49102","cve":"CVE-2022-49102","aliases":[],"title":"habanalabs kernel driver (MMU shadow teardown): A memory leak on the habanalabs MMU teardown path. Each affected teardown leaks host-resident shadow…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"habanalabs kernel driver (MMU shadow teardown)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory leak on the habanalabs MMU teardown path. Each affected teardown leaks host-resident shadow page-table memory, so a tenant that repeatedly opens and closes the device can grind the node into OOM over a long-running shift. Slow-burn availability problem, not a confidentiality one.","attack_vector":"Local user with the habanalabs device node - repeated device open/close cycles are enough.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49102","https://git.kernel.org/stable/c/12e49aefda2e04b07604f13e03f40027cbeb0dc6"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49133","cve":"CVE-2022-49133","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: svm range restore work deadlock when process exit","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49133","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49135","cve":"CVE-2022-49135","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix memory leak","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49135","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49137","cve":"CVE-2022-49137","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amd/amdgpu/amdgpu_cs): A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amd/amdgpu/amdgpu_cs)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/amdgpu/amdgpu_cs: fix refcount leak of a dma_fence obj","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49137","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49232","cve":"CVE-2022-49232","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix a NULL pointer dereference in amdgpu_dm_connector_add_common_modes()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49232","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49233","cve":"CVE-2022-49233","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Call dc_stream_release for remove link enc assignment","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49233","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49294","cve":"CVE-2022-49294","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Check if modulo is 0 before dividing.","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49294","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49335","cve":"CVE-2022-49335","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/cs): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/cs)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu/cs: make commands with 0 chunks illegal behaviour.","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49335","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49365","cve":"CVE-2022-49365","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Off by one in dm_dmub_outbox1_low_irq()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49365","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49529","cve":"CVE-2022-49529","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: fix the null pointer while the smu is disabled","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49529","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49773","cve":"CVE-2022-49773","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Fix optc2_configure warning on dcn314","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49773","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49784","cve":"CVE-2022-49784","aliases":[],"title":"Linux perf/x86/amd/uncore - memory leak in the events array: Per-CPU northbridge and last-level-cache uncore contexts are allocated but not freed when a CPU comes online…","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Linux perf/x86/amd/uncore - memory leak in the events array","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Per-CPU northbridge and last-level-cache uncore contexts are allocated but not freed when a CPU comes online, leaking memory on every hotplug. Uncore counters are what you use to measure memory bandwidth and cache behaviour on AMD - i.e. the telemetry an AI operator actually cares about - so this leaks in proportion to how much you monitor.","attack_vector":"Local, driven by CPU hotplug with uncore perf events in use.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49784"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49864","cve":"CVE-2022-49864","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix NULL pointer dereference in svm_migrate_to_ram()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49864","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49965","cve":"CVE-2022-49965","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: add missing ->fini_xxxx interfaces for some SMU13 asics","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49965","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49966","cve":"CVE-2022-49966","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: add missing ->fini_microcode interface for Sienna Cichlid","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49966","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-49971","cve":"CVE-2022-49971","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/pm: Fix a potential gpu_metrics_table memory leak","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49971","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50479","cve":"CVE-2022-50479","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd: fix potential memory leak","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50479","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50515","cve":"CVE-2022-50515","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix memory leak in hpd_rx_irq_create_workqueue()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50515","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50527","cve":"CVE-2022-50527","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu kernel driver core. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Fix size validation for non-exclusive domains (v4)","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50527","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50535","cve":"CVE-2022-50535","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential null-deref in dm_resume","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50535","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0188","cve":"CVE-2023-0188","aliases":[],"title":"GPU Display Driver: DoS (stack overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (stack overflow)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0188","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-119"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0190","cve":"CVE-2023-0190","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0190","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0197","cve":"CVE-2023-0197","aliases":[],"title":"vGPU Manager (Cloud Gaming): Guest-triggered host DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager (Cloud Gaming)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Guest-triggered host DoS (null deref)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager on hypervisor; migrate/evict guest VMs, reboot host","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0197","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-1018","cve":"CVE-2023-1018","aliases":[],"title":"TPM 2.0 reference implementation: Out-of-bounds read in the same routine — disclosure of TPM-resident data","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"TPM 2.0 reference implementation","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the same routine — disclosure of TPM-resident data","attack_vector":"Local, low privilege","remediation":"Same TPM firmware update and the same key-loss problem","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-1018"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-31021","cve":"CVE-2023-31021","aliases":[],"title":"vGPU software: Guest-triggered host DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU software","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Guest-triggered host DoS (null deref)","attack_vector":"Tenant VM guest","remediation":"Upgrade vGPU Manager; migrate VMs, reboot host","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-31022","cve":"CVE-2023-31022","aliases":[],"title":"GPU Display Driver / vGPU: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver / vGPU","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Tenant container or VM guest","remediation":"Driver + vGPU Manager upgrade; rolling reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-31023","cve":"CVE-2023-31023","aliases":[],"title":"GPU Display Driver (Windows): DoS (untrusted pointer deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver (Windows)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (untrusted pointer deref)","attack_vector":"Local user","remediation":"Upgrade Oct-2023 driver branch","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5491/5491.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-822"]},{"id":"CVE-2023-34327","cve":"CVE-2023-34327","aliases":[],"title":"Xen on AMD - debug extensions (DBEXT) exposure to guests: MULTI-TENANT ISOLATION: AMD CPUs since roughly 2014 carry extensions to x86 debugging that Xen exposed to…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD - debug extensions (DBEXT) exposure to guests","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: AMD CPUs since roughly 2014 carry extensions to x86 debugging that Xen exposed to guests without properly managing the associated MSR state across context switches. A guest can use debug facilities to observe or interfere with state belonging to another context - debug hardware is designed to see everything, which is precisely why leaking it across a VM boundary matters.","attack_vector":"From inside a guest VM on AMD hardware under Xen.","remediation":"Fixed in Xen (XSA-444). Update the hypervisor and reboot the host.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34327","https://xenbits.xen.org/xsa/advisory-444.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-34328","cve":"CVE-2023-34328","aliases":[],"title":"Xen on AMD - debug extensions (DBEXT) exposure to guests: MULTI-TENANT ISOLATION: Companion to the other XSA-444 debug-extension issue on AMD. Guest-accessible debug…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD - debug extensions (DBEXT) exposure to guests","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Companion to the other XSA-444 debug-extension issue on AMD. Guest-accessible debug state that is not correctly isolated across context switches.","attack_vector":"From inside a guest VM on AMD hardware under Xen.","remediation":"Fixed in Xen (XSA-444). Hypervisor update plus host reboot; patch both XSA-444 CVEs together.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-34328","https://xenbits.xen.org/xsa/advisory-444.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-38575","cve":"CVE-2023-38575","aliases":[],"title":"Intel processors (return predictor target sharing): Return predictor targets are shared non-transparently between contexts, giving an authorised local user an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (return predictor target sharing)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Return predictor targets are shared non-transparently between contexts, giving an authorised local user an information-disclosure channel. Fixed in the same microcode wave as the 2024 return-predictor advisories.","attack_vector":"Local authorised code.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-38575","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00982.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2023-40238","cve":"CVE-2023-40238","aliases":["LogoFAIL"],"title":"Insyde InsydeH2O BmpDecoderDxe: Crafted BMP logo copies data to a chosen address during DXE — arbitrary write before Secure Boot","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O BmpDecoderDxe","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Crafted BMP logo copies data to a chosen address during DXE — arbitrary write before Secure Boot","attack_vector":"Local, ESP write","remediation":"Insyde kernel update shipped through each OEM; the CVSS understates it because the outcome is a firmware implant","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40238"],"status":"curated"},{"id":"CVE-2023-40549","cve":"CVE-2023-40549","aliases":["shim 15.8 batch"],"title":"shim (verify_buffer_authenticode): Out-of-bounds read on a malformed PE file crashes shim and blocks boot. Same operational cost as the other…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (verify_buffer_authenticode)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read on a malformed PE file crashes shim and blocks boot. Same operational cost as the other shim DoS - a node that will not boot needs hands or BMC per box.","attack_vector":"A malformed EFI binary in the boot path.","remediation":"shim package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40549","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-40550","cve":"CVE-2023-40550","aliases":["shim 15.8 batch"],"title":"shim (verify_buffer_sbat): Out-of-bounds read in SBAT verification discloses adjacent boot-time memory to an attacker who can already…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (verify_buffer_sbat)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read in SBAT verification discloses adjacent boot-time memory to an attacker who can already run in the boot path. Reconnaissance value rather than direct compromise.","attack_vector":"Crafted SBAT metadata in a binary shim is asked to verify.","remediation":"shim package update + reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40550","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-4211","cve":"CVE-2023-4211","aliases":[],"title":"Arm Mali GPU kernel driver: Use-after-free via improper GPU memory processing - local non-privileged user reaches freed memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Arm Mali GPU kernel driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":true,"impact":"Use-after-free via improper GPU memory processing - local non-privileged user reaches freed memory; exploited in the wild [KEV]","attack_vector":"Local user with GPU device access","remediation":"Driver update + reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4211"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-4327","cve":"CVE-2023-4327","aliases":["CVE-2023-4328"],"title":"Broadcom LSI Storage Authority (LSA) - on-disk credential/key storage on Linux and Windows: The keys LSA uses to encrypt its stored secrets sit in files any local user can read. On a bare-metal node…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Broadcom LSI Storage Authority (LSA) - on-disk credential/key storage on Linux and Windows","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The keys LSA uses to encrypt its stored secrets sit in files any local user can read. On a bare-metal node that means anyone who lands non-root code on the host - a container escape, a compromised monitoring agent, a tenant with shell on a shared management jumpbox - recovers the credentials LSA uses to talk to the RAID controller and, in Gateway deployments, to other nodes. That turns a low-privilege foothold into control of the storage controller on the fleet, which is the layer that decides whether the next tenant sees the previous tenant's blocks.","attack_vector":"Any unprivileged local account on a host running the LSA agent (Linux or Windows). No network exposure required, no controller access required first.","remediation":"Upgrade LSA to the fixed 7.017.011.000 build. If your OEM has not shipped it, tighten the file permissions on the LSA install directory and its key material by hand and re-apply after every OEM tooling update, since OEM installers reset them. Rotate any controller/gateway credentials that were stored under the old keys - patching alone does not invalidate what has already leaked. No reboot, no array impact.","references":["https://www.broadcom.com/support/resources/product-security-center","https://nvd.nist.gov/vuln/detail/CVE-2023-4327"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2023-52460","cve":"CVE-2023-52460","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL pointer dereference at hibernate","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52460","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52485","cve":"CVE-2023-52485","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Wake DMCUB before sending a command","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52485","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52585","cve":"CVE-2023-52585","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix possible NULL dereference in amdgpu_ras_query_error_status_helper()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52585","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52625","cve":"CVE-2023-52625","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Refactor DMCUB enter/exit idle interface","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52625","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52632","cve":"CVE-2023-52632","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Fix lock dependency warning with srcu","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52632","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52634","cve":"CVE-2023-52634","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix disable_otg_wa logic","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52634","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52659","cve":"CVE-2023-52659","aliases":[],"title":"Linux x86/mm - pfn_to_kaddr() 64-bit input handling (SNP support code): On 64-bit platforms the pfn_to_kaddr() macro dropped high address bits when handed a narrower type, producing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/mm - pfn_to_kaddr() 64-bit input handling (SNP support code)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On 64-bit platforms the pfn_to_kaddr() macro dropped high address bits when handed a narrower type, producing wrong kernel addresses in the SEV-SNP support paths that use it. Wrong addresses in code that manages confidential-guest page state means operating on memory that is not the memory intended - a correctness failure right underneath the mechanism enforcing guest isolation.","attack_vector":"Local, in the host kernel's SNP page-management paths.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52659"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52671","cve":"CVE-2023-52671","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix hang/underflow when transitioning to ODM4:1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52671","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52673","cve":"CVE-2023-52673","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix a debugfs null pointer error","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52673","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52695","cve":"CVE-2023-52695","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check writeback connectors in create_validate_stream_for_sink","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52695","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52753","cve":"CVE-2023-52753","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid NULL dereference of timing generator","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52753","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52773","cve":"CVE-2023-52773","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix a NULL pointer dereference in amdgpu_dm_i2c_xfer()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52773","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52814","cve":"CVE-2023-52814","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix potential null pointer derefernce","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52814","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52815","cve":"CVE-2023-52815","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vkms): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu/vkms)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/vkms: fix a possible null pointer dereference","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52815","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52817","cve":"CVE-2023-52817","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix a null pointer access when the smc_rreg pointer is NULL","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52908","cve":"CVE-2023-52908","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix potential NULL dereference","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52908","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-52912","cve":"CVE-2023-52912","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Fixed bug on error when unloading amdgpu","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52912","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53036","cve":"CVE-2023-53036","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Fix call trace warning and hang when removing amdgpu device","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53036","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53042","cve":"CVE-2023-53042","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Do not set DRR on pipe Commit","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53042","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53073","cve":"CVE-2023-53073","aliases":[],"title":"Linux perf/x86/amd/core - overflow status not cleared for unhandled indices: Unhandled overflow bits are left set in the PMU status register, so stale overflow state persists and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux perf/x86/amd/core - overflow status not cleared for unhandled indices","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unhandled overflow bits are left set in the PMU status register, so stale overflow state persists and subsequent interrupts are misattributed. The visible effect is corrupted performance data and spurious NMIs - which on a GPU cluster means your capacity and efficiency measurements are quietly wrong.","attack_vector":"Local, through perf counter overflow handling.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53073"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53074","cve":"CVE-2023-53074","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A race condition or locking defect in the amdgpu RAS / GPU reset and recovery path. Concurrent paths touch…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu RAS / GPU reset and recovery path. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix ttm_bo calltrace warning in psp_hw_fini","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53074","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53152","cve":"CVE-2023-53152","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A race condition or locking defect in the amdgpu power management (SMU/powerplay). Concurrent paths touch…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu power management (SMU/powerplay). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix calltrace warning in amddrm_buddy_fini","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53152","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53193","cve":"CVE-2023-53193","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix amdgpu_irq_put call trace in gmc_v10_0_hw_fini","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53193","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53228","cve":"CVE-2023-53228","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: drop redundant sched job cleanup when cs is aborted","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53228","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53237","cve":"CVE-2023-53237","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix amdgpu_irq_put call trace in gmc_v11_0_hw_fini","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53237","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53248","cve":"CVE-2023-53248","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: install stub fence into potential unused fence pointers","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53248","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53258","cve":"CVE-2023-53258","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Fix possible underflow for displays with large vblank","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53258","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53351","cve":"CVE-2023-53351","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the DRM…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/sched: Check scheduler work queue before calling timeout handling","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53351","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53352","cve":"CVE-2023-53352","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/ttm): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the DRM scheduler /…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/ttm)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/ttm: check null pointer before accessing when swapping","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53352","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53370","cve":"CVE-2023-53370","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix memory leak in mes self test","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53370","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53438","cve":"CVE-2023-53438","aliases":[],"title":"Linux x86/MCE - CS register not saved on AMD Zen Instruction Fetch Poison errors: On AMD Zen systems, the Instruction Fetch unit does not report the address of a poisoned…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/MCE - CS register not saved on AMD Zen Instruction Fetch Poison errors","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On AMD Zen systems, the Instruction Fetch unit does not report the address of a poisoned (uncorrectable-error) instruction fetch, and the kernel was not saving the CS register either - so when memory corruption hits an instruction fetch, you lose the information needed to attribute it. On a GPU training fleet where uncorrectable memory errors are a routine operational event, losing attribution means you cannot tell which tenant's job hit the bad memory or which DIMM to replace.","attack_vector":"Not attacker-driven. This is an observability failure on the machine-check path.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot. Worth taking on any fleet where you do RAS-driven node retirement - without it your machine-check records are missing the field you need to act on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53438"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53498","cve":"CVE-2023-53498","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential null dereference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53498","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53547","cve":"CVE-2023-53547","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A race condition or locking defect in the amdgpu firmware, ACPI and IP-block initialisation. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu firmware, ACPI and IP-block initialisation. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Fix sdma v4 sw fini error","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53547","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53563","cve":"CVE-2023-53563","aliases":[],"title":"Linux cpufreq/amd-pstate-ut - kernel panic when loading the unit-test driver: Loading the amd-pstate unit-test module panics the kernel. A test module shipped in production kernels that…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux cpufreq/amd-pstate-ut - kernel panic when loading the unit-test driver","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Loading the amd-pstate unit-test module panics the kernel. A test module shipped in production kernels that crashes the host when loaded - relevant because automated hardware-validation tooling on a new GPU fleet is exactly the thing that loads every available module to see what happens.","attack_vector":"Local, requires the ability to load the amd_pstate_ut module - so root, or automated burn-in tooling running as root.","remediation":"Distro kernel update plus reboot. Meanwhile, keep amd_pstate_ut out of any module-loading sweep in your node acceptance testing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53563"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53625","cve":"CVE-2023-53625","aliases":[],"title":"Linux i915 GVT-g mediated GPU virtualisation: MULTI-TENANT ISOLATION: Unsafe cleanup of per-vGPU debugfs state when a mediated vGPU is destroyed. GVT-g is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux i915 GVT-g mediated GPU virtualisation","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Unsafe cleanup of per-vGPU debugfs state when a mediated vGPU is destroyed. GVT-g is the mechanism that carves one physical Intel GPU into vGPUs handed to different VMs, so bugs in its lifecycle paths sit directly on the tenant boundary. Practical effect here is a host kernel crash triggered by a vGPU teardown.","attack_vector":"Reachable by whoever can cause a vGPU to be created and destroyed - the virtualisation control plane, or a tenant that can start and stop VMs.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53625","https://git.kernel.org/stable/c/44c0e07e3972e3f2609d69ad873d4f342f8a68ec"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53628","cve":"CVE-2023-53628","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2023","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: drop gfx_v11_0_cp_ecc_error_irq_funcs","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53628","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-6548","cve":"CVE-2023-6548","aliases":[],"title":"Citrix NetScaler ADC/Gateway: Code injection on the management interface","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Citrix NetScaler ADC/Gateway","year":"2023","cvss_score":5.5,"severity":"medium","kev":true,"impact":"[KEV] Code injection on the management interface -> authenticated RCE via NSIP/CLIP/SNIP","attack_vector":"Adjacent network","remediation":"Control-plane: patch; management IPs must never be internet-reachable","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6548"],"status":"curated"},{"id":"CVE-2024-0086","cve":"CVE-2024-0086","aliases":[],"title":"vGPU Manager: Host DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS (null deref)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0086","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"]},{"id":"CVE-2024-0088","cve":"CVE-2024-0088","aliases":[],"title":"Triton Inference Server: DoS (stack buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (stack buffer overflow)","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0088","https://github.com/NVIDIA/product-security/tree/main/2024/5535"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:L/I:L/A:H","cwe":["CWE-119"]},{"id":"CVE-2024-0092","cve":"CVE-2024-0092","aliases":[],"title":"GPU Display Driver: DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0092","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","fleet":{"ubiquity":"Universal - same driver","remediation_pain":"`node-reboot`","pain_class":"node-reboot","why_fleet_wide":"Improper exception handling gives a local tenant a reliable node-level DoS: one customer can knock an 8-GPU box offline for everyone on it"},"cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-703"]},{"id":"CVE-2024-0094","cve":"CVE-2024-0094","aliases":[],"title":"vGPU Manager: Host resource exhaustion / DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host resource exhaustion / DoS","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0094","https://github.com/NVIDIA/product-security/tree/main/2024/5551"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-799"]},{"id":"CVE-2024-0098","cve":"CVE-2024-0098","aliases":[],"title":"ChatRTX: Unencrypted credential storage","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ChatRTX","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unencrypted credential storage","attack_vector":"Local user","remediation":"Consumer app; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0098","https://github.com/NVIDIA/product-security/tree/main/2024/5533"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-319"]},{"id":"CVE-2024-0137","cve":"CVE-2024-0137","aliases":[],"title":"Container Toolkit / GPU Operator: Host DoS / info disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS / info disclosure","attack_vector":"Any tenant with a container","remediation":"Bump toolkit + restart runtime","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0137","https://github.com/NVIDIA/product-security/tree/main/2025/5599"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:R/S:C/C:L/I:L/A:L","cwe":["CWE-653"]},{"id":"CVE-2024-0147","cve":"CVE-2024-0147","aliases":[],"title":"GPU Display Driver: DoS (use-after-free)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (use-after-free)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0147","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-1441","cve":"CVE-2024-1441","aliases":[],"title":"libvirt: Off-by-one in udevListInterfacesByStatus() - libvirtd crash / info leak","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"libvirt","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Off-by-one in udevListInterfacesByStatus() - libvirtd crash / info leak","attack_vector":"Local user or API client with libvirt access","remediation":"Package update + libvirtd restart; running domains survive","references":["https://access.redhat.com/security/cve/CVE-2024-1441"],"status":"curated"},{"id":"CVE-2024-26629","cve":"CVE-2024-26629","aliases":[],"title":"Linux nfsd (NFS server): Broken RELEASE_LOCKOWNER handling in nfsd causing state corruption","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux nfsd (NFS server)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Broken RELEASE_LOCKOWNER handling in nfsd causing state corruption","attack_vector":"Local","remediation":"Data-plane: kernel update batched into the next reboot window","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26629"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26647","cve":"CVE-2024-26647","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix late derefrence 'dsc' check in 'link_set_dsc_pps_packet()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26647","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26648","cve":"CVE-2024-26648","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix variable deferencing before NULL check in edp_setup_replay()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26648","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26649","cve":"CVE-2024-26649","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix the null pointer when load rlc firmware","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26649","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26657","cve":"CVE-2024-26657","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched): A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/sched)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/sched: fix null-ptr-deref in init entity","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26657","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26661","cve":"CVE-2024-26661","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add NULL test for 'timing generator' in 'dcn21_set_pipe()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26661","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26662","cve":"CVE-2024-26662","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix 'panel_cntl' could be null in 'dcn21_set_backlight_level()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26662","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26700","cve":"CVE-2024-26700","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix MST Null Ptr for RV","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26700","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26729","cve":"CVE-2024-26729","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential null pointer dereference in dc_dmub_srv","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26729","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26767","cve":"CVE-2024-26767","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fixed integer types and null check locations","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26767","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26817","cve":"CVE-2024-26817","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: amdkfd: use calloc instead of kzalloc to avoid integer overflow","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26817","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26833","cve":"CVE-2024-26833","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix memory leak in dm_sw_fini()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26833","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26915","cve":"CVE-2024-26915","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu RAS / GPU reset and recovery path - a length…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdgpu RAS / GPU reset and recovery path - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Reset IH OVERFLOW_CLEAR bit","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26915","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26948","cve":"CVE-2024-26948","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add a dc_state NULL check in dc_state_release","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26948","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26949","cve":"CVE-2024-26949","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: Fix NULL pointer dereference when get power limit","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26949","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-26986","cve":"CVE-2024-26986","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: Fix memory leak in create_process failure","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-26986","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27041","cve":"CVE-2024-27041","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix NULL checks for adev->dm.dc in amdgpu_dm_fini()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27041","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27044","cve":"CVE-2024-27044","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix potential NULL pointer dereferences in 'dcn10_set_output_transfer_func()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27044","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-31584","cve":"CVE-2024-31584","aliases":[],"title":"PyTorch (flatbuffer loader): Out-of-bounds read parsing flatbuffer model","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"PyTorch (flatbuffer loader)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Out-of-bounds read parsing flatbuffer model","attack_vector":"Customer-supplied model file","remediation":"Ship torch >= 2.2.0","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31584"],"status":"curated"},{"id":"CVE-2024-35795","cve":"CVE-2024-35795","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix deadlock while reading mqd from debugfs","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35795","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-35799","cve":"CVE-2024-35799","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Prevent crash when disable stream","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35799","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36026","cve":"CVE-2024-36026","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/pm: fixes a random hang in S4 for SMU v13.0.4/11","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36026","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36284","cve":"CVE-2024-36284","aliases":[],"title":"Intel Neural Compressor: Input-validation failure reachable by an authenticated user, ending in privilege escalation inside the Neural…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Neural Compressor","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Input-validation failure reachable by an authenticated user, ending in privilege escalation inside the Neural Compressor service.","attack_vector":"Authenticated user of the service.","remediation":"Upgrade to v3.0 or later.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36284","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01219.html"],"status":"curated"},{"id":"CVE-2024-36316","cve":"CVE-2024-36316","aliases":[],"title":"AMD Graphics Driver - integer overflow bypassing size checks: An integer overflow in the AMD graphics driver lets an attacker wrap a size calculation and slip past the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Graphics Driver - integer overflow bypassing size checks","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An integer overflow in the AMD graphics driver lets an attacker wrap a size calculation and slip past the driver's own bounds checks. The advisory scopes the outcome to denial of service, but overflow-defeats-size-check is the standard front half of a heap corruption chain, so treat the ceiling as higher than the score.","attack_vector":"Local, via the graphics driver interface.","remediation":"Update the AMD graphics driver and reload or reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36316","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36897","cve":"CVE-2024-36897","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Atom Integrated System Info v2_2 for DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36897","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36951","cve":"CVE-2024-36951","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: range check cp bad op exception interrupts","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36951","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-36969","cve":"CVE-2024-36969","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Fix division by zero in setup_dsc_config","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36969","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-40987","cve":"CVE-2024-40987","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: fix UBSAN warning in kv_dpm.c","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40987","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-40997","cve":"CVE-2024-40997","aliases":[],"title":"Linux cpufreq/amd-pstate - memory leak on CPU EPP exit: The amd-pstate driver leaks its per-CPU allocation when a CPU's energy-performance-preference path exits.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux cpufreq/amd-pstate - memory leak on CPU EPP exit","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The amd-pstate driver leaks its per-CPU allocation when a CPU's energy-performance-preference path exits. Each CPU hotplug or governor transition drops memory on the floor - slow, but on a long-lived host that cycles power states under variable AI load it accumulates into unreclaimable kernel memory and eventually pressures every workload on the node.","attack_vector":"Local, driven by CPU hotplug and power-governor transitions rather than by an attacker directly - though a tenant that can influence CPU frequency governors can accelerate it.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot; no firmware step. Group it with the other amd-pstate fixes rather than scheduling a window for it alone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40997"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-41093","cve":"CVE-2024-41093","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: avoid using null object of framebuffer","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-41093","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42122","cve":"CVE-2024-42122","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL pointer check for kzalloc","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42122","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43827","cve":"CVE-2024-43827","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check before access structs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43827","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43886","cve":"CVE-2024-43886","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check in resource_log_pipe_topology_update","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43886","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43899","cve":"CVE-2024-43899","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix null pointer deref in dcn20_resource.c","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43899","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43901","cve":"CVE-2024-43901","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL pointer dereference for DTN log in DCN401","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43901","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43902","cve":"CVE-2024-43902","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null checker before passing variables","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43902","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43904","cve":"CVE-2024-43904","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null checks for 'stream' and 'plane' before dereferencing","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43904","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43905","cve":"CVE-2024-43905","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: Fix the null pointer dereference for vega10_hwmgr","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43905","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43907","cve":"CVE-2024-43907","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: Fix the null pointer dereference in apply_state_adjust_rules","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43907","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43908","cve":"CVE-2024-43908","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix the null pointer dereference to ras_manager","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43908","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-43909","cve":"CVE-2024-43909","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/pm: Fix the null pointer dereference for smu7","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-43909","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-44961","cve":"CVE-2024-44961","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Forward soft recovery errors to userspace","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44961","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-44980","cve":"CVE-2024-44980","aliases":[],"title":"Linux drm/xe GPU kernel driver (display opregion): A resource leak in the xe driver's display opregion handling. xe is the newer Intel GPU driver that Data…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux drm/xe GPU kernel driver (display opregion)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A resource leak in the xe driver's display opregion handling. xe is the newer Intel GPU driver that Data Center GPU Max and later parts move to, so this is worth tracking as the xe driver takes over from i915 in production images.","attack_vector":"Local, on driver load/unload cycles.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-44980","https://git.kernel.org/stable/c/f4b2a0ae1a31fd3d1b5ca18ee08319b479cf9b5f"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46694","cve":"CVE-2024-46694","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: avoid using null object of framebuffer","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46694","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46714","cve":"CVE-2024-46714","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Skip wbscl_set_scaler_filter if filter is null","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46714","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46720","cve":"CVE-2024-46720","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix dereference after null check","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46720","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46726","cve":"CVE-2024-46726","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Ensure index calculation will not overflow","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46726","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46727","cve":"CVE-2024-46727","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add otg_master NULL check within resource_log_pipe_topology_update","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46727","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46728","cve":"CVE-2024-46728","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Check index for aux_rd_interval before using","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46728","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46732","cve":"CVE-2024-46732","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Assign linear_pitch_alignment even for VM","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46732","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46772","cve":"CVE-2024-46772","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check denominator crb_pipes before used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46772","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46773","cve":"CVE-2024-46773","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check denominator pbn_div before used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46773","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46775","cve":"CVE-2024-46775","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Validate function returns","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46775","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46776","cve":"CVE-2024-46776","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Run DC_LOG_DC after checking link->link_enc","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46776","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46778","cve":"CVE-2024-46778","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Check UnboundedRequestEnabled's value","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46778","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46802","cve":"CVE-2024-46802","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: added NULL check at start of dc_validate_stream","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46802","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46805","cve":"CVE-2024-46805","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix the waring dereferencing hive","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46805","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46806","cve":"CVE-2024-46806","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Fix the warning division or modulo by zero","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46806","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46807","cve":"CVE-2024-46807","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu kernel…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu kernel driver core. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/amdgpu: Check tbo resource pointer","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46807","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46808","cve":"CVE-2024-46808","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add missing NULL pointer check within dpcd_extend_address_range","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46808","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46809","cve":"CVE-2024-46809","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check BIOS images before it is used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46809","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46819","cve":"CVE-2024-46819","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: the warning dereferencing obj for nbio_v7_4","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46819","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46835","cve":"CVE-2024-46835","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Fix smatch static checker warning","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46835","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46896","cve":"CVE-2024-46896","aliases":[],"title":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu): A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel DRM scheduler / TTM / dma-buf shared layer used by amdgpu (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the DRM scheduler / TTM / dma-buf shared layer used by amdgpu. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: don't access invalid sched","attack_vector":"Local. Reachable by any process that can submit GPU work or import/export a dma-buf - i.e. any ROCm or graphics tenant on the node. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46896","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47661","cve":"CVE-2024-47661","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Avoid overflow from uint32_t to uint8_t","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47661","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47662","cve":"CVE-2024-47662","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Remove register from DCN35 DMCUB diagnostic collection","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47662","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47683","cve":"CVE-2024-47683","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Skip Recompute DSC Params if no Stream on Link","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47683","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47704","cve":"CVE-2024-47704","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check link_res->hpo_dp_link_enc before using it","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47704","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47720","cve":"CVE-2024-47720","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for set_output_gamma in dcn30_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47720","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49890","cve":"CVE-2024-49890","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: ensure the fw_info is not null before using it","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49890","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49892","cve":"CVE-2024-49892","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Initialize get_bytes_per_element's default to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49892","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49893","cve":"CVE-2024-49893","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check stream_status before it is used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49893","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49896","cve":"CVE-2024-49896","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check stream before comparing them","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49896","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49897","cve":"CVE-2024-49897","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check phantom_stream before it is used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49897","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49898","cve":"CVE-2024-49898","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null-initialized variables","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49898","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49899","cve":"CVE-2024-49899","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Initialize denominators' default to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49899","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49904","cve":"CVE-2024-49904","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: add list empty check to avoid null pointer issue","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49904","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49905","cve":"CVE-2024-49905","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for 'afb' in amdgpu_dm_plane_handle_cursor_update (v2)","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49905","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49906","cve":"CVE-2024-49906","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check null pointer before try to access it","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49906","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49907","cve":"CVE-2024-49907","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointers before using dc->clk_mgr","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49907","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49908","cve":"CVE-2024-49908","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for 'afb' in amdgpu_dm_update_cursor (v2)","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49908","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49909","cve":"CVE-2024-49909","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for function pointer in dcn32_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49909","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49910","cve":"CVE-2024-49910","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for function pointer in dcn401_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49910","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49911","cve":"CVE-2024-49911","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for function pointer in dcn20_set_output_transfer_func","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49911","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49912","cve":"CVE-2024-49912","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Handle null 'stream_status' in 'planes_changed_for_existing_stream'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49912","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49913","cve":"CVE-2024-49913","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for top_pipe_to_program in commit_planes_for_stream","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49913","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49914","cve":"CVE-2024-49914","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add null check for pipe_ctx->plane_state in dcn20_program_pipe","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49914","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49915","cve":"CVE-2024-49915","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL check for clk_mgr in dcn32_init_hw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49915","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49916","cve":"CVE-2024-49916","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add NULL check for clk_mgr and clk_mgr->funcs in dcn401_init_hw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49916","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49917","cve":"CVE-2024-49917","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Add NULL check for clk_mgr and clk_mgr->funcs in dcn30_init_hw","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49917","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49918","cve":"CVE-2024-49918","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for head_pipe in dcn32_acquire_idle_pipe_for_head_pipe_in_layer","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49918","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49919","cve":"CVE-2024-49919","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null check for head_pipe in dcn201_acquire_free_pipe_for_layer","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49919","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49920","cve":"CVE-2024-49920","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointers before multiple uses","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49920","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49921","cve":"CVE-2024-49921","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Check null pointers before used","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49921","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49922","cve":"CVE-2024-49922","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointers before using them","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49922","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49923","cve":"CVE-2024-49923","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Pass non-null to dcn20_validate_apply_pipe_split_flags","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49923","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49970","cve":"CVE-2024-49970","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: Implement bounds check for stream encoder creation in DCN401","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49970","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49971","cve":"CVE-2024-49971","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Increase array size of dummy_boolean","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49971","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-49972","cve":"CVE-2024-49972","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Memory is handed to a consumer without being initialised or cleared in the amdgpu display core (DC/DM).…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory is handed to a consumer without being initialised or cleared in the amdgpu display core (DC/DM). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amd/display: Deallocate DML memory if allocation fails","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-49972","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50003","cve":"CVE-2024-50003","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix system hang while resume with TBT monitor","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50003","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50004","cve":"CVE-2024-50004","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: update DML2 policy EnhancedPrefetchScheduleAccelerationFinal DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50004","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50009","cve":"CVE-2024-50009","aliases":[],"title":"Linux cpufreq/amd-pstate - unchecked cpufreq_cpu_get() return value: cpufreq_cpu_get() can return NULL and amd-pstate did not check it, giving a kernel NULL dereference and a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux cpufreq/amd-pstate - unchecked cpufreq_cpu_get() return value","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"cpufreq_cpu_get() can return NULL and amd-pstate did not check it, giving a kernel NULL dereference and a host panic. Availability failure in always-running platform code.","attack_vector":"Local, on the amd-pstate path.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50009"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50049","cve":"CVE-2024-50049","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Check null pointer before dereferencing se","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50049","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50108","cve":"CVE-2024-50108","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Disable PSR-SU on Parade 08-01 TCON too","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50108","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50172","cve":"CVE-2024-50172","aliases":[],"title":"Linux bnxt_re RoCE driver (chip context memory leak): Memory leak in the Broadcom RoCE driver when doorbell BAR mapping fails. Low severity on its own, but…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_re RoCE driver (chip context memory leak)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Memory leak in the Broadcom RoCE driver when doorbell BAR mapping fails. Low severity on its own, but doorbell-BAR mapping is the mechanism that gives a userspace RDMA process direct hardware access, and leaks in its error path are worth tracking on a fabric where that mapping is the tenant boundary.","attack_vector":"Local, via repeated RDMA device setup failures.","remediation":"Kernel/driver upgrade plus host reboot; bundle with the other bnxt_re fixes rather than scheduling separately.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50172"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-50177","cve":"CVE-2024-50177","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix a UBSAN warning in DML2.1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-50177","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53060","cve":"CVE-2024-53060","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: prevent NULL pointer dereference if ATIF is not supported","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53060","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53072","cve":"CVE-2024-53072","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (platform/x86/amd/pmc): A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (platform/x86/amd/pmc)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: platform/x86/amd/pmc: Detect when STB is not available","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53072","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53200","cve":"CVE-2024-53200","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix null check for pipe_ctx->plane_state in hwss_setup_dpp","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53200","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53201","cve":"CVE-2024-53201","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix null check for pipe_ctx->plane_state in dcn20_program_pipe","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53201","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53869","cve":"CVE-2024-53869","aliases":[],"title":"GPU Display Driver: Info disclosure (missing initialization)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Info disclosure (missing initialization)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53869","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-459"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53881","cve":"CVE-2024-53881","aliases":[],"title":"vGPU Manager: Host DoS (missing error handling)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS (missing error handling)","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade; evacuate VMs","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53881","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-459"]},{"id":"CVE-2024-55881","cve":"CVE-2024-55881","aliases":[],"title":"Linux KVM x86 - hypercall completion for protected guests: KVM used the wrong helper to decide whether a hypercall was 64-bit when completing it, which misbehaves for…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM x86 - hypercall completion for protected guests","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM used the wrong helper to decide whether a hypercall was 64-bit when completing it, which misbehaves for guests with protected state such as SEV-ES and SEV-SNP where the host cannot see guest registers. The result is host-side state confusion driven by a confidential guest - a crash or incorrect emulation on the hypervisor path that every other guest on the node shares.","attack_vector":"From inside a protected (SEV-ES/SNP) guest issuing hypercalls.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-55881"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56542","cve":"CVE-2024-56542","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: fix a memleak issue when driver is removed","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56542","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56594","cve":"CVE-2024-56594","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: set the right AMDGPU sg segment limitation","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56594","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56666","cve":"CVE-2024-56666","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Dereference null return value","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56666","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56697","cve":"CVE-2024-56697","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Fix the memory allocation issue in amdgpu_discovery_get_nps_info()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56697","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-56753","cve":"CVE-2024-56753","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx9): A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx9)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/gfx9: Add Cleaner Shader Deinitialization in gfx_v9_0 Module","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-56753","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57897","cve":"CVE-2024-57897","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Correct the migration DMA map direction","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57897","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57919","cve":"CVE-2024-57919","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: fix divide error in DM plane scale calcs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57919","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57922","cve":"CVE-2024-57922","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Add check for granularity in dml ceil/floor helpers","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57922","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-57950","cve":"CVE-2024-57950","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Initialize denominator defaults to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-57950","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-58052","cve":"CVE-2024-58052","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2024","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Fix potential NULL pointer dereference in atomctrl_get_smc_sclk_range_table","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-58052","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21645","cve":"CVE-2025-21645","aliases":[],"title":"Linux platform/x86/amd/pmc - IRQ1 wakeup disabled unconditionally: The AMD PMC driver disabled IRQ1 wakeup in cases where i8042 had never enabled it, corrupting interrupt…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux platform/x86/amd/pmc - IRQ1 wakeup disabled unconditionally","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The AMD PMC driver disabled IRQ1 wakeup in cases where i8042 had never enabled it, corrupting interrupt wakeup configuration. Low-severity platform-driver correctness bug; included because the amd/pmc driver autoloads on AMD hosts whether or not the platform needs it, and unnecessary loaded drivers are unnecessary attack surface.","attack_vector":"Local, through platform power-management paths.","remediation":"Distro kernel update plus reboot, or blacklist the module on server images where power management is handled by the BMC and BIOS anyway.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21645"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21784","cve":"CVE-2025-21784","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: bail out when failed to load fw in psp_init_cap_microcode()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21784","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21841","cve":"CVE-2025-21841","aliases":[],"title":"Linux cpufreq/amd-pstate - cpufreq_policy reference counting: amd_pstate_update_limits() takes a cpufreq_policy reference and never drops it. A leaked reference means the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux cpufreq/amd-pstate - cpufreq_policy reference counting","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"amd_pstate_update_limits() takes a cpufreq_policy reference and never drops it. A leaked reference means the policy object can never be freed, so CPU hotplug and driver unbind hang or leak indefinitely - the sort of defect that makes a node impossible to cleanly reconfigure without a reboot.","attack_vector":"Local, through the amd-pstate limits-update path.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21841"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21940","cve":"CVE-2025-21940","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix NULL Pointer Dereference in KFD queue","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21940","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21941","cve":"CVE-2025-21941","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix null check for pipe_ctx->plane_state in resource_build_scaling_params","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21941","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21956","cve":"CVE-2025-21956","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Assign normalized_pix_clk when color depth = 14","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21956","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21987","cve":"CVE-2025-21987","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu kernel driver core. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: init return value in amdgpu_ttm_clear_buffer","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21987","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21989","cve":"CVE-2025-21989","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix missing .is_two_pixels_per_container","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21989","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21990","cve":"CVE-2025-21990","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu firmware, ACPI and IP-block initialisation reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: NULL-check BO's backing store when determining GFX12 PTE flags","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21990","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-22093","cve":"CVE-2025-22093","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: avoid NPD when ASIC does not support DMUB","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22093","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23137","cve":"CVE-2025-23137","aliases":[],"title":"Linux cpufreq/amd-pstate - missing NULL check in amd_pstate_update: amd_pstate_update() dereferences the cpufreq policy without checking it for NULL, panicking the host. The CPU…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux cpufreq/amd-pstate - missing NULL check in amd_pstate_update","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"amd_pstate_update() dereferences the cpufreq policy without checking it for NULL, panicking the host. The CPU frequency driver runs constantly on every AMD node, so a NULL dereference here takes the machine down and every GPU job on it with no warning and no attacker involvement.","attack_vector":"Local, on the amd-pstate update path. Reachable through normal frequency-governor activity.","remediation":"Distro kernel update plus reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23137"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23245","cve":"CVE-2025-23245","aliases":[],"title":"vGPU Manager: Local privesc via improper file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Local privesc via improper file permissions","attack_vector":"Local operator on the hypervisor","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23245","https://github.com/NVIDIA/product-security/tree/main/2025/5630"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-732"]},{"id":"CVE-2025-23246","cve":"CVE-2025-23246","aliases":[],"title":"vGPU Manager: Host DoS via resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Host DoS via resource exhaustion","attack_vector":"Tenant VM guest","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23246","https://github.com/NVIDIA/product-security/tree/main/2025/5630"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-400"]},{"id":"CVE-2025-23261","cve":"CVE-2025-23261","aliases":[],"title":"Cumulus Linux: Sensitive data exposure on the switch","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Cumulus Linux","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Sensitive data exposure on the switch","attack_vector":"Authenticated switch user","remediation":"Upgrade Cumulus Linux; rolling switch upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23261","https://github.com/NVIDIA/product-security/tree/main/2025/5655"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-532"]},{"id":"CVE-2025-23285","cve":"CVE-2025-23285","aliases":[],"title":"vGPU Manager: Local privesc via file permissions","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"vGPU Manager","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Local privesc via file permissions","attack_vector":"Local operator on the hypervisor","remediation":"vGPU Manager upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23285","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-732"]},{"id":"CVE-2025-23300","cve":"CVE-2025-23300","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Allocating a specific GPU memory resource drives the kernel driver into a null-pointer dereference, crashing…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Allocating a specific GPU memory resource drives the kernel driver into a null-pointer dereference, crashing the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23300","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-23330","cve":"CVE-2025-23330","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A null-pointer dereference in the Linux display driver crashes the node from an unprivileged local account.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in the Linux display driver crashes the node from an unprivileged local account. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23330","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-27249","cve":"CVE-2025-27249","aliases":[],"title":"Intel Gaudi software suite (SynapseAI stack): An authenticated local user can drive the Gaudi software stack into unbounded resource consumption and deny…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi software suite (SynapseAI stack)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An authenticated local user can drive the Gaudi software stack into unbounded resource consumption and deny the accelerator to everything else on the node. On a bin-packed training cluster that is one tenant stalling a whole 8-card box.","attack_vector":"A local authenticated user on the node, which on most Gaudi deployments means anyone with a container that has the Gaudi devices mapped in.","remediation":"Upgrade the Gaudi software suite to 1.21.0 or later. Userspace and driver package update; plan a node drain because the habanalabs driver has to be reloaded, but no BIOS or accelerator firmware flash is required.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-27249","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01374.html"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-33237","cve":"CVE-2025-33237","aliases":[],"title":"GPU Display Driver: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33237","https://github.com/NVIDIA/product-security/tree/main/2026/5747"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37754","cve":"CVE-2025-37754","aliases":[],"title":"Linux i915 GPU kernel driver (HuC firmware load): The HuC delayed-loading fence is not released when probe fails early, leaving the driver in a broken state.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux i915 GPU kernel driver (HuC firmware load)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The HuC delayed-loading fence is not released when probe fails early, leaving the driver in a broken state. Matters on nodes where the HuC firmware blob is missing or mismatched, which is a common outcome of an incomplete linux-firmware package in a slim container host image.","attack_vector":"Triggered on driver probe - so on node boot, not by a tenant.","remediation":"Kernel update plus reboot. Also confirm the linux-firmware package on the node actually contains the matching GuC/HuC blobs for the installed silicon; a missing blob is what exposes this path in the first place.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37754","https://git.kernel.org/stable/c/4bd4bf79bcfe101f0385ab81dbabb6e3f7d96c00"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37766","cve":"CVE-2025-37766","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37766","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37767","cve":"CVE-2025-37767","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37767","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37768","cve":"CVE-2025-37768","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37768","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37769","cve":"CVE-2025-37769","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm/smu11): A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm/smu11)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm/smu11: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37769","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37770","cve":"CVE-2025-37770","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37770","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37771","cve":"CVE-2025-37771","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu power management (SMU/powerplay), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/pm: Prevent division by zero","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37771","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37852","cve":"CVE-2025-37852","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: handle amdgpu_cgs_create_device() errors in amd_powerplay_create()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37852","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37853","cve":"CVE-2025-37853","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: debugfs hang_hws skip GPU with MES","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37853","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37855","cve":"CVE-2025-37855","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Guard Possible Null Pointer Dereference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37855","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37870","cve":"CVE-2025-37870","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: prevent hang on link training fail","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37870","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-37965","cve":"CVE-2025-37965","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix invalid context error in dml helper","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37965","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38011","cve":"CVE-2025-38011","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: csa unmap use uninterruptible lock","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38011","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38021","cve":"CVE-2025-38021","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix null check of pipe_ctx->plane_state for update_dchubp_dpp","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38021","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38145","cve":"CVE-2025-38145","aliases":[],"title":"ASPEED LPC snoop driver (drivers/soc/aspeed/aspeed-lpc-snoop.c): Under memory pressure an allocation in the LPC snoop setup path returns NULL and is dereferenced, oopsing the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED LPC snoop driver (drivers/soc/aspeed/aspeed-lpc-snoop.c)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Under memory pressure an allocation in the LPC snoop setup path returns NULL and is dereferenced, oopsing the BMC kernel. LPC snoop is the channel that captures host POST codes and BIOS progress, so the practical loss is the BMC panicking during host boot - exactly the window where an operator is watching POST codes to diagnose a node that will not come up. Availability only, but on a large fleet the BMC dropping out mid-boot means losing remote power control on a node you now have to touch physically.","attack_vector":"Local on the BMC, and needs the BMC to be under memory pressure when the snoop channel is enabled. Practically this surfaces as a reliability bug rather than something an external attacker drives, though anything that inflates BMC memory usage (a pre-auth bmcweb allocation bug, for example) raises the odds.","remediation":"Kernel one-liner, backported to stable. Reaches nodes only via a BMC firmware image update: per-node, out-of-band, ODM-rebase dependent. No config workaround worth the tradeoff - disabling LPC snoop costs you host POST-code visibility, which is one of the main reasons the BMC is there. Low enough severity that batching it into your next scheduled firmware refresh is the right call.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38145","https://git.kernel.org/stable/c/c550999f939b529d28a914d5034cc4290066aea6"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-38205","cve":"CVE-2025-38205","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A division by zero in the amdgpu display core (DC/DM), reachable with attacker-influenced parameters. The kernel takes a divide fault and the node goes down, taking every co-resident job with it. Upstream fix: drm/amd/display: Avoid divide by zero by initializing dummy pitch to 1","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38205","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38254","cve":"CVE-2025-38254","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add sanity checks for drm_edid_raw()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38254","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38319","cve":"CVE-2025-38319","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pp): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pp)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pp: Fix potential NULL pointer dereference in atomctrl_initialize_mc_reg_table","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38319","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38360","cve":"CVE-2025-38360","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add more checks for DSC / HUBP ONO guarantees","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38360","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38362","cve":"CVE-2025-38362","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null pointer check for get_first_active_display()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38362","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38426","cve":"CVE-2025-38426","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: Add basic validation for RAS header","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38426","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38487","cve":"CVE-2025-38487","aliases":[],"title":"ASPEED LPC snoop driver channel teardown (drivers/soc/aspeed/aspeed-lpc-snoop.c): Unbinding the LPC snoop driver tears down channels that were never brought up, dereferencing NULL and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED LPC snoop driver channel teardown (drivers/soc/aspeed/aspeed-lpc-snoop.c)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Unbinding the LPC snoop driver tears down channels that were never brought up, dereferencing NULL and panicking the BMC kernel. The reproducer is a single write to the driver's sysfs unbind file. Anyone with root on the BMC can hard-crash the management processor on demand; more usefully for an operator, it fires during ordinary driver reload and platform-teardown sequences, so it shows up as BMC instability on ASPEED platforms that only wire up a subset of the snoop channels.","attack_vector":"Root on the BMC (write access to the platform driver's sysfs bind/unbind), or any BMC-side maintenance flow that unbinds the driver. Not reachable from the host or the network on its own.","remediation":"Kernel patch, backported to stable. Delivered only in a new BMC firmware image - per-node, out-of-band flash, gated on the ODM. Low priority as a standalone item; treat it as one more reason not to run BMC firmware images that are years behind upstream, and roll it in with the other lpc-snoop and video-engine fixes in a single flash rather than a dedicated campaign.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38487","https://git.kernel.org/stable/c/9e1d2b97f5e2a36a2fd30a8bd30ead9dac5e3a51"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-38506","cve":"CVE-2025-38506","aliases":[],"title":"Linux KVM - CPU soft lockup setting per-page memory attributes on large SNP guests: Running an SEV-SNP guest with a large memory footprint - 1TB and up, which is exactly the shape of an AI…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - CPU soft lockup setting per-page memory attributes on large SNP guests","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Running an SEV-SNP guest with a large memory footprint - 1TB and up, which is exactly the shape of an AI training VM - drives the host into CPU soft lockups while KVM walks per-page memory attributes without rescheduling. The host stalls, the watchdog fires, and co-resident workloads suffer. A tenant does not need to attack anything: simply asking for a big confidential VM is enough.","attack_vector":"Triggered by a guest with a very large memory allocation. Tenant-reachable through normal VM sizing, which makes it as much a capacity-planning hazard as a security one.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Directly relevant to AI hosts, where terabyte-class confidential VMs are the norm rather than the exception - do not dismiss this as a corner case on a GPU fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38506"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38518","cve":"CVE-2025-38518","aliases":[],"title":"Linux x86/CPU/AMD - INVLPGB on Zen 2 (Cyan Skillfish): Using broadcast TLB invalidation (INVLPGB) on affected Zen 2 parts oopses the system. TLB invalidation is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/CPU/AMD - INVLPGB on Zen 2 (Cyan Skillfish)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Using broadcast TLB invalidation (INVLPGB) on affected Zen 2 parts oopses the system. TLB invalidation is core memory-management machinery, so a defect here is both a stability problem and, in principle, a correctness problem for the mappings that separate address spaces.","attack_vector":"Local, triggered by normal kernel memory management on affected silicon rather than by an attacker.","remediation":"Fixed in the Linux kernel by disabling INVLPGB on affected parts. Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38518"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38520","cve":"CVE-2025-38520","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Don't call mmput from MMU notifier callback","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38520","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38705","cve":"CVE-2025-38705","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: fix null pointer access","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38705","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39675","cve":"CVE-2025-39675","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add null pointer check in mod_hdcp_hdcp1_create_session()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39675","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39693","cve":"CVE-2025-39693","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid a NULL pointer dereference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39693","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39705","cve":"CVE-2025-39705","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: fix a Null pointer dereference vulnerability","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39705","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39706","cve":"CVE-2025-39706","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Destroy KFD debugfs after destroy KFD wq","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39706","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39707","cve":"CVE-2025-39707","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: check if hubbub is NULL in debugfs/amdgpu_dm_capabilities","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39707","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39762","cve":"CVE-2025-39762","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: add null check","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39762","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-39936","cve":"CVE-2025-39936","aliases":[],"title":"Linux crypto/ccp - SEV platform shutdown error handling: The ccp driver's SEV/SNP platform shutdown path could be called without a valid error pointer, dereferencing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux crypto/ccp - SEV platform shutdown error handling","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The ccp driver's SEV/SNP platform shutdown path could be called without a valid error pointer, dereferencing it and panicking the host. Since this is on the SEV platform teardown path, the crash lands during operations like driver unload or SNP re-initialisation - the exact moments you are already doing maintenance on a confidential-computing host.","attack_vector":"Local, in the host's SEV platform management path; reachable by whatever drives SEV init/shutdown, i.e. host administration rather than tenants.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-39936"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-57275","cve":"CVE-2025-57275","aliases":["SPDK NVMe-oF target buffer overflow","SPDK lib/nvmf"],"title":"SPDK (Storage Performance Development Kit) 25.05 - NVMe-oF target, lib/nvmf: TENANT ISOLATION: A buffer overflow in the NVMe-oF target component of SPDK 25.05. SPDK is the userspace…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"SPDK (Storage Performance Development Kit) 25.05 - NVMe-oF target, lib/nvmf","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: A buffer overflow in the NVMe-oF target component of SPDK 25.05. SPDK is the userspace, poll-mode NVMe-oF target most commonly deployed by neoclouds and storage vendors precisely because it outperforms the kernel target, so it fronts tenant namespaces in exactly the environments this database is aimed at. NVD scores it as requiring high privileges with limited integrity impact plus availability loss, so the realistic outcome is a target crash - taking every attached tenant's I/O with it - rather than a clean takeover. Worth tracking because NeVerMore separately verified seven NVMe-oF protocol attacks against SPDK, so the target's overall exposure is broader than this single defect.","attack_vector":"Reached through the NVMe-oF target path in lib/nvmf on SPDK 25.05. NVD's vector puts it at network-adjacent reachability with high privileges required, which in practice means an authenticated or otherwise privileged initiator context rather than an anonymous peer. Public detail is thin - the advisory text is a one-line description with no reproducer.","remediation":"Upgrade SPDK past 25.05 and restart the target process - no kernel change, no host reboot, but the restart drops all NVMe-oF connections, so run it behind multipath initiators or during a maintenance window per storage node. Combine with the NVMe-oF hardening in the discovery-controller entry: in-band DH-HMAC-CHAP, per-subsystem host allow-lists, and discovery on a management-only interface. Since SPDK is a library embedded in vendor and in-house appliances, check with your storage vendor which SPDK release their firmware ships rather than assuming the host package version is what is running.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-57275","https://arxiv.org/abs/2202.08080"],"status":"curated"},{"id":"CVE-2025-71293","cve":"CVE-2025-71293","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/ras): A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/ras)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/ras: Move ras data alloc before bad page check","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71293","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-71294","cve":"CVE-2025-71294","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix NULL pointer issue buffer funcs","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-71294","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-8404","cve":"CVE-2025-8404","aliases":[],"title":"A shared library inside Supermicro BMC firmware that parses request headers: An authenticated attacker overflows a stack buffer during header parsing and executes code in the BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"A shared library inside Supermicro BMC firmware that parses request headers","year":"2025","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An authenticated attacker overflows a stack buffer during header parsing and executes code in the BMC firmware operating system. The shared-library location is what makes this worth flagging separately: fixing one web endpoint does not fix it, and the same primitive is likely reachable from whichever BMC service an operator has left enabled. Outcome is the usual BMC-root outcome - out-of-band power, console, virtual media and firmware persistence under the host. Because it is shared, the same overflow is reachable from more than one front-end service on the controller rather than from a single CGI endpoint.","attack_vector":"Any authenticated session that reaches a BMC service using this library over the network. Because it is shared code, restricting one interface does not close it.","remediation":"Firmware flash from Supermicro's November 2025 BMC/IPMI batch. Disabling individual BMC services is a weaker mitigation than usual here, since the bug lives in shared parsing code rather than one handler - so treat network isolation of the management VLAN plus per-node unique BMC credentials as the interim control, and prioritise the flash. Expect the fixed image to be board-specific.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-8404","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/8xxx/CVE-2025-8404.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-12480","cve":"CVE-2026-12480","aliases":[],"title":"Keras (HDF5 ExternalLink, incomplete fix): Arbitrary HDF5 file read","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Keras (HDF5 ExternalLink, incomplete fix)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Arbitrary HDF5 file read; incomplete fix for CVE-2026-1669","attack_vector":"Customer-supplied model file","remediation":"Upgrade past 3.13.2","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-12480"],"status":"curated"},{"id":"CVE-2026-22703","cve":"CVE-2026-22703","aliases":[],"title":"cosign / sigstore: A crafted bundle verifies successfully even though the embedded Rekor entry does not reference the artifact","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A crafted bundle verifies successfully even though the embedded Rekor entry does not reference the artifact; signature policy bypass","attack_vector":"Malicious image","remediation":"Upgrade cosign to 2.6.2/3.0.4+; re-verify admitted images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-22703"],"status":"curated"},{"id":"CVE-2026-23163","cve":"CVE-2026-23163","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix NULL pointer dereference in amdgpu_gmc_filter_faults_remove","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23163","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23213","cve":"CVE-2026-23213","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu power management (SMU/powerplay) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/pm: Disable MMIO access during SMU Mode 1 reset","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23213","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23338","cve":"CVE-2026-23338","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A race condition or locking defect in the amdgpu user-mode queues (doorbell submission path). Concurrent…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu user-mode queues (doorbell submission path). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu/userq: Do not allow userspace to trivially triger kernel warnings","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23338","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23358","cve":"CVE-2026-23358","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu RAS…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu RAS / GPU reset and recovery path. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: Fix error handling in slot reset","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23358","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23435","cve":"CVE-2026-23435","aliases":[],"title":"Linux perf/x86 - event pointer setup ordering in x86_pmu_enable(): A NULL pointer dereference in the x86 PMU enable path, reported from a production AMD EPYC system.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux perf/x86 - event pointer setup ordering in x86_pmu_enable()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the x86 PMU enable path, reported from a production AMD EPYC system. Performance counters are enabled by every profiling and observability agent on a fleet, so this crashes hosts through the monitoring stack rather than through anything a tenant did - and it takes co-resident GPU jobs with it.","attack_vector":"Local, through perf event enablement. Reachable by whatever has perf access, which on many clusters includes node-level observability agents and, if perf_event_paranoid is relaxed, tenants.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. Check perf_event_paranoid on GPU nodes: if you have loosened it so tenants can profile their own kernels - which is a reasonable thing to want on an AI cluster - you have also widened who can reach this.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23435"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23468","cve":"CVE-2026-23468","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: Limit BO list entry count to prevent resource exhaustion","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23468","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-24160","cve":"CVE-2026-24160","aliases":[],"title":"TensorRT-LLM: DoS (null deref in tensor ops)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT-LLM","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"DoS (null deref in tensor ops)","attack_vector":"Malicious inference input","remediation":"Bump TensorRT-LLM; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24160","https://github.com/NVIDIA/product-security/tree/main/2026/5805"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-690"]},{"id":"CVE-2026-31460","cve":"CVE-2026-31460","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: check if ext_caps is valid in BL setup","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31460","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31461","cve":"CVE-2026-31461","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: Fix drm_edid leak in amdgpu_dm","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31461","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31462","cve":"CVE-2026-31462","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: prevent immediate PASID reuse case","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31462","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31540","cve":"CVE-2026-31540","aliases":[],"title":"Linux i915 GPU kernel driver (submission backend setup): i915 dereferences the submission backend before checking it is set, which happens when the GuC/HuC firmware…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (submission backend setup)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"i915 dereferences the submission backend before checking it is set, which happens when the GuC/HuC firmware binaries are absent. Node panics on GPU init instead of degrading. Common on minimal host images that omit linux-firmware.","attack_vector":"No attacker; triggered by a node image missing GPU firmware blobs.","remediation":"Kernel update plus reboot, and fix the node image so linux-firmware carries the GuC/HuC blobs for the installed GPUs.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31540","https://git.kernel.org/stable/c/0162ab3220bac870e43e229e6e3024d1a21c3f26"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31591","cve":"CVE-2026-31591","aliases":[],"title":"Linux KVM/SEV - vCPU locking when synchronizing VMSAs for SNP launch finish: KVM did not lock all vCPUs while synchronising and encrypting VMSAs at SNP launch finish, so vCPU state could…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM/SEV - vCPU locking when synchronizing VMSAs for SNP launch finish","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM did not lock all vCPUs while synchronising and encrypting VMSAs at SNP launch finish, so vCPU state could change underneath the encryption step. The VMSA is what the launch measurement covers; if it can move while being measured, the attestation you hand the tenant does not necessarily describe the VM that actually ran.","attack_vector":"Through the KVM SNP launch path, from the VMM process.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31591"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31593","cve":"CVE-2026-31593","aliases":[],"title":"Linux KVM - VMSA sync on an already-launched SEV vCPU: KVM allowed synchronising vCPU state into the VMSA after the VMSA had already been encrypted and the guest…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - VMSA sync on an already-launched SEV vCPU","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM allowed synchronising vCPU state into the VMSA after the VMSA had already been encrypted and the guest launched. Writing to an encrypted VMSA behind the guest's back corrupts the confidential vCPU state that attestation covered - the guest is no longer the thing that was measured, and the failure is silent rather than loud.","attack_vector":"Via the KVM ioctl surface, from the VMM process managing the guest.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31593"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31689","cve":"CVE-2026-31689","aliases":[],"title":"Linux EDAC/mc - error path ordering in edac_mc_alloc(): When a private-data allocation fails in edac_mc_alloc(), the error path unwinds in the wrong order and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux EDAC/mc - error path ordering in edac_mc_alloc()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"When a private-data allocation fails in edac_mc_alloc(), the error path unwinds in the wrong order and touches memory it has already released. EDAC is the memory error-detection and correction subsystem - the thing telling you which DIMM is going bad on a node full of expensive HBM-adjacent DRAM - so a defect in its allocation path costs you both stability and the RAS visibility you were relying on.","attack_vector":"Local, on the EDAC allocation error path - hit under memory pressure rather than by an attacker.","remediation":"Distro kernel update plus reboot; no firmware step. Worth taking on any fleet where you drive node retirement off EDAC data.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31689"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-31765","cve":"CVE-2026-31765","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: Change AMDGPU_VA_RESERVED_TRAP_SIZE to 64KB","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-31765","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43034","cve":"CVE-2026-43034","aliases":[],"title":"Linux bnxt_en driver (backing store type from firmware response): A second firmware-controlled-index bug in the same driver: the backing-store type returned in a firmware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux bnxt_en driver (backing store type from firmware response)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A second firmware-controlled-index bug in the same driver: the backing-store type returned in a firmware response is stored and later used to index a fixed array. Same class as the DBG_BUF_PRODUCER issue and listed separately because it is fixed by a different commit — an operator matching only one CVE will still be running the other.","attack_vector":"The NIC firmware's response content.","remediation":"Kernel/driver upgrade plus host reboot; same patch cycle as the other bnxt_en firmware-input issues.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43034"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43130","cve":"CVE-2026-43130","aliases":[],"title":"Linux iommu/vt-d (dev-IOTLB flush in scalable mode): TENANT ISOLATION: the scalable-mode half of the device-IOTLB invalidation problem — ATS invalidation is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux iommu/vt-d (dev-IOTLB flush in scalable mode)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: the scalable-mode half of the device-IOTLB invalidation problem — ATS invalidation is issued or skipped based on device accessibility, and the earlier fix in this area left a gap. Scalable mode is what modern Intel platforms use for PASID-based device assignment, so this is the current-generation path for handing NIC and accelerator functions to tenants. Same underlying concern: a device retaining stale translations after the host revoked them.","attack_vector":"A tenant with a scalable-mode-assigned, ATS-capable PCIe function.","remediation":"Kernel upgrade plus host reboot, rolling across passthrough-capable nodes. Verify after patching that IOMMU is in enforcing (not passthrough/`iommu=pt`) mode for tenant-assigned devices — a surprising number of performance-tuned GPU hosts run with IOMMU translation effectively disabled, which makes this class of bug moot only because the isolation was never there.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43130"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43131","cve":"CVE-2026-43131","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm): A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amd/pm)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu power management (SMU/powerplay). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/pm: Fix null pointer dereference issue","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43131","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43161","cve":"CVE-2026-43161","aliases":[],"title":"Linux iommu/vt-d (dev-IOTLB flush for passed-through PCIe devices): TENANT ISOLATION: the Intel IOMMU driver skips device-IOTLB invalidation for PCIe endpoints that are…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux iommu/vt-d (dev-IOTLB flush for passed-through PCIe devices)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: the Intel IOMMU driver skips device-IOTLB invalidation for PCIe endpoints that are ATS-enabled and passed through to userspace — the exact configuration used for SR-IOV NIC VFs handed to a VM, and for DPDK and RDMA userspace drivers on a GPU node. The device-IOTLB is the device's own cached copy of the address translations; if it is not flushed when a mapping is revoked, the device can keep reaching memory the kernel believes it has taken away. That is the core mechanism DMA isolation depends on, and passthrough NICs are precisely the devices you hand to untrusted tenants. Companion issue CVE-2026-43130 covers the scalable-mode variant.","attack_vector":"A tenant holding a passed-through, ATS-enabled PCIe device — an SR-IOV VF assigned to their VM, or a DPDK/RDMA userspace-bound NIC — able to exercise the device after a mapping has been torn down.","remediation":"Kernel upgrade plus host reboot — rolling across every node that does device passthrough, which on a GPU cloud is all of them. Nothing to flash. If you cannot patch immediately, the meaningful mitigation is disabling ATS on passed-through endpoints (a BIOS/kernel-parameter change, at a measurable performance cost) or not passing devices through to untrusted tenants at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43161","https://nvd.nist.gov/vuln/detail/CVE-2026-43130"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43191","cve":"CVE-2026-43191","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Adjust PHY FSM transition to TX_EN-to-PLL_ON for TMDS on DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43191","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43195","cve":"CVE-2026-43195","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate user queue size constraints","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43195","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43243","cve":"CVE-2026-43243","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Add signal type check for dcn401 get_phyd32clk_src","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43243","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43298","cve":"CVE-2026-43298","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: Skip vcn poison irq release on VF","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43298","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43305","cve":"CVE-2026-43305","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Fix mismatched unlock for DMUB HW lock in HWSS fast path","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43305","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43318","cve":"CVE-2026-43318","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: fix sync handling in amdgpu_dma_buf_move_notify","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43318","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43320","cve":"CVE-2026-43320","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix dsc eDP issue","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43320","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43337","cve":"CVE-2026-43337","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Fix NULL pointer dereference in dcn401_init_hw()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43337","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43367","cve":"CVE-2026-43367","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd: Fix a few more NULL pointer dereference in device cleanup","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43367","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43369","cve":"CVE-2026-43369","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amd): A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu kernel driver core. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd: Fix NULL pointer dereference in device cleanup","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43369","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43398","cve":"CVE-2026-43398","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: add upper bound check on user inputs in wait ioctl","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43398","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43399","cve":"CVE-2026-43399","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A memory or reference-count leak in the amdgpu user-mode queues (doorbell submission path). Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu user-mode queues (doorbell submission path). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/userq: Fix reference leak in amdgpu_userq_wait_ioctl","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43399","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43400","cve":"CVE-2026-43400","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: add upper bound check on user inputs in signal ioctl","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43400","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-43444","cve":"CVE-2026-43444","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Unreserve bo if queue update failed","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-43444","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-45947","cve":"CVE-2026-45947","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix memory leak in amdgpu_acpi_enumerate_xcc()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45947","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-45976","cve":"CVE-2026-45976","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A memory or reference-count leak in the amdgpu RAS / GPU reset and recovery path. Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu RAS / GPU reset and recovery path. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix memory leak in amdgpu_ras_init()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45976","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-45979","cve":"CVE-2026-45979","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: clean up the amdgpu_cs_parser_bos","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-45979","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46220","cve":"CVE-2026-46220","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/sdma4): A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu/sdma4)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu GEM/VM/command-submission ioctl surface reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma4: replace BUG_ON with WARN_ON in fence emission","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46220","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46229","cve":"CVE-2026-46229","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdkfd…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdkfd (KFD compute driver, /dev/kfd). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdkfd: Clear VRAM on allocation to prevent stale data exposure","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46229","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46245","cve":"CVE-2026-46245","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix dc_link NULL handling in HPD init","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46245","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-46276","cve":"CVE-2026-46276","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu RAS…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu RAS / GPU reset and recovery path. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: fix zero-size GDS range init on RDNA4","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-46276","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-47630","cve":"CVE-2026-47630","aliases":[],"title":"NVIDIA Triton Inference Server: An absolute path traversal reachable from a local low-privileged account reaches code execution. On a shared…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"An absolute path traversal reachable from a local low-privileged account reaches code execution. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5865. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47630","https://github.com/NVIDIA/product-security/tree/main/2026/5865"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-36"]},{"id":"CVE-2026-53121","cve":"CVE-2026-53121","aliases":[],"title":"Linux amd-pstate - memory leak in amd_pstate_epp_cpu_init(): On failure to set the energy-performance preference, amd_pstate_epp_cpu_init() returns without freeing what…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux amd-pstate - memory leak in amd_pstate_epp_cpu_init()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"On failure to set the energy-performance preference, amd_pstate_epp_cpu_init() returns without freeing what it allocated. Another slow leak in the AMD CPU frequency driver, hit on the error path - which is the path a misconfigured or partially-supported platform takes repeatedly rather than once.","attack_vector":"Local, on the amd-pstate initialisation error path.","remediation":"Distro kernel update plus reboot. Batch with the other amd-pstate fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53121"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53135","cve":"CVE-2026-53135","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL deref and buffer over-read in SDP debugfs","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53135","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53142","cve":"CVE-2026-53142","aliases":[],"title":"Linux drm/xe GPU kernel driver (suspend/shutdown without display): The xe driver oopses on suspend or shutdown on configurations with no display probed - which is exactly the…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux drm/xe GPU kernel driver (suspend/shutdown without display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"The xe driver oopses on suspend or shutdown on configurations with no display probed - which is exactly the headless datacenter GPU configuration. Effect is a dirty shutdown rather than a clean one, which risks filesystem and checkpoint state on nodes that reboot under automation.","attack_vector":"No attacker; fires on normal shutdown of a headless xe node.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53142","https://git.kernel.org/stable/c/0f68ddfaaebfbb5581ee931779757d31f4dc9e24"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53144","cve":"CVE-2026-53144","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: fix NULL dereference in get_queue_ids()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53144","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53283","cve":"CVE-2026-53283","aliases":[],"title":"Linux iommu/amd - devid bounds check in __rlookup_amd_iommu(): MULTI-TENANT ISOLATION: The AMD IOMMU driver looked up device IDs without bounds-checking them, so a device…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux iommu/amd - devid bounds check in __rlookup_amd_iommu()","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The AMD IOMMU driver looked up device IDs without bounds-checking them, so a device ID outside the expected range indexes past the array. Device enumeration walks every device on the PCI bus, and on a dense GPU node that bus is crowded - many accelerators, switches, NICs and bridges. An out-of-bounds read in the IOMMU's device lookup is a kernel memory-safety issue in the component enforcing DMA isolation.","attack_vector":"Local, triggered during IOMMU device registration and lookup. Influenced by what is on the PCI bus, so a malicious or malfunctioning device - or a device presented by a compromised BMC - can reach it.","remediation":"Fixed in the Linux kernel. Distro kernel update plus a node reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53283"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53285","cve":"CVE-2026-53285","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu display core (DC/DM). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amd/display: Wrap DCN32 phantom-plane allocation in DC_RUN_WITH_PREEMPTION_ENABLED","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53285","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53293","cve":"CVE-2026-53293","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu GEM/VM/command-submission ioctl surface. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: fix AMDGPU_INFO_READ_MMR_REG","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53293","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53313","cve":"CVE-2026-53313","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid NULL dereference in dc_dmub_srv error paths","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53313","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53315","cve":"CVE-2026-53315","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras): A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/ras: Fix NULL deref in ras_core_get_utc_second_timestamp()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53315","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53316","cve":"CVE-2026-53316","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras): A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amd/ras)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/ras: Fix NULL deref in ras_core_ras_interrupt_detected()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53316","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53345","cve":"CVE-2026-53345","aliases":[],"title":"Linux KVM - dirty-page tracking without a vCPU on a dying VM: KVM warned (and on panic_on_warn hosts, panicked) when a page was marked dirty without an associated vCPU…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM - dirty-page tracking without a vCPU on a dying VM","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"KVM warned (and on panic_on_warn hosts, panicked) when a page was marked dirty without an associated vCPU while a VM was being torn down. A guest that can arrange the teardown timing turns a warning into a host crash on any fleet running panic_on_warn - which plenty of hardened kernels do.","attack_vector":"From inside a guest, by timing VM teardown. Tenant-reachable.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance. If you run panic_on_warn on GPU hosts for crash-dump fidelity, note that it converts warning-class kernel bugs like this into fleet availability incidents - worth reviewing that setting alongside the patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53345"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-53376","cve":"CVE-2026-53376","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":5.5,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdkfd (KFD compute driver, /dev/kfd). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdkfd: Add upper bound check for num_of_nodes","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-53376","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-14456","cve":"CVE-2019-14456","aliases":[],"title":"Opengear console server (serial port logging): Stored XSS injected from a device *connected to* a serial port — a compromised switch can attack the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Opengear console server (serial port logging)","year":"2019","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Stored XSS injected from a device *connected to* a serial port — a compromised switch can attack the operator's console-server UI, inverting the expected trust direction","attack_vector":"Local device to OOB management UI","remediation":"Console-server firmware upgrade to 4.5.0+; notable as an example of the OOB network being attackable from the devices it manages","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14456"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-8558","cve":"CVE-2020-8558","aliases":[],"title":"Kubernetes (kubelet/kube-proxy): Node's 127.0.0.1-bound services reachable from adjacent hosts and pods","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet/kube-proxy)","year":"2020","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Node's 127.0.0.1-bound services reachable from adjacent hosts and pods; commonly reaches an unauthenticated kubelet or etcd","attack_vector":"Any pod on the node, or an adjacent host on the node network","remediation":"Rolling kubelet/kube-proxy upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8558"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-21818","cve":"CVE-2022-21818","aliases":[],"title":"NVIDIA License System - DLS virtual appliance: Installation scripts on the DLS appliance leave other users' credentials readable to any signed-in portal…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - DLS virtual appliance","year":"2022","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Installation scripts on the DLS appliance leave other users' credentials readable to any signed-in portal user, giving lateral privilege escalation inside your licensing infrastructure.","attack_vector":"Network, authenticated as any portal user. Anyone you gave a licensing-portal account to.","remediation":"Patch the DLS appliance per bulletin 5319 and rotate every credential the appliance held - the patch does not undo the exposure. Cost: appliance restart; vGPU guests keep running on cached licences through a short DLS outage.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21818","https://github.com/NVIDIA/product-security/tree/main/2022/5319"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-312"]},{"id":"CVE-2024-0103","cve":"CVE-2024-0103","aliases":[],"title":"Triton Inference Server: Insufficient access-control granularity","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Insufficient access-control granularity","attack_vector":"Authenticated inference client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0103","https://github.com/NVIDIA/product-security/tree/main/2024/5546"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-1419"]},{"id":"CVE-2024-9526","cve":"CVE-2024-9526","aliases":[],"title":"Kubeflow (Pipelines UI): Stored XSS in the pipeline view","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"Kubeflow (Pipelines UI)","year":"2024","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Stored XSS in the pipeline view","attack_vector":"Tenant-supplied pipeline definition rendered to another user","remediation":"Upgrade; multi-tenant Kubeflow UIs share an origin","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-9526"],"status":"curated"},{"id":"CVE-2025-48515","cve":"CVE-2025-48515","aliases":[],"title":"AMD Secure Processor bootloader - SPIROM upgrade path: MULTI-TENANT ISOLATION: An attacker who can drive the SPIROM upgrade path can pass unsanitised parameters to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor bootloader - SPIROM upgrade path","year":"2025","cvss_score":5.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An attacker who can drive the SPIROM upgrade path can pass unsanitised parameters to the ASP bootloader and overwrite memory, reaching arbitrary code execution in the secure processor. This is the classic firmware-update-as-attack-surface problem: the mechanism you use to patch the platform is itself the way in.","attack_vector":"Local, requires access to the SPI ROM upgrade mechanism - typically root plus flash write, or a compromised BMC that can drive host SPI.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Worth pairing with BMC hardening: on most server designs the BMC can write host SPI, so a BMC compromise reaches this directly. Restrict who can invoke firmware updates and require signed update packages end to end.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48515","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-7623","cve":"CVE-2025-7623","aliases":[],"title":"Supermicro BMC SMASH-CLP shell on MBD-X13SEDW-F: Full control of the instruction pointer inside the BMC's firmware OS from a shell that operators routinely…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC SMASH-CLP shell on MBD-X13SEDW-F","year":"2025","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Full control of the instruction pointer inside the BMC's firmware OS from a shell that operators routinely hand to junior staff and monitoring tooling. The attacker converts a limited management login into arbitrary code on the controller, and from there into the standard BMC prize set: power control, console capture, virtual media, and firmware-level persistence. The published CVSS understates this - the vector was scored conservatively, but the described primitive is return-address control. A stack buffer overflow reached by a crafted SMASH command, with control of the saved return address and registers.","attack_vector":"An authenticated low-privilege BMC account with SSH access to the controller. Any operator-tier credential works; no administrator role is needed.","remediation":"Firmware flash from Supermicro's November 2025 BMC/IPMI advisory batch, matched to the board SKU. The cheap and immediate mitigation is config-only: turn off SSH/SMASH on the BMC if your management path is Redfish or IPMI-over-LAN. That single change also covers CVE-2026-3821 and the rest of the SMASH overflow cluster, so it is the highest-leverage action available before a flash window opens.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-7623","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/7xxx/CVE-2025-7623.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-32677","cve":"CVE-2026-32677","aliases":[],"title":"Intel Gaudi / gaudi-container-runtime: MULTI-TENANT ISOLATION: A path-traversal bug in the container runtime shim that wires Gaudi devices into…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Intel Gaudi / gaudi-container-runtime","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A path-traversal bug in the container runtime shim that wires Gaudi devices into containers lets a tenant who controls the container spec cause the runtime to touch host paths outside the container root. On a shared Gaudi box the runtime executes as root on the host, so this is the classic accelerator-runtime escape shape: tenant container -> host filesystem -> every other tenant's job on that node.","attack_vector":"Any tenant who can launch a container on a Gaudi node through the normal scheduler. No host account and no physical access required - the container spec is the attack surface.","remediation":"Upgrade gaudi-container-runtime to 1.24.0 or later across every Gaudi node. This is a host-side userspace package, so no BIOS, firmware or microcode update is involved, but the runtime binary is in the path of every new container start - roll it per node and restart the container engine, which means draining running jobs on that node.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-32677","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01489.html"],"status":"curated"},{"id":"CVE-2026-33726","cve":"CVE-2026-33726","aliases":[],"title":"Cilium: Ingress NetworkPolicies not enforced for pod traffic to L7 services","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"Ingress NetworkPolicies not enforced for pod traffic to L7 services","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-33726"],"status":"curated"},{"id":"CVE-2026-39350","cve":"CVE-2026-39350","aliases":[],"title":"Istio: serviceAccounts and notServiceAccounts in AuthorizationPolicy are evaluated incorrectly","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"serviceAccounts and notServiceAccounts in AuthorizationPolicy are evaluated incorrectly","attack_vector":"Any pod on the mesh","remediation":"Rolling istiod upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39350"],"status":"curated"},{"id":"CVE-2026-56743","cve":"CVE-2026-56743","aliases":[],"title":"Cilium: CIDR ipBlock rules without selectors generate a wildcard, over-permitting traffic","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2026","cvss_score":5.4,"severity":"medium","kev":false,"impact":"CIDR ipBlock rules without selectors generate a wildcard, over-permitting traffic","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-56743"],"status":"curated"},{"id":"CVE-2018-10892","cve":"CVE-2018-10892","aliases":[],"title":"Docker / moby: Default OCI spec does not mask /proc/acpi, so a container can change host hardware state","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2018","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Default OCI spec does not mask /proc/acpi, so a container can change host hardware state","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine or add explicit maskedPaths","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-10892"],"status":"curated"},{"id":"CVE-2019-9836","cve":"CVE-2019-9836","aliases":[],"title":"AMD Platform Security Processor - SEV key derivation (PSP firmware <= 0.17 build 11): MULTI-TENANT ISOLATION: The SEV implementation in PSP firmware used a broken elliptic-curve parameter check…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Platform Security Processor - SEV key derivation (PSP firmware <= 0.17 build 11)","year":"2019","cvss_score":5.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The SEV implementation in PSP firmware used a broken elliptic-curve parameter check, so an attacker with host privileges can run an invalid-curve attack against the platform Diffie-Hellman exchange and recover the SEV endorsement key. With that key the host can decrypt a guest's launch secret and read the encrypted VM's memory outright. If you sold confidential computing on top of SEV on affected firmware, the guarantee was not there.","attack_vector":"Local, requires hypervisor/host administrator privilege - which is exactly the party SEV is supposed to defend the guest against. No guest cooperation needed.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Because this sits inside the SEV-SNP trust boundary, the update also moves the platform's reported TCB version: after patching you must refresh VCEK certificates from AMD's KDS and update whatever attestation policy your tenants (or your own confidential-VM control plane) pin against, or every guest launch will start failing validation. Fixed in PSP/SEV firmware 0.17 build 22 and later. Any guest launched or attested on older firmware should be treated as having had no confidentiality guarantee - rotate the secrets those guests held rather than just patching forward.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-9836","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-29662","cve":"CVE-2020-29662","aliases":[],"title":"Harbor: Catalog registry API exposed on an unauthenticated path","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Catalog registry API exposed on an unauthenticated path; full image inventory disclosure","attack_vector":"Unauthenticated network","remediation":"Upgrade Harbor; put the registry behind auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-29662"],"status":"curated"},{"id":"CVE-2020-7202","cve":"CVE-2020-7202","aliases":["HPESBHF04069"],"title":"HPE iLO 4 / iLO 5 (unauthenticated information disclosure): An unauthenticated remote request pulls back the server serial number and other identifying detail from the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE iLO 4 / iLO 5 (unauthenticated information disclosure)","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An unauthenticated remote request pulls back the server serial number and other identifying detail from the iLO. Low impact on its own, high value as reconnaissance: it lets an attacker who can reach a management network enumerate exactly what hardware sits behind each iLO address, fingerprint generations, and pick which nodes are worth a real exploit - all without a single failed login to show up in an audit log. Covers ProLiant, Apollo, Synergy compute modules and Converged Systems, which is most of the HPE fleet shape a GPU operator would run.","attack_vector":"Anything routable to the iLO on the out-of-band management VLAN, unauthenticated. If any iLO is inadvertently internet-exposed, this is what a mass scanner harvests first.","remediation":"Flash iLO 5 to v2.31 or later and iLO 4 to v2.76 or later. Out-of-band, per-node, no host reboot and no drain. Given the low direct impact, most operators should fold this into the next scheduled iLO firmware campaign rather than running a dedicated one - but do treat any internet-reachable iLO as an emergency independent of this CVE.","references":["https://support.hpe.com/hpsc/doc/public/display?docLocale=en_US&docId=emr_na-hpesbhf04069en_us","https://nvd.nist.gov/vuln/detail/CVE-2020-7202"],"status":"curated"},{"id":"CVE-2020-8552","cve":"CVE-2020-8552","aliases":[],"title":"Kubernetes (kube-apiserver): Successful API requests can DoS the apiserver","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Successful API requests can DoS the apiserver","attack_vector":"Any authenticated cluster user","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8552"],"status":"curated"},{"id":"CVE-2021-1055","cve":"CVE-2021-1055","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Improper access control in the escape handler leaks information and lets an unprivileged caller crash the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Improper access control in the escape handler leaks information and lets an unprivileged caller crash the driver.","attack_vector":"Any local user with GPU device access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1055"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-22815","cve":"CVE-2021-22815","aliases":["SEVD-2021-313-03"],"title":"APC/Schneider Electric UPS, PDU, and cooling products using NMC2/NMC3 (Smart-UPS, Symmetra, Galaxy, rack PDUs, InRow cooling, NetBotz): The card's troubleshooting archive — a diagnostic bundle…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"APC/Schneider Electric UPS, PDU, and cooling products using NMC2/NMC3","year":"2021","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The card's troubleshooting archive — a diagnostic bundle that can contain configuration and log details — can be pulled off the device by someone who shouldn't have access to it, giving an attacker reconnaissance data useful for planning further attacks on that power/cooling unit.","attack_vector":"Network access to the card's web interface; the advisory describes this as an information-exposure issue reachable without full administrative rights.","remediation":"Firmware upgrade per Schneider's SEVD-2021-313-03 advisory (fixed AOS versions vary by card generation — NMC2 vs NMC3). This spans a very wide product line (UPS, rack PDUs, cooling, NetBotz), so treat it as a fleet-wide inventory-and-patch exercise rather than a one-off fix.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2021-313-03"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-45925","cve":"CVE-2021-45925","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC login): The login flow answers differently for real and fake usernames, so an unauthenticated attacker can enumerate…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC login)","year":"2021","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The login flow answers differently for real and fake usernames, so an unauthenticated attacker can enumerate every valid BMC account. The value to an attacker is targeting: sweep the management range, learn which nodes still carry the ODM's default account or the provisioning template's service account, and aim credential-stuffing only at those. It turns a noisy brute-force into a quiet, low-attempt campaign that will not trip lockout thresholds.","attack_vector":"Unauthenticated network access to the BMC web login. Anything that can reach the BMC's HTTP/HTTPS port on the management VLAN.","remediation":"Firmware flash to SPx_12-update-7.00 / SPx_13-update-5.00 or later; low urgency on its own, fold it into whatever BMC flash campaign you are already running. The config-only work carries most of the value and costs nothing: delete vendor default accounts, avoid a fleet-wide shared username in the provisioning template, and enable BMC account lockout plus authentication logging to your SIEM so the enumeration sweep itself becomes visible.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-45925/","https://nvd.nist.gov/vuln/detail/CVE-2021-45925"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-27652","cve":"CVE-2022-27652","aliases":[],"title":"CRI-O: Containers started with non-empty default inheritable capabilities","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Containers started with non-empty default inheritable capabilities","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27652"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-34684","cve":"CVE-2022-34684","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An off-by-one error in nvidia.ko permits data tampering or information disclosure across the ioctl boundary.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An off-by-one error in nvidia.ko permits data tampering or information disclosure across the ioctl boundary. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34684","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-36109","cve":"CVE-2022-36109","aliases":[],"title":"Docker / moby: Supplementary groups not set up properly","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Supplementary groups not set up properly; access to group-readable files inside container","attack_vector":"Any tenant workload","remediation":"Upgrade Docker Engine; restart containers","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-36109"],"status":"curated"},{"id":"CVE-2022-40258","cve":"CVE-2022-40258","aliases":[],"title":"AMI MegaRAC: Weak MD5 password hashing for BMC accounts","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Weak MD5 password hashing for BMC accounts; offline cracking of captured hashes","attack_vector":"Local/offline after hash disclosure","remediation":"BMC firmware update plus credential rotation, since previously-hashed passwords must be considered recoverable","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-40258"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-4122","cve":"CVE-2022-4122","aliases":[],"title":"Buildah: Symlink following when reading .containerignore/.dockerignore discloses host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Buildah","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Symlink following when reading .containerignore/.dockerignore discloses host files","attack_vector":"Malicious build context","remediation":"Upgrade Buildah on build hosts","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-4122"],"status":"curated"},{"id":"CVE-2022-42254","cve":"CVE-2022-42254","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An out-of-bounds array access in nvidia.ko gives an unprivileged local user a crash, a memory leak, or data…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An out-of-bounds array access in nvidia.ko gives an unprivileged local user a crash, a memory leak, or data corruption. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42254","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-125"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42255","cve":"CVE-2022-42255","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A second out-of-bounds array access path in nvidia.ko with the same unprivileged-local reach. Everything with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A second out-of-bounds array access path in nvidia.ko with the same unprivileged-local reach. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42255","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-787"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42256","cve":"CVE-2022-42256","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in index validation defeats the driver's own bounds check, giving an unprivileged user an…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An integer overflow in index validation defeats the driver's own bounds check, giving an unprivileged user an out-of-bounds access in kernel context. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42256","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42257","cve":"CVE-2022-42257","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in nvidia.ko produces information disclosure, data tampering or a node crash. Everything…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An integer overflow in nvidia.ko produces information disclosure, data tampering or a node crash. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42257","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42258","cve":"CVE-2022-42258","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): Another integer overflow path in nvidia.ko reachable by any local GPU user. Everything with a GPU allocation…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Another integer overflow path in nvidia.ko reachable by any local GPU user. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42258","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42265","cve":"CVE-2022-42265","aliases":[],"title":"GPU Display Driver / vGPU guest driver: DoS / info disclosure (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver / vGPU guest driver","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"DoS / info disclosure (integer overflow)","attack_vector":"Any tenant with a container or vGPU guest","remediation":"Driver + vGPU Manager upgrade; rolling node reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42265","https://github.com/NVIDIA/product-security/tree/main/2024/5520"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-42288","cve":"CVE-2022-42288","aliases":[],"title":"DGX servers BMC: Info exposure from BMC","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX servers BMC","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Info exposure from BMC","attack_vector":"Network-adjacent","remediation":"Flash BMC 2.09.00+","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42288","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-208"]},{"id":"CVE-2022-48698","cve":"CVE-2022-48698","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2022","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: fix memory leak when using debugfs_lookup()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-48698","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-20582","cve":"CVE-2023-20582","aliases":[],"title":"AMD IOMMU - nested page table entry faults bypass SEV-SNP RMP checks: MULTI-TENANT ISOLATION: The IOMMU mishandles invalid nested page table entries, letting a privileged attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU - nested page table entry faults bypass SEV-SNP RMP checks","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The IOMMU mishandles invalid nested page table entries, letting a privileged attacker induce PTE faults that bypass SEV-SNP RMP enforcement and tamper with confidential guest memory. Same failure shape as the DTE variant and disclosed alongside it - the IOMMU's error paths are where RMP enforcement leaks.","attack_vector":"Hypervisor-privileged attacker driving DMA through the IOMMU.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Patch alongside the DTE variant; they ship together.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20582","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20584","cve":"CVE-2023-20584","aliases":[],"title":"AMD IOMMU - invalid device table entries bypass SEV-SNP RMP checks: MULTI-TENANT ISOLATION: The IOMMU mishandles certain special address ranges when the device table entry is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU - invalid device table entries bypass SEV-SNP RMP checks","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The IOMMU mishandles certain special address ranges when the device table entry is invalid, and a privileged attacker with a compromised hypervisor can induce those DTE faults to skip the reverse-map table checks that enforce SEV-SNP page ownership. RMP checks are the mechanism that stops a device DMA from landing in a confidential guest's memory; bypassing them via the IOMMU means DMA-based guest tampering.","attack_vector":"Requires hypervisor privilege plus the ability to drive DMA - so a compromised host with a device (or a device it controls) it can point at guest pages.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Particularly relevant on GPU nodes, where large numbers of devices do high-rate DMA and the IOMMU is doing real work rather than sitting idle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20584","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-20585","cve":"CVE-2023-20585","aliases":[],"title":"AMD IOMMU host buffer access - insufficient RMP checks (AMD-SB-3016): MULTI-TENANT ISOLATION: Insufficient RMP checking on IOMMU host buffer access produces an out-of-bounds…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU host buffer access - insufficient RMP checks (AMD-SB-3016)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient RMP checking on IOMMU host buffer access produces an out-of-bounds write. This is the most operationally expensive item in the RMP family, because of what fixing it costs rather than what it does.","attack_vector":"Privileged attacker with hypervisor control, via IOMMU host buffer operations.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. **This one needs more than a BIOS flash**: AMD requires the firmware update, an OS update, *and* a full SNP guest shutdown and platform re-initialization. In practice that means draining every confidential VM off the host, tearing down SNP, updating, re-initializing and re-admitting - a materially longer maintenance window than the rest of the batch, and one you cannot overlap with normal rolling reboots. Plan capacity for it explicitly.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20585","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3016.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-20592","cve":"CVE-2023-20592","aliases":[],"title":"AMD SEV-ES (CacheWarp): CacheWarp: INVD lets a malicious hypervisor revert SEV-ES guest memory writes, breaking guest integrity and…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD SEV-ES (CacheWarp)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"CacheWarp: INVD lets a malicious hypervisor revert SEV-ES guest memory writes, breaking guest integrity and enabling auth bypass inside the VM","attack_vector":"Malicious/compromised host against a tenant confidential VM","remediation":"Microcode/AGESA update + reboot. Undermines the trust story of any SEV-based confidential GPU-VM offering - must be reflected in attestation policy, not just patching","references":["https://access.redhat.com/security/cve/CVE-2023-20592"],"status":"curated","fleet":{"ubiquity":"common - SEV-SNP is the CPU-side TEE that anchors \"confidential GPU\" offerings on EPYC Naples/Rome/Milan hosts","remediation_pain":"microcode+reboot for Milan; Naples/Rome are effectively unpatchable-mitigate-only, so affected nodes must be retired from any confidential-compute SKU","pain_class":"unpatchable / mitigate-only","why_fleet_wide":"Breaks the integrity guarantee of confidential VMs from a malicious hypervisor - the exact threat model a neocloud invokes when it tells a customer their weights are safe from the operator."}},{"id":"CVE-2023-25173","cve":"CVE-2023-25173","aliases":[],"title":"containerd: Supplementary groups not set up correctly inside containers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Supplementary groups not set up correctly inside containers; unexpected access to group-readable files","attack_vector":"Any tenant workload","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25173"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-25192","cve":"CVE-2023-25192","aliases":[],"title":"AMI MegaRAC SPX (Redfish): User enumeration through Redfish — maps valid BMC accounts fleet-wide ahead of a credential attack","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPX (Redfish)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"User enumeration through Redfish — maps valid BMC accounts fleet-wide ahead of a credential attack","attack_vector":"Network / Redfish, unauthenticated","remediation":"BMC firmware update to SPx12-update-7.00 / SPx13-update-5.00; low severity alone, but it is the reconnaissance leg of the MegaRAC chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25192"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25512","cve":"CVE-2023-25512","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: An out-of-bounds read on a malformed input file reaches limited code execution and information disclosure.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An out-of-bounds read on a malformed input file reaches limited code execution and information disclosure. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25512","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-125"]},{"id":"CVE-2023-25513","cve":"CVE-2023-25513","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A second out-of-bounds read path on malformed input with the same limited code-execution reach. The realistic…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A second out-of-bounds read path on malformed input with the same limited code-execution reach. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25513","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-125"]},{"id":"CVE-2023-25514","cve":"CVE-2023-25514","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A third out-of-bounds read path on malformed input, fixed in the same bulletin. The realistic exposure is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A third out-of-bounds read path on malformed input, fixed in the same bulletin. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25514","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-125"]},{"id":"CVE-2023-30633","cve":"CVE-2023-30633","aliases":["INSYDE-SA-2023045"],"title":"Insyde InsydeH2O (TrEEConfigDriver, TPM PCR reporting): Low CVSS, high operational consequence. The driver can report false TPM Platform Configuration Register…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (TrEEConfigDriver, TPM PCR reporting)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Low CVSS, high operational consequence. The driver can report false TPM Platform Configuration Register values, which means the measurements a node presents during remote attestation do not reflect what actually booted. Every control an operator layers on top of measured boot - proving to a customer that their node runs the firmware and image you claim, gating access to model weights or key material on an attestation quote, detecting a bootkit left by the previous tenant - silently returns a pass on a compromised node. This is the bug class that turns every other firmware CVE in this list from detectable into invisible.","attack_vector":"Local attacker on the host able to influence the platform configuration the driver reports, then any subsequent attestation. The attack is against the verifier's trust, not against the node's availability.","remediation":"OEM BIOS update on the fixed Insyde kernel; flash + reboot per node. No config workaround, because the whole point is that the reported state is wrong. Until patched, do not treat PCR-based attestation from affected platforms as authoritative for tenant-isolation or key-release decisions, and cross-check firmware integrity out of band (BMC-side SPI measurement, offline flash comparison) rather than trusting the node's own quote.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-30633","https://www.insyde.com/security-pledge/SA-2023045"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-32280","cve":"CVE-2023-32280","aliases":["INTEL-SA-00922"],"title":"Intel Server OpenBMC firmware (before egs-1.05) - credential storage: Credentials are insufficiently protected and an unauthenticated caller on the network can read them. Scored…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server OpenBMC firmware (before egs-1.05) - credential storage","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Credentials are insufficiently protected and an unauthenticated caller on the network can read them. Scored only 5.3 because what leaks is partial, but credential material recovered without authentication is worth more than the score suggests on a fleet that reuses BMC credentials across nodes - the standard practice for anyone who provisioned racks from a template. Read once, reuse everywhere.","attack_vector":"Unauthenticated, over the network, to the BMC management interface.","remediation":"Fixed in Intel Server OpenBMC egs-1.05 and later - per-node out-of-band BMC firmware update. The compensating control is per-node unique BMC credentials, which is config-only, costs a provisioning change, and blunts every credential-disclosure bug in this cluster rather than just this one. If your fleet currently shares one BMC password, fixing that is higher leverage than this specific patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32280","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00922.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-34344","cve":"CVE-2023-34344","aliases":["AMI-SA-2023005","NVIDIA OSR review"],"title":"AMI MegaRAC SPx (IPMI handler): Timing and response differences in the IPMI handler let an unauthenticated attacker confirm which usernames…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (IPMI handler)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Timing and response differences in the IPMI handler let an unauthenticated attacker confirm which usernames exist on the BMC. On its own it leaks nothing but names; in a fleet it is the reconnaissance step that makes credential-stuffing efficient, because it tells the attacker which nodes still carry the vendor default account or the provisioning template's service account before they spend any attempts.","attack_vector":"Network-reachable IPMI service, no credentials, no interaction. Any host that can reach UDP/623 on the BMC can enumerate accounts across the whole management range in a single sweep.","remediation":"Firmware flash to SPx_12.7 / SPx_13.5, out-of-band per node, ODM-gated. Genuinely low urgency for the flash itself. The config-only work is what matters: remove or rename vendor default accounts, ensure no username is shared across the fleet by the provisioning template, and disable IPMI-over-LAN where your tooling allows it - that closes the enumeration surface outright with no reboot.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2023005.pdf","https://nvd.nist.gov/vuln/detail/CVE-2023-34344"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-36844","cve":"CVE-2023-36844","aliases":[],"title":"Juniper Junos OS J-Web (EX): **[KEV]** PHP external variable modification","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS J-Web (EX)","year":"2023","cvss_score":5.3,"severity":"medium","kev":true,"impact":"**[KEV]** PHP external variable modification; part of the exploited J-Web chain","attack_vector":"Network, unauthenticated","remediation":"Junos upgrade with switch failover","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36844"],"status":"curated"},{"id":"CVE-2023-36847","cve":"CVE-2023-36847","aliases":[],"title":"Juniper Junos OS J-Web (EX): **[KEV]** Missing authentication on `installAppPackage.php` — unauthenticated file upload to the switch…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS J-Web (EX)","year":"2023","cvss_score":5.3,"severity":"medium","kev":true,"impact":"**[KEV]** Missing authentication on `installAppPackage.php` — unauthenticated file upload to the switch filesystem","attack_vector":"Network, unauthenticated","remediation":"Junos upgrade; combined with CVE-2023-36845 this is a pre-auth RCE chain against management switches","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-36847"],"status":"curated"},{"id":"CVE-2023-38958","cve":"CVE-2023-38958","aliases":["CVE-2023-38954","CVE-2023-38955","CVE-2023-38956"],"title":"ZKTeco BioAccess IVS v3.3.1 access control platform: An unauthenticated attacker can open and close any door the platform manages by sending a crafted web…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ZKTeco BioAccess IVS v3.3.1 access control platform","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An unauthenticated attacker can open and close any door the platform manages by sending a crafted web request. The CVSS score badly understates this: a 5.3 labelled 'access control issue' is, in the physical world, a remote door-open primitive with no credentials. The same disclosure set adds unauthenticated device enumeration (IP addresses and names of every reader and controller), unauthenticated path traversal for arbitrary file read, and SQL injection - so the attacker can map the site, find the door serving the GPU cage specifically, and then open it. ZKTeco gear is common in cost-sensitive and fast-built deployments, which describes a lot of newer GPU hosting sites, and it is often installed by a general contractor rather than a security integrator. What follows from an open cage door is the usual list: drives with model weights and customer data walk out, a console or USB device gets attached to a running node, or someone plugs into the out-of-band switch and reaches every BMC in the row.","attack_vector":"Unauthenticated HTTP request to the BioAccess IVS platform. It is a web-managed server; where it is reachable from the corporate network or, worse, published for remote administration, the attack is a single request from anywhere. Device enumeration first means the attacker does not need prior knowledge of your site layout.","remediation":"Upgrade past 3.3.1 to a fixed BioAccess IVS release - ZKTeco's patch cadence and advisory quality are weak, so verify with the vendor that the specific door-control endpoint is fixed rather than trusting a version bump. Given the vendor track record, the stronger recommendation for a datacenter is to treat this product as unsuitable for a hall or cage boundary and plan replacement with a platform that has a real security program. Meanwhile: take the platform entirely off any network reachable from outside the security VLAN, put it behind a jump host, and add a compensating physical control on the doors that actually matter - a mechanical lock, a mantrap with a second independent system, or a guard - because a remote unlock primitive with no authentication is not something a network ACL fully mitigates if the ACL is ever wrong.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-38958","https://claroty.com/team82/disclosure-dashboard/cve-2023-38958","https://nvd.nist.gov/vuln/detail/CVE-2023-38956"],"status":"curated"},{"id":"CVE-2023-4155","cve":"CVE-2023-4155","aliases":[],"title":"Linux KVM - SEV-ES/SEV-SNP VMGEXIT double-fetch race: A KVM guest running SEV-ES or SEV-SNP with several vCPUs can trigger a double-fetch race in the host's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM - SEV-ES/SEV-SNP VMGEXIT double-fetch race","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A KVM guest running SEV-ES or SEV-SNP with several vCPUs can trigger a double-fetch race in the host's VMGEXIT handler and drive it into recursion. Repeated invocation exhausts the host kernel stack and panics the hypervisor - a confidential guest taking down the host it runs on, and with it every other tenant on that machine. This is guest-to-host denial of service, which is the direction operators care about.","attack_vector":"From inside a guest VM using SEV-ES or SEV-SNP with multiple vCPUs. Tenant-reachable - no host privilege needed.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Prioritise on any host that admits tenant-controlled confidential VMs; the attacker prerequisite is just 'has a VM here'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4155","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-4693","cve":"CVE-2023-4693","aliases":[],"title":"GRUB2 (NTFS filesystem parser): Out-of-bounds read in the same NTFS path leaks GRUB heap memory. On its own it is an information leak, but it…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (NTFS filesystem parser)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Out-of-bounds read in the same NTFS path leaks GRUB heap memory. On its own it is an information leak, but it is the ASLR-defeating half that makes the paired write bug reliably exploitable.","attack_vector":"Attacker-supplied NTFS volume, physical or via BMC virtual media.","remediation":"grub2 package update + reboot. Same as its sibling - dropping the NTFS module from your build is the durable answer on a Linux-only fleet.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-4693","https://access.redhat.com/security/cve/CVE-2023-4693"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-48299","cve":"CVE-2023-48299","aliases":[],"title":"TorchServe (model/workflow API): Information disclosure of files on the serving host","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TorchServe (model/workflow API)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Information disclosure of files on the serving host","attack_vector":"Network to the management API","remediation":"Patch to 0.9.0+; firewall the management plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-48299"],"status":"curated"},{"id":"CVE-2023-52738","cve":"CVE-2023-52738","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/fence): A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/fence)","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu RAS / GPU reset and recovery path. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/fence: Fix oops due to non-matching drm_sched init/fini","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52738","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-6032","cve":"CVE-2023-6032","aliases":["SEVD-2023-318-03"],"title":"Schneider Electric Galaxy VS / VL / VXL three-phase UPS, Network Management Card over HTTPS: Path traversal lets an attacker enumerate and download files from the management card on a large…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Schneider Electric Galaxy VS / VL / VXL three-phase UPS, Network Management Card over HTTPS","year":"2023","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Path traversal lets an attacker enumerate and download files from the management card on a large three-phase UPS - the class of unit that sits between utility power and an entire GPU hall, not a single rack. What leaks is configuration and credential material for the power estate. Treat this as reconnaissance that precedes a physical availability attack rather than as a data-loss event in itself.","attack_vector":"Anyone who can reach the NMC's HTTPS interface. Galaxy-class UPS management cards live on the facility network, which in a leased colo is usually the landlord's network, not yours.","remediation":"Firmware update to the card, per SEVD-2023-318-03. Non-disruptive to the load, but on a leased site you may not own the equipment - in that case the real remediation is contractual: require the landlord to evidence the patch level of every UPS management card that feeds your halls, and require the facility network to be segmented from anything you run.","references":["https://download.schneider-electric.com/files?p_Doc_Ref=SEVD-2023-318-03&p_enDocType=Security+and+Safety+Notice&p_File_Name=SEVD-2023-318-03.pdf"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21944","cve":"CVE-2024-21944","aliases":[],"title":"AMD SEV-SNP (BadRAM): BadRAM: improper validation of DIMM SPD metadata lets an attacker with physical access or ring0 on a…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD SEV-SNP (BadRAM)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"BadRAM: improper validation of DIMM SPD metadata lets an attacker with physical access or ring0 on a non-compliant DIMM overwrite guest memory and forge SNP attestation","attack_vector":"Physical access / compromised host firmware against a confidential tenant VM","remediation":"AGESA/BIOS firmware update + reboot, plus DIMM SPD lockdown at the supply-chain level. Cannot be fixed in software - a genuine constraint on any \"we cannot see your data\" confidential-GPU claim","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21944"],"status":"curated","fleet":{"ubiquity":"common - 3rd/4th-gen EPYC (Milan, Milan-X, Genoa, Bergamo, Genoa-X, Siena) hosts under GPU nodes","remediation_pain":"firmware-flash - AMD's fix validates SPD metadata at boot, so it is a BIOS/AGESA update per node; the underlying attack needs ~$10 of hardware and physical access, which colo and bare-metal-rental models do not exclude","pain_class":"physical access","why_fleet_wide":"Forges SEV-SNP attestation reports and inserts undetectable backdoors into confidential VMs, i.e. every attestation a customer verified on affected hosts is retroactively meaningless."}},{"id":"CVE-2024-23650","cve":"CVE-2024-23650","aliases":[],"title":"BuildKit: Malicious client or frontend crashes the BuildKit daemon","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Malicious client or frontend crashes the BuildKit daemon","attack_vector":"Anyone with build API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23650"],"status":"curated"},{"id":"CVE-2024-27891","cve":"CVE-2024-27891","aliases":["Arista Security Advisory 0102"],"title":"Arista EOS (MACsec with egress ACLs): TENANT ISOLATION: on interfaces with both MACsec and egress ACLs configured, the egress ACL is not enforced…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (MACsec with egress ACLs)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: on interfaces with both MACsec and egress ACLs configured, the egress ACL is not enforced for packets leaving those ports. The combination — link encryption plus egress filtering — is exactly what you deploy on inter-site or inter-pod links carrying multiple tenants, so the failure lands on the highest-trust links in the build.","attack_vector":"Traffic egressing an interface configured with both MACsec and an egress ACL. No attacker capability needed.","remediation":"EOS upgrade plus reload. Interim: move the filtering to the ingress direction on the far side of the link, which is a live config change and restores enforcement without touching MACsec.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27891","https://www.arista.com/en/support/advisories-notices/security-advisory/19908-security-advisory-0102"],"status":"curated"},{"id":"CVE-2024-31157","cve":"CVE-2024-31157","aliases":[],"title":"Intel UEFI firmware (OutOfBandXML module): Improper initialisation in the OutOfBandXML UEFI module allows a privileged user to disclose information from…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel UEFI firmware (OutOfBandXML module)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Improper initialisation in the OutOfBandXML UEFI module allows a privileged user to disclose information from firmware. The out-of-band XML path is part of remote platform configuration, so it is reachable in the management workflows operators actually automate.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in platform BIOS/UEFI firmware. That means an OEM release, a per-node drain, a flash and a cold reboot - and OEM availability commonly lags the Intel advisory by quarters on server boards. There is no microcode or OS-level shortcut for this class; budget it as a fleet-wide maintenance campaign, not a patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-31157","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01139.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-37152","cve":"CVE-2024-37152","aliases":[],"title":"Argo CD: /api/v1/settings exposes sensitive settings without authentication","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"/api/v1/settings exposes sensitive settings without authentication","attack_vector":"Unauthenticated network","remediation":"Rolling Argo CD upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-37152"],"status":"curated"},{"id":"CVE-2024-39707","cve":"CVE-2024-39707","aliases":["INSYDE-SA-2024007"],"title":"Insyde InsydeH2O (IHISI function 0x49, UEFI variable factory reset): IHISI function 0x49 restores certain UEFI variables to factory defaults with no authentication by default.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Insyde InsydeH2O (IHISI function 0x49, UEFI variable factory reset)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"IHISI function 0x49 restores certain UEFI variables to factory defaults with no authentication by default. That is a rollback primitive: an attacker resets security-relevant firmware settings that an operator had hardened - and on affected platforms can wind protections back to a weaker known state without needing to exploit anything. For a fleet where node hardening is applied once at provisioning and then assumed, this quietly undoes it, and nothing in the OS logs the change.","attack_vector":"Local attacker on the host OS able to invoke the IHISI interface. No authentication required on affected platforms.","remediation":"OEM BIOS update carrying the Insyde fix, which makes the function require authentication. Firmware flash, reboot per node. Because the impact is settings rollback rather than code execution, the practical operator control is detection: baseline your BIOS settings at provisioning and re-verify them (via the OEM's redfish/BIOS-attribute API) on every node before it re-enters the tenant pool, rather than assuming a hardened node stays hardened.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-39707","https://www.insyde.com/security-pledge/SA-2024007"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-42478","cve":"CVE-2024-42478","aliases":[],"title":"llama.cpp (RPC backend): Arbitrary address read via `rpc_tensor.data`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"llama.cpp (RPC backend)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Arbitrary address read via `rpc_tensor.data`","attack_vector":"Unauthenticated network to the RPC port","remediation":"Rebuild; network-isolate","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42478"],"status":"curated"},{"id":"CVE-2024-43101","cve":"CVE-2024-43101","aliases":[],"title":"Intel Data Center GPU Flex Series - Windows driver software: Improper access control in the Flex Series Windows driver software allows an authenticated user to cause…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Data Center GPU Flex Series - Windows driver software","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Improper access control in the Flex Series Windows driver software allows an authenticated user to cause denial of service.","attack_vector":"Local, authenticated, on a Windows host running Flex Series driver software before 31.0.101.4255.","remediation":"Update to 31.0.101.4255 or later. Cost: node reboot after driver replacement.","references":["https://www.intel.com/content/www/us/en/security-center/default.html","https://nvd.nist.gov/vuln/detail/CVE-2024-43101"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-45783","cve":"CVE-2024-45783","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (HFS+ filesystem parser): A reference count can be decremented twice, producing a use-after-free. Lower severity on its own but a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (HFS+ filesystem parser)","year":"2024","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A reference count can be decremented twice, producing a use-after-free. Lower severity on its own but a usable link in a chain with the write primitives in the same batch.","attack_vector":"Attacker-supplied HFS+ volume.","remediation":"grub2 package update + reboot; or drop the module.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45783","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23028","cve":"CVE-2025-23028","aliases":[],"title":"Cilium: Denial of service in the Cilium dataplane","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Denial of service in the Cilium dataplane","attack_vector":"Any pod on the cluster network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23028"],"status":"curated"},{"id":"CVE-2025-29934","cve":"CVE-2025-29934","aliases":[],"title":"AMD CPU - stale TLB entries in SEV-SNP guests: MULTI-TENANT ISOLATION: A silicon bug lets a local admin-privileged attacker run an SEV-SNP guest against…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU - stale TLB entries in SEV-SNP guests","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A silicon bug lets a local admin-privileged attacker run an SEV-SNP guest against stale TLB entries. The guest executes with translations that no longer reflect the real page mappings, which the host can steer - a data-integrity attack on a confidential VM that leaves no trace in the guest, since from inside the VM the memory simply reads wrong.","attack_vector":"Local, admin-privileged host attacker targeting a confidential guest.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29934","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33185","cve":"CVE-2025-33185","aliases":[],"title":"NVIDIA AIStore - AuthN: An unauthenticated user extracts information from the AIStore authentication component. AIStore fronts…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA AIStore - AuthN","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"An unauthenticated user extracts information from the AIStore authentication component. AIStore fronts training datasets, so what leaks here is metadata about your data estate.","attack_vector":"Network, unauthenticated, no user interaction. Anyone who can reach the AuthN endpoint.","remediation":"Upgrade AIStore per bulletin 5724 and roll the AuthN pods. Cost: rolling restart of the storage control plane; data path is unaffected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33185","https://github.com/NVIDIA/product-security/tree/main/2025/5724"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-862"],"fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2025-5197","cve":"CVE-2025-5197","aliases":[],"title":"HuggingFace transformers: ReDoS in `convert_tf_weight_name_to_pt_weight_name`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"HuggingFace transformers","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"ReDoS in `convert_tf_weight_name_to_pt_weight_name`","attack_vector":"Customer-supplied TF checkpoint converted at load","remediation":"Upgrade transformers","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-5197"],"status":"curated"},{"id":"CVE-2025-52534","cve":"CVE-2025-52534","aliases":[],"title":"AMD CPU microcode - bound check: MULTI-TENANT ISOLATION: An improper bound check inside AMD CPU microcode lets a malicious **guest** write…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU microcode - bound check","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An improper bound check inside AMD CPU microcode lets a malicious **guest** write into host memory. This is a straight VM-escape primitive expressed in silicon microcode rather than in the hypervisor: the tenant does not need a QEMU or KVM bug, the CPU itself lets the write through. Once a guest can write host memory the host is compromised, and from there every other VM on the box.","attack_vector":"From inside a guest VM - no host privilege required. That makes this materially more dangerous operationally than the local-admin microcode issues: your tenants are the attackers in the threat model.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly. Prioritise this above the local-admin microcode issues on any node that runs untrusted guest VMs - the attacker prerequisite here is simply 'is a tenant'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52534","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-54511","cve":"CVE-2025-54511","aliases":[],"title":"AMD Secure Processor - privilege check on write path: The ASP accepts an input value and performs a write without confirming the caller had sufficient privilege to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - privilege check on write path","year":"2025","cvss_score":5.3,"severity":"medium","kev":false,"impact":"The ASP accepts an input value and performs a write without confirming the caller had sufficient privilege to ask for it. The result is an integrity loss inside the secure processor - a caller that should have been refused gets its write.","attack_vector":"Local, through the ASP's callable interface.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54511","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24208","cve":"CVE-2026-24208","aliases":[],"title":"Triton Inference Server: Info disclosure via path traversal on model files","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Info disclosure via path traversal on model files","attack_vector":"Network client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24208","https://github.com/NVIDIA/product-security/tree/main/2026/5828"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-22"]},{"id":"CVE-2026-24227","cve":"CVE-2026-24227","aliases":[],"title":"TensorRT: DoS via resource exhaustion","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"TensorRT","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"DoS via resource exhaustion","attack_vector":"Malicious model/engine input","remediation":"Bump TensorRT; rebuild serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24227","https://github.com/NVIDIA/product-security/tree/main/2026/5855"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-502"]},{"id":"CVE-2026-47262","cve":"CVE-2026-47262","aliases":[],"title":"containerd: Crafted image causes DoS during container creation","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Crafted image causes DoS during container creation","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47262"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-47622","cve":"CVE-2026-47622","aliases":[],"title":"NVIDIA Dynamo: Error messages from Dynamo leak sensitive information to an unauthenticated network caller - typically…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Dynamo","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Error messages from Dynamo leak sensitive information to an unauthenticated network caller - typically internal paths, versions and configuration that shorten the next stage of an attack.","attack_vector":"Network, unauthenticated. Anyone who can reach the Dynamo serving endpoint and provoke an error.","remediation":"Upgrade Dynamo per bulletin 5842 and roll the deployment. Cost: rolling restart of the serving tier.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47622","https://github.com/NVIDIA/product-security/tree/main/2026/5842"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-209"],"fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2026-55686","cve":"CVE-2026-55686","aliases":[],"title":"Podman: Malicious image WORKDIR symlink creates directories or changes ownership on the host filesystem","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Podman","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Malicious image WORKDIR symlink creates directories or changes ownership on the host filesystem","attack_vector":"Malicious image","remediation":"Upgrade Podman to 5.7.1+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-55686"],"status":"curated"},{"id":"NCVD-2020-002-rdma-fabric-remote-dram-bank-con","cve":null,"aliases":["Bankrupt","RDMA memory-bank covert channel","Ustiugov et al., arXiv:2006.03854"],"title":"RDMA fabric + remote DRAM bank contention (cross-node covert channel): TENANT ISOLATION: Bankrupt establishes a 74 Kb/s covert channel between two processes on different machines…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RDMA fabric + remote DRAM bank contention (cross-node covert channel)","year":"2020","cvss_score":5.3,"severity":"medium","kev":false,"impact":"TENANT ISOLATION: Bankrupt establishes a 74 Kb/s covert channel between two processes on different machines in an RDMA network, by steering RDMA packets to addresses that map to a single DRAM bank on a shared intermediary node and timing the resulting queuing. It remained undetectable to existing monitoring - CPU and NIC performance counters showed nothing. For an operator, this defeats the assumption that network segmentation between tenants prevents exfiltration: a compromised process inside an isolated enclave can signal out to a colluding receiver anywhere on the same RDMA fabric, without opening a connection between them. Any data-loss-prevention story that relies on egress controls at the IP layer is bypassed.","attack_vector":"Spy and receiver each allocate their own private memory region on a common intermediary machine - a normal thing for any RDMA tenant to do. The spy issues RDMA operations to a chosen set of remote addresses, causing deep queuing at one memory bank; the receiver probes addresses mapped to the same bank in its own region and reads the timing. Both sides only ever touch memory they legitimately own, which is why nothing flags it.","remediation":"No patch. Mitigation is placement and monitoring: avoid a shared intermediary that both a sensitive tenant and an untrusted tenant can target (scheduler policy change), and where memory-bank interleaving is configurable, randomise the physical-address-to-bank mapping per tenant (BIOS/firmware setting, requires a host reboot). Detection needs per-QP latency-distribution telemetry rather than counters - a monitoring build-out, not a config toggle. For genuinely sensitive workloads, a dedicated fabric is the only reliable answer.","references":["https://arxiv.org/abs/2006.03854"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-006-dmtf-libspdm-csr-generation-unde","cve":null,"aliases":["GHSA-j54w-759w-xj3m","DMTF-2026-0002"],"title":"DMTF libspdm CSR generation under the mbedTLS crypto backend (cryptlib_mbedtls): Stack corruption inside the firmware component that generates certificate signing requests - meaning the code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DMTF libspdm CSR generation under the mbedTLS crypto backend (cryptlib_mbedtls)","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"Stack corruption inside the firmware component that generates certificate signing requests - meaning the code path that establishes a device's cryptographic identity is the one that can be smashed. At minimum this is a denial of service against device provisioning; at worst, stack corruption in a firmware context of this privilege is a candidate for control-flow hijack. On a GPU fleet the affected components are the accelerators, NICs and management controllers that participate in device identity provisioning, and a device whose identity provisioning can be attacked is a device whose attestation claims an operator should not rely on. An oversized Common Name in the RequesterInfo field of a GET_CSR request writes past the end of a stack array. Reachable only where the responder advertises CSR_CAP and uses libspdm_gen_x509_csr().","attack_vector":"An SPDM requester able to send GET_CSR to a responder that has CSR_CAP enabled and uses the mbedTLS backend. That is a narrow configuration, but it is the configuration used by embedded firmware that cannot carry OpenSSL - which describes a lot of BMC and device firmware.","remediation":"Update libspdm and rebuild affected firmware; there is no CVE, so this will not surface through NVD-based scanning and you have to track the DMTF advisory series directly. The narrowing conditions are useful operationally: ask vendors whether their SPDM responder is built with CSR_CAP and against mbedTLS, and if CSR generation is not a capability you use, having it compiled out removes the path entirely. As with everything in libspdm, the actual rollout is per-vendor firmware images and per-device flashes.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-j54w-759w-xj3m"],"status":"curated"},{"id":"NCVD-2026-007-dmtf-libspdm-responder-handling","cve":null,"aliases":["GHSA-m4wc-xmvg-369f","DMTF-2026-0001"],"title":"DMTF libspdm responder handling of GET_MEASUREMENT_EXTENSION_LOG: A requester reads memory it was never authorised to read, out of the responder - which in a datacenter is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DMTF libspdm responder handling of GET_MEASUREMENT_EXTENSION_LOG","year":"2026","cvss_score":5.3,"severity":"medium","kev":false,"impact":"A requester reads memory it was never authorised to read, out of the responder - which in a datacenter is typically the firmware of a GPU, NIC, or the BMC acting as an SPDM responder. What leaks is whatever sits adjacent to the measurement extension log in that component's address space: potentially keys, measurement state, or other firmware data. Because the responder is a device that multiple hosts or tenants may talk to over the life of a node, this is a cross-boundary read from a component that is supposed to be the root of the attestation story rather than a target of it. The Offset and Length fields are added with wrapping arithmetic before the bounds check, so a crafted pair overflows and passes validation. Requires the responder to advertise MEL_CAP and CHUNK_CAP.","attack_vector":"Anything that can act as an SPDM requester to the affected responder - a host driver, a management controller, or a peer device on the fabric. On bare metal that includes the tenant's own host software if the platform lets host drivers speak SPDM to accelerators, which most do.","remediation":"This one has no CVE assigned, only a GHSA and a DMTF advisory number, so it will not appear in NVD-driven scanning at all - operators tracking firmware risk off CVE feeds alone will simply never see it. Remediation is a libspdm update followed by per-vendor firmware rebuilds and per-device flashes, on each vendor's own timeline. If your responders do not need the measurement extension log, having the vendor build with MEL_CAP disabled removes the reachable path, but that is a firmware build option rather than something an operator can toggle.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-m4wc-xmvg-369f"],"status":"curated"},{"id":"CVE-2020-15257","cve":"CVE-2020-15257","aliases":[],"title":"containerd: containerd-shim abstract-socket API exposed to host-network containers","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2020","cvss_score":5.2,"severity":"medium","kev":false,"impact":"containerd-shim abstract-socket API exposed to host-network containers; container escape to host root","attack_vector":"Any tenant pod running with hostNetwork:true","remediation":"Upgrade containerd; drain node to restart shims. Also ban hostNetwork for tenant pods via policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15257"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-46746","cve":"CVE-2021-46746","aliases":[],"title":"AMD Secure Processor TEE - Secure OS stack overrun (AMD-SB-3003): A stack overrun in the ASP Secure OS trusted execution environment, denying service to the secure processor.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor TEE - Secure OS stack overrun (AMD-SB-3003)","year":"2021","cvss_score":5.2,"severity":"medium","kev":false,"impact":"A stack overrun in the ASP Secure OS trusted execution environment, denying service to the secure processor. With the ASP down, the platform loses fTPM services, SEV key operations and attestation - so on a confidential-computing host this is not a cosmetic crash, it takes your CVM capacity offline until the node is power-cycled.","attack_vector":"Local, through the ASP TEE interface.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46746","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-31011","cve":"CVE-2023-31011","aliases":[],"title":"DGX H100 BMC (REST): DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX H100 BMC (REST)","year":"2023","cvss_score":5.2,"severity":"medium","kev":false,"impact":"DoS","attack_vector":"Network-adjacent BMC REST client","remediation":"Flash BMC 23.08.18 out-of-band","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5473/5473.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:H/UI:N/S:U/C:L/I:H/A:N","cwe":["CWE-20"]},{"id":"CVE-2023-31189","cve":"CVE-2023-31189","aliases":["INTEL-SA-00922"],"title":"Intel Server OpenBMC firmware (before egs-1.09) - authentication logic: An authenticated low-privilege user escalates privilege on the BMC, with a scope change - the escalation…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Server OpenBMC firmware (before egs-1.09) - authentication logic","year":"2023","cvss_score":5.2,"severity":"medium","kev":false,"impact":"An authenticated low-privilege user escalates privilege on the BMC, with a scope change - the escalation crosses a security boundary rather than staying inside the account model. On fleets that hand out constrained BMC accounts to tenants, support staff or monitoring systems (read-only sensor scraping is a common one), this converts any of those accounts into something with more control over the node. The lesson for a GPU operator: a read-only BMC account issued to a customer or a monitoring vendor is not a safe thing to hand out on this firmware.","attack_vector":"Local access with a low-privileged authenticated BMC account. Needs an existing credential, so exposure tracks how widely you distribute BMC accounts.","remediation":"Fixed in Intel Server OpenBMC egs-1.09 and later; delivery is a per-node out-of-band BMC firmware update through Intel's platform packages, with the usual OEM rebase lag on non-Intel-badged boards using the same base. Config-only compensation: stop issuing BMC accounts to third parties, proxy sensor and telemetry reads through your own collector rather than giving monitoring vendors direct BMC credentials.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31189","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00922.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-20397","cve":"CVE-2024-20397","aliases":[],"title":"Cisco NX-OS (bootloader / image signature verification): Secure boot on the switch is defeatable: an attacker with physical access or admin credentials can make a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Cisco NX-OS (bootloader / image signature verification)","year":"2024","cvss_score":5.2,"severity":"medium","kev":false,"impact":"Secure boot on the switch is defeatable: an attacker with physical access or admin credentials can make a Nexus load an unsigned NX-OS image. That is how a fabric compromise becomes permanent — a modified image keeps root across reloads and reflashes, and nothing in your config management will notice. Affects Nexus 3000/7000/9000, MDS 9000 and UCS 6400/6500 fabric interconnects, so it covers both the Ethernet and the storage fabric.","attack_vector":"Physical access to the switch (console/bootloader prompt) or an existing administrative account. Realistic threat model for colocation, shared cages, and any switch that has passed through a supply chain or an RMA.","remediation":"BIOS update on both the primary and the alternate BIOS bank — either through `install all` with a fixed NX-OS release or Cisco's release-independent BIOS upgrade script. This is a firmware flash, needs a reload, and must be applied per-device; you cannot fix it with config. Pair it with physical access control on the console ports.","references":["https://sec.cloudapps.cisco.com/security/center/content/CiscoSecurityAdvisory/cisco-sa-nxos-image-sig-bypas-pQDRQvjL","https://nvd.nist.gov/vuln/detail/CVE-2024-20397"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-33660","cve":"CVE-2024-33660","aliases":["AMI-SA-2024004"],"title":"AMI AptioV UEFI BIOS (SPI flash integrity verification): An actor with physical access can modify the SPI flash without the modification being detected. The score is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI AptioV UEFI BIOS (SPI flash integrity verification)","year":"2024","cvss_score":5.2,"severity":"medium","kev":false,"impact":"An actor with physical access can modify the SPI flash without the modification being detected. The score is modest because it needs hands on the hardware, but the operator consequence is that your firmware integrity story has no floor: a node that passed through an untrusted physical environment can carry an undetectable implant. That matters concretely for GPU fleets in leased colo where remote hands are third-party staff, for hardware shipped internationally, for anything bought on the secondary market during a supply crunch, and for RMA units returning from a vendor depot.","attack_vector":"Physical access to the machine, no credentials needed. Anyone who can open the chassis and reach the SPI flash - datacenter remote-hands staff, shipping and logistics handling, a vendor's repair depot, or a hosting provider's own technicians.","remediation":"BIOS update to BKC_5.37 or later - firmware flash plus reboot per node, vendor-rebase-gated. Because the threat model is physical rather than network, patching is only part of it: the process controls are what actually help. Take a firmware measurement baseline per node at commissioning, re-measure after any physical service event or RMA return, use chassis intrusion detection and seal logging, and treat any node that came back from third-party hands without a verified measurement as needing a reflash before it rejoins the pool.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024004.pdf","https://nvd.nist.gov/vuln/detail/CVE-2024-33660"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-47389","cve":"CVE-2021-47389","aliases":[],"title":"Linux KVM/SVM - missing sev_decommission in sev_receive_start: KVM failed to DECOMMISSION the current SEV context when binding an ASID fails after RECEIVE_START. The…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux KVM/SVM - missing sev_decommission in sev_receive_start","year":"2021","cvss_score":5.1,"severity":"medium","kev":false,"impact":"KVM failed to DECOMMISSION the current SEV context when binding an ASID fails after RECEIVE_START. The firmware-side SEV context is left allocated with nothing owning it, exhausting the limited pool of SEV contexts the platform supports. Repeat the failure enough times and the host can no longer launch confidential VMs at all - a resource-exhaustion denial of service against your confidential-computing capacity, reachable through the guest-import path.","attack_vector":"Through the SEV guest receive/import path - so a tenant or control-plane action that fails repeatedly, deliberately or otherwise.","remediation":"Fixed in the Linux kernel. Distro kernel update plus a host reboot. Note that recovering exhausted SEV contexts on an unpatched host generally means an SNP platform shutdown/init cycle, which requires draining every confidential guest anyway.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-47389"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-3172","cve":"CVE-2022-3172","aliases":[],"title":"Kubernetes (kube-apiserver): Aggregated API server can redirect apiserver clients","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2022","cvss_score":5.1,"severity":"medium","kev":false,"impact":"Aggregated API server can redirect apiserver clients; SSRF from the control plane","attack_vector":"Whoever controls an aggregated APIService backend","remediation":"Rolling control-plane upgrade; audit registered APIServices","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2023-40551","cve":"CVE-2023-40551","aliases":["shim 15.8 batch"],"title":"shim (MZ/PE header parser): Out-of-bounds read parsing MZ binaries. Leaks memory contents from the pre-boot environment, which can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"shim (MZ/PE header parser)","year":"2023","cvss_score":5.1,"severity":"medium","kev":false,"impact":"Out-of-bounds read parsing MZ binaries. Leaks memory contents from the pre-boot environment, which can include key material the firmware has not yet cleared.","attack_vector":"Crafted MZ binary loaded via shim.","remediation":"shim package update + reboot. Bundle with the rest of the shim 15.8 batch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40551","https://www.openwall.com/lists/oss-security/2024/01/26/1"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-47973","cve":"CVE-2024-47973","aliases":["Solidigm SA-000563","over-provisioning data disclosure"],"title":"Solidigm DC SSDs (D3-S4510/S4520/S4610/S4620, D5-P5316, D7-P5520/P5620, DC S4500/S4600) - over-provisioned NAND not covered by sanitization: A defect in how the drive handles its over-provisioned…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSDs (D3-S4510/S4520/S4610/S4620, D5-P5316, D7-P5520/P5620, DC S4500/S4600) - over-provisioned NAND not…","year":"2024","cvss_score":5.1,"severity":"medium","kev":false,"impact":"A defect in how the drive handles its over-provisioned capacity leaks data to an attacker. Over-provisioned blocks are the spare NAND the FTL keeps outside the host-visible LBA range - they hold real tenant data that was written and then remapped, and they are invisible to every host-side wipe. BREAKS TENANT HANDOFF: dd-over-the-whole-device, blkdiscard, mkfs and any 'overwrite the visible address space' routine cannot reach these blocks by construction, so a previous tenant's data survives a reclaim that looks completely thorough from the host. This is the concrete, CVE'd instance of the general wear-levelling problem operators are usually told about only in the abstract.","attack_vector":"A tenant with local/root access on the bare-metal host after reclaim, reading back data the previous tenant wrote. Low privilege required on the host; no physical access needed.","remediation":"Firmware flash, per SKU, drive offline: ACV10340 for D5-P5316, XCV10151/XC311151 for D3-S4510/S4610, 7CV10111 for D3-S4520/S4620, 9CV10410 for D7-P5520/P5620, YCV10200 for D5-P5530, applied with Solidigm Storage Tool. Note the sting in the advisory: for DC S4500 and DC S4600 Solidigm states it has NO plans to ship a fix unless a customer explicitly requests one - on those SKUs this is effectively UNPATCHABLE, and the only safe reclaim is physical destruction or accepting that host-side wipes do not cover OP blocks. Because no host-side tool can reach over-provisioned NAND, the durable control is again encryption you own: LUKS/dm-crypt per tenant with the key in your KMS, so remapped ciphertext in OP blocks is worthless after you delete the key. Flashing a 10,000-drive fleet is a rolling drain across weeks with per-SKU tooling and per-SKU firmware images; expect the inventory step (working out which of five Solidigm families each node actually has) to take as long as the flashing.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47973","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-7881","cve":"CVE-2024-7881","aliases":["TFV-13","data memory-dependent prefetch leak"],"title":"Arm Neoverse V2 / V3 / V3AE, Cortex-X3 / X4 / X925, C1-series; mitigated in Trusted Firmware-A v2.2-v2.12 and LTS 2.8/2.10: Unprivileged code can steer a data-memory-dependent prefetcher into…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Neoverse V2 / V3 / V3AE, Cortex-X3 / X4 / X925, C1-series; mitigated in Trusted Firmware-A v2.2-v2.12 and LTS…","year":"2024","cvss_score":5.1,"severity":"medium","kev":false,"impact":"Unprivileged code can steer a data-memory-dependent prefetcher into loading a privileged address and then dereferencing its contents, turning the prefetcher into a read oracle for kernel, hypervisor or secure-world memory. On a shared training or inference node this is a slow but real cross-boundary leak - the kind of thing that gets you key material and pointers for a follow-on exploit rather than an instant escape. Neoverse V2 is the Grace core, so GH200 and GB200 head nodes are in scope.","attack_vector":"Any unprivileged process on the host, or unprivileged code inside a guest. No devices, no network, no physical access. A container tenant on a shared Arm node is enough.","remediation":"EL3 firmware sets CPUACTLR6_EL1[41]=1 (or IMP_CPUECTLR_EL1[49]=1 on C1-Pro) to disable the offending prefetcher, and exposes it via SMCCC_ARCH_WORKAROUND_4. It is enabled by default in fixed TF-A on vulnerable cores, so the operator task is 'take the OEM firmware build and flash it', with the usual reboot and drain. Disabling a prefetcher is not free - expect single-digit percent regression on pointer-chasing workloads. There is no software-only mitigation you can apply without the firmware update.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-13.html","https://nvd.nist.gov/vuln/detail/CVE-2024-7881"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-54388","cve":"CVE-2025-54388","aliases":[],"title":"Docker / moby: On firewalld reload, published container ports become reachable from outside despite the intended restriction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2025","cvss_score":5.1,"severity":"medium","kev":false,"impact":"On firewalld reload, published container ports become reachable from outside despite the intended restriction","attack_vector":"Unauthenticated network","remediation":"Upgrade Docker Engine; verify iptables/nftables rules after any firewalld reload","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54388"],"status":"curated"},{"id":"CVE-2019-14891","cve":"CVE-2019-14891","aliases":[],"title":"CRI-O: All pod processes share one memory cgroup, so a workload OOM kills conmon and destabilises the node","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2019","cvss_score":5,"severity":"medium","kev":false,"impact":"All pod processes share one memory cgroup, so a workload OOM kills conmon and destabilises the node","attack_vector":"Any tenant workload","remediation":"Upgrade CRI-O; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-14891"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-32760","cve":"CVE-2021-32760","aliases":[],"title":"containerd: Crafted image can change Unix file permissions of existing host files during extraction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2021","cvss_score":5,"severity":"medium","kev":false,"impact":"Crafted image can change Unix file permissions of existing host files during extraction","attack_vector":"Malicious image","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32760"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-21701","cve":"CVE-2022-21701","aliases":[],"title":"Istio: A user with CREATE on Gateway API resources escalates privilege in istiod","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2022","cvss_score":5,"severity":"medium","kev":false,"impact":"A user with CREATE on Gateway API resources escalates privilege in istiod","attack_vector":"Cluster user with namespace access and Gateway API rights","remediation":"Rolling istiod upgrade; restrict Gateway creation","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21701"],"status":"curated"},{"id":"CVE-2022-42292","cve":"CVE-2022-42292","aliases":[],"title":"GeForce Experience installer: Local privesc via symlink following","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce Experience installer","year":"2022","cvss_score":5,"severity":"medium","kev":false,"impact":"Local privesc via symlink following","attack_vector":"Local user","remediation":"Consumer-only; no DC action","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42292","https://github.com/NVIDIA/product-security/tree/main/2023/5384"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:N/I:L/A:H","cwe":["CWE-59"]},{"id":"CVE-2023-0203","cve":"CVE-2023-0203","aliases":[],"title":"ConnectX-5/6/6-DX NIC firmware: NIC DoS (insufficient access-control granularity)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX-5/6/6-DX NIC firmware","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"NIC DoS (insufficient access-control granularity)","attack_vector":"Any unprivileged tenant with a VF","remediation":"Flash NIC firmware 35.1012+; node reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5459/5459.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-1220"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0205","cve":"CVE-2023-0205","aliases":[],"title":"ConnectX-5/6/6-DX NIC firmware: NIC DoS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"ConnectX-5/6/6-DX NIC firmware","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"NIC DoS","attack_vector":"Any unprivileged tenant with a VF","remediation":"Flash NIC firmware 35.1012+; node reboot","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5459/5459.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-1220"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0264","cve":"CVE-2023-0264","aliases":[],"title":"Keycloak: OIDC authentication flaw - attacker reusing data from a same-realm request impersonates a user and mints…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"OIDC authentication flaw - attacker reusing data from a same-realm request impersonates a user and mints session tokens","attack_vector":"Network (remote)","remediation":"Control-plane: URGENT - Keycloak fronts the tenant portal; upgrade + invalidate all sessions","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0264"],"status":"curated"},{"id":"CVE-2023-25809","cve":"CVE-2023-25809","aliases":[],"title":"runc: Rootless runc leaves /sys/fs/cgroup writable inside the container","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2023","cvss_score":5,"severity":"medium","kev":false,"impact":"Rootless runc leaves /sys/fs/cgroup writable inside the container","attack_vector":"Any tenant workload under rootless runc","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25809"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2024-42934","cve":"CVE-2024-42934","aliases":[],"title":"OpenIPMI before 2.0.36: Where this bites an operator is in test and CI infrastructure rather than production nodes: ipmi_sim is what…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenIPMI before 2.0.36","year":"2024","cvss_score":5,"severity":"medium","kev":false,"impact":"Where this bites an operator is in test and CI infrastructure rather than production nodes: ipmi_sim is what teams use to emulate BMCs when developing provisioning automation and firmware tooling. An attacker who can reach a simulator instance crashes it or, with luck, bypasses its authentication - and a compromised CI runner that builds and signs your provisioning images is a supply-chain foothold into the real fleet. The low probability of code execution is the honest read; the availability impact on a build pipeline is the likely one. An out-of-bounds array access on the authentication type in the ipmi_sim BMC simulator, with denial of service the likely outcome and authentication bypass or code execution possible at low probability.","attack_vector":"Network access to a running ipmi_sim instance. In most environments that is a developer workstation or a CI runner rather than the production management VLAN, but CI runners are frequently more reachable than people assume.","remediation":"Package update to OpenIPMI 2.0.36 or later on any host running ipmi_sim - a distribution package update, no firmware and no reboot. The broader hygiene point: BMC simulators used in CI should not be reachable from anything but the test harness, and should not run on a host that also holds production BMC credentials or signing keys.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42934","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2024/42xxx/CVE-2024-42934.json"],"status":"curated"},{"id":"CVE-2024-48936","cve":"CVE-2024-48936","aliases":[],"title":"Slurm: Authentication-handling mistake in stepmgr lets an attacker execute processes under other users' jobs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Slurm","year":"2024","cvss_score":5,"severity":"medium","kev":false,"impact":"Authentication-handling mistake in stepmgr lets an attacker execute processes under other users' jobs; cross-tenant job compromise on a shared GPU cluster","attack_vector":"Any user who can submit a job","remediation":"Upgrade Slurm to 24.05.4+; disable --stepmgr if not needed","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-48936"],"status":"curated"},{"id":"CVE-2025-23260","cve":"CVE-2025-23260","aliases":[],"title":"AIS Operator (AIStore): Improper access control on cluster storage","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"AIS Operator (AIStore)","year":"2025","cvss_score":5,"severity":"medium","kev":false,"impact":"Improper access control on cluster storage","attack_vector":"Tenant with cluster network access","remediation":"Upgrade the operator Helm chart","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23260","https://github.com/NVIDIA/product-security/tree/main/2025/5660"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:C/C:L/I:N/A:N","cwe":["CWE-266"]},{"id":"CVE-2025-23332","cve":"CVE-2025-23332","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): A null-pointer dereference in a Linux driver kernel module, reachable locally, panics the node. Everything…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2025","cvss_score":5,"severity":"medium","kev":false,"impact":"A null-pointer dereference in a Linux driver kernel module, reachable locally, panics the node. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5703. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23332","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-476"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-41413","cve":"CVE-2026-41413","aliases":[],"title":"Istio: A RequestAuthentication jwksUri pointed at an internal service makes istiod issue an unauthenticated request","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Istio","year":"2026","cvss_score":5,"severity":"medium","kev":false,"impact":"A RequestAuthentication jwksUri pointed at an internal service makes istiod issue an unauthenticated request; SSRF from the control plane","attack_vector":"Cluster user with namespace access","remediation":"Rolling istiod upgrade to 1.28.6/1.29.2+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41413"],"status":"curated"},{"id":"CVE-2018-9279","cve":"CVE-2018-9279","aliases":[],"title":"Eaton UPS 9PX 8000 SP web interface: The device's own web page contains the user password in cleartext in the page source. Anyone who gets a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Eaton UPS 9PX 8000 SP web interface","year":"2018","cvss_score":4.9,"severity":"medium","kev":false,"impact":"The device's own web page contains the user password in cleartext in the page source. Anyone who gets a single authenticated view - or a screenshot, or a saved page in a support ticket - has the credential. The companion issue CVE-2018-9280 does the same for the SNMPv3 read and write user passwords, which is worse, because the SNMP write community is a control channel.","attack_vector":"Anyone who can load the UPS web interface, or who obtains a saved copy of the page.","remediation":"Firmware update where available. Rotate the UPS and SNMPv3 credentials, and specifically rotate the SNMP write credential - and then ask whether SNMP write needs to be enabled at all, because on most UPS deployments it does not.","references":["https://www.bishopfox.com/news/2018/10/eaton-ups-9px-8000-sp-multiple-vulnerabilities/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-11245","cve":"CVE-2019-11245","aliases":[],"title":"Kubernetes (kubelet): Container restart runs as uid 0 despite mustRunAsNonRoot","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2019","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Container restart runs as uid 0 despite mustRunAsNonRoot","attack_vector":"Any tenant workload","remediation":"Rolling kubelet upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11245"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2020-11484","cve":"CVE-2020-11484","aliases":[],"title":"NVIDIA DGX BMC (AMI firmware): An administrative BMC user can pull the hash of the BMC/IPMI user password. In practice that means one…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVIDIA DGX BMC (AMI firmware)","year":"2020","cvss_score":4.9,"severity":"medium","kev":false,"impact":"An administrative BMC user can pull the hash of the BMC/IPMI user password. In practice that means one compromised BMC admin session yields offline-crackable credentials that are frequently reused across an entire DGX fleet - a lateral movement multiplier across every node in the rack. DGX-1 before BMC 3.38.30.","attack_vector":"An attacker who already has administrative access to one BMC, including via the hard-coded credentials in the same advisory.","remediation":"Flash the DGX BMC firmware from NVIDIA's DGX firmware update container (DGX-1 to 3.38.30 or later, DGX-2 to 1.06.06 or later; DGX A100 per the bulletin's table). A BMC flash does not require the host OS to reboot but drops out-of-band management for several minutes and NVIDIA recommends a host power cycle afterwards, so treat it as a per-node maintenance window. Rotate every BMC and IPMI credential after the flash - flashing does not invalidate secrets an attacker already pulled. Keep BMCs on an isolated management VLAN with no route from tenant or job networks.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-11484"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-44769","cve":"CVE-2021-44769","aliases":["AMI-SA-2022001","Nozomi Labs BMC firmware research"],"title":"AMI MegaRAC SPx 12 / SPx 13 (BMC TLS certificate generation): Malformed input to the BMC's certificate-generation function permanently wedges the controller, and the only…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx 12 / SPx 13 (BMC TLS certificate generation)","year":"2021","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Malformed input to the BMC's certificate-generation function permanently wedges the controller, and the only documented recovery is a factory reset. On a GPU fleet that is worse than a normal DoS: you lose out-of-band power control, console and virtual media on the affected nodes, and getting them back means an on-site factory reset that also wipes your BMC configuration - accounts, certificates, network settings, alert destinations - so every affected node has to be re-provisioned by hand. A malicious insider or a compromised management account can do this across a rack in minutes and cost you days of remote-hands work.","attack_vector":"Network access to the BMC with high privileges - an administrative BMC account. Because BMC admin credentials are so often cloned across a fleet by the provisioning system, one leaked credential scales this to every node that shares it.","remediation":"Firmware flash to SPx_12-update-5.00 / SPx_13-update-3.00 or later, out-of-band per node, ODM-gated. Config-only risk reduction: unique per-node BMC admin credentials so one leak cannot sweep the fleet, and - critically - keep an exported, version-controlled copy of every BMC's configuration so that if you do have to factory-reset, restoring is automated rather than a per-node manual rebuild.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/AMI-SA-2022001.pdf","https://www.nozominetworks.com/labs/vulnerability-advisories/cve-2021-44769/","https://nvd.nist.gov/vuln/detail/CVE-2021-44769"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-21417","cve":"CVE-2022-21417","aliases":[],"title":"MySQL Server: InnoDB flaw allowing a high-privileged network attacker to cause a repeatable DoS","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MySQL Server","year":"2022","cvss_score":4.9,"severity":"medium","kev":false,"impact":"InnoDB flaw allowing a high-privileged network attacker to cause a repeatable DoS","attack_vector":"Network (remote)","remediation":"Control-plane: quarterly CPU patch","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21417"],"status":"curated"},{"id":"CVE-2022-22488","cve":"CVE-2022-22488","aliases":["IBM X-Force 226337"],"title":"IBM OpenBMC OP910 / OP940 certificate handling (phosphor-certificate-manager lineage): A privileged BMC user who uploads or deletes CA certificates rapidly enough takes the BMC down. Low severity…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IBM OpenBMC OP910 / OP940 certificate handling (phosphor-certificate-manager lineage)","year":"2022","cvss_score":4.9,"severity":"medium","kev":false,"impact":"A privileged BMC user who uploads or deletes CA certificates rapidly enough takes the BMC down. Low severity because it needs an admin account, but the fleet-relevant version is not malice: it is your own certificate-rotation automation. An operator scripting CA distribution across a few hundred BMCs can trip this and take out out-of-band management fleet-wide during what was supposed to be a routine hygiene job.","attack_vector":"Authenticated BMC administrator over the network - including your own automation holding admin credentials.","remediation":"Fixed in later OP910/OP940 firmware; per-node system firmware update with a maintenance window, and not worth a dedicated campaign at this severity. The operational fix is free: rate-limit and serialize certificate operations in your BMC automation, and stagger fleet-wide certificate pushes rather than fanning out at full concurrency.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-22488","https://www.ibm.com/support/pages/node/6840155","https://exchange.xforce.ibmcloud.com/vulnerabilities/226337"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-22084","cve":"CVE-2023-22084","aliases":[],"title":"MySQL Server: InnoDB flaw - a high-privileged network attacker can hang or repeatedly crash the server","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"MySQL Server","year":"2023","cvss_score":4.9,"severity":"medium","kev":false,"impact":"InnoDB flaw - a high-privileged network attacker can hang or repeatedly crash the server","attack_vector":"Network (remote)","remediation":"Control-plane: quarterly CPU patch on managed MySQL","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-22084"],"status":"curated"},{"id":"CVE-2023-29153","cve":"CVE-2023-29153","aliases":[],"title":"Intel SPS firmware: Uncontrolled resource consumption in SPS firmware lets a privileged user deny service to the management…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SPS firmware","year":"2023","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Uncontrolled resource consumption in SPS firmware lets a privileged user deny service to the management engine over the network. Losing the management engine means losing out-of-band recovery on that node - which is exactly when you need it.","attack_vector":"Privileged user with network access to the affected interface.","remediation":"Fixed in Intel CSME/SPS firmware, which reaches you as an OEM BIOS or firmware package - not as a microcode or OS update. That means: wait for your server vendor to ship it, drain the node, flash, and reboot. OEM availability is the long pole and routinely lags the Intel advisory by one or more quarters on server platforms. Track it per platform SKU, because vendors ship these unevenly across their own product lines.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-29153","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01003.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-46118","cve":"CVE-2023-46118","aliases":[],"title":"RabbitMQ: HTTP API enforces no request body limit","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"RabbitMQ","year":"2023","cvss_score":4.9,"severity":"medium","kev":false,"impact":"HTTP API enforces no request body limit -> authenticated user exhausts node memory (DoS)","attack_vector":"Network (remote)","remediation":"Control-plane: broker upgrade; set max_message_size","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46118"],"status":"curated"},{"id":"CVE-2024-0116","cve":"CVE-2024-0116","aliases":[],"title":"NVIDIA Triton Inference Server: An out-of-bounds read triggered by releasing a shared memory region while it is still in use crashes the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2024","cvss_score":4.9,"severity":"medium","kev":false,"impact":"An out-of-bounds read triggered by releasing a shared memory region while it is still in use crashes the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5565. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0116","https://github.com/NVIDIA/product-security/tree/main/2024/5565"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-125"]},{"id":"CVE-2024-23444","cve":"CVE-2024-23444","aliases":[],"title":"Elasticsearch: elasticsearch-certutil --csr writes the private key to disk unencrypted despite --pass","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Elasticsearch","year":"2024","cvss_score":4.9,"severity":"medium","kev":false,"impact":"elasticsearch-certutil --csr writes the private key to disk unencrypted despite --pass","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + reissue any cert whose key was generated this way","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-23444"],"status":"curated"},{"id":"CVE-2024-53880","cve":"CVE-2024-53880","aliases":[],"title":"Triton Inference Server: DoS (integer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2024","cvss_score":4.9,"severity":"medium","kev":false,"impact":"DoS (integer overflow)","attack_vector":"Client of the inference endpoint","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53880","https://github.com/NVIDIA/product-security/tree/main/2025/5612"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-190"]},{"id":"CVE-2025-26482","cve":"CVE-2025-26482","aliases":["DSA-2025-046"],"title":"Dell PowerEdge Server BIOS + iDRAC9 (information disclosure): Information disclosure spanning both the BIOS and iDRAC9 on a very large PowerEdge model list - and that list…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell PowerEdge Server BIOS + iDRAC9 (information disclosure)","year":"2025","cvss_score":4.9,"severity":"medium","kev":false,"impact":"Information disclosure spanning both the BIOS and iDRAC9 on a very large PowerEdge model list - and that list explicitly includes the GPU platforms: XE9680, XE9680L, XE9640 and XE8640. Low severity in isolation; what makes it worth tracking is coverage. If you run Dell GPU nodes, this one almost certainly applies to your exact SKUs, and disclosed platform/firmware detail is the reconnaissance that makes a later BMC or BIOS exploit reliable rather than a guess.","attack_vector":"A high-privilege attacker with remote access - i.e. someone who already holds an administrative iDRAC credential. This is a post-compromise information-leak rather than an entry point, which is why the score is moderate.","remediation":"Two separate rollouts. The iDRAC9 side is an out-of-band firmware flash, per-node, no host reboot, no drain. The BIOS side needs a System BIOS update that applies only on the next reboot, so it costs a drain of running training jobs - for XE9680-class nodes that is a real scheduling problem, since those are the machines you least want to take down. Sequence the iDRAC flash immediately and batch the BIOS update into the next planned maintenance window. Per-model version floors are in the advisory (e.g. 2.5.4 for the R660/R760 family, 1.2.6 for the R470/R570/R670/R770 family).","references":["https://www.dell.com/support/kbdoc/en-us/000370138/dsa-2025-046-security-update-for-dell-poweredge-server-and-dell-idrac9-for-information-disclosure-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2025-26482"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-11251","cve":"CVE-2019-11251","aliases":[],"title":"Kubernetes (kubectl): Double-symlink in tar output escapes the kubectl cp destination","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2019","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Double-symlink in tar output escapes the kubectl cp destination","attack_vector":"Malicious image","remediation":"Upgrade kubectl on operator and CI machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-11251"],"status":"curated"},{"id":"CVE-2019-6195","cve":"CVE-2019-6195","aliases":[],"title":"Lenovo XClarity Controller (XCC): Authorization bypass — a low-privilege authenticated user gains read access beyond their role","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC)","year":"2019","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Authorization bypass — a low-privilege authenticated user gains read access beyond their role","attack_vector":"Network / XCC web","remediation":"XCC firmware update; low CVSS but relevant where BMC access is delegated to tenants or remote-hands staff","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-6195"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-15211","cve":"CVE-2020-15211","aliases":[],"title":"TensorFlow Lite (flatbuffer models): Out-of-bounds via duplicate tensor indices in flatbuffer models","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"TensorFlow Lite (flatbuffer models)","year":"2020","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Out-of-bounds via duplicate tensor indices in flatbuffer models","attack_vector":"Customer-supplied TFLite model","remediation":"Patch; edge/CPU inference tiers only","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-15211"],"status":"curated"},{"id":"CVE-2022-3466","cve":"CVE-2022-3466","aliases":[],"title":"CRI-O: Shipped OpenShift CRI-O builds regressed the CVE-2022-2995 fix","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Shipped OpenShift CRI-O builds regressed the CVE-2022-2995 fix; supply-chain style reintroduction of a fixed bug","attack_vector":"Any tenant workload on an affected OCP build","remediation":"Verify CRI-O build version explicitly, not just the OCP version; upgrade and drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-3466"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-31339","cve":"CVE-2023-31339","aliases":[],"title":"ARM Trusted Firmware in AMD Zynq UltraScale+ MPSoC/RFSoC: Improper input validation in the ARM Trusted Firmware used on AMD's Zynq UltraScale+ parts allows…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ARM Trusted Firmware in AMD Zynq UltraScale+ MPSoC/RFSoC","year":"2023","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Improper input validation in the ARM Trusted Firmware used on AMD's Zynq UltraScale+ parts allows out-of-bounds reads and data leakage. Relevant to datacenter operators through the side door: Zynq and Versal parts show up as SmartNIC, DPU, storage-controller and management-plane silicon inside servers, so this is firmware running on your network path rather than on your compute path.","attack_vector":"Local to the device, requires privileged access to the ATF interface on the Zynq part.","remediation":"Fixed in updated ARM Trusted Firmware from AMD/Xilinx. Delivery depends entirely on who integrated the part - a SmartNIC vendor, a storage OEM, your own board team - so tracing the update path is often harder than applying it. Requires a device firmware update and a reset of the affected card. Inventory which AMD/Xilinx adaptive SoCs are in your servers; most operators cannot answer that question, which is the real finding.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31339","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-7258","cve":"CVE-2023-7258","aliases":[],"title":"gVisor: Reference-counting bug in mount-point tracking panics the sandbox","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"gVisor","year":"2023","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Reference-counting bug in mount-point tracking panics the sandbox","attack_vector":"A tenant running as root inside the sandbox with mount permission","remediation":"Upgrade runsc; restart sandboxed pods","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-7258"],"status":"curated"},{"id":"CVE-2025-24513","cve":"CVE-2025-24513","aliases":[],"title":"ingress-nginx: auth-secret file path traversal in the controller","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2025","cvss_score":4.8,"severity":"medium","kev":false,"impact":"auth-secret file path traversal in the controller","attack_vector":"Cluster user who can create Ingress objects","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-29949","cve":"CVE-2025-29949","aliases":[],"title":"AMD Secure Processor bootloader - legacy recovery mode: Insufficient input sanitisation in the ASP bootloader's legacy recovery path lets an attacker write out of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor bootloader - legacy recovery mode","year":"2025","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Insufficient input sanitisation in the ASP bootloader's legacy recovery path lets an attacker write out of bounds and corrupt Secure DRAM, bricking the boot flow. The outcome is denial of service at the firmware level - a node that will not come back up, which on a GPU fleet means an RMA-shaped hole rather than a reboot.","attack_vector":"Local, and only through legacy recovery mode - so it needs an attacker who can force the platform into recovery, which usually means firmware-level or physical access.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. Where the platform allows it, disabling legacy ASP recovery mode in SBIOS removes the reachable path without waiting for the OEM.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29949","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24147","cve":"CVE-2026-24147","aliases":[],"title":"Triton Inference Server: Unauthorized file access via path traversal","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Unauthorized file access via path traversal","attack_vector":"Authenticated inference client","remediation":"Upgrade Triton; redeploy serving images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24147","https://github.com/NVIDIA/product-security/tree/main/2026/5816"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:L","cwe":["CWE-22"]},{"id":"CVE-2026-41174","cve":"CVE-2026-41174","aliases":[],"title":"Traefik: Cross-namespace isolation not enforced in the Kubernetes CRD provider","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Traefik","year":"2026","cvss_score":4.8,"severity":"medium","kev":false,"impact":"Cross-namespace isolation not enforced in the Kubernetes CRD provider","attack_vector":"Cluster user with namespace access","remediation":"Rolling Traefik upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41174"],"status":"curated"},{"id":"CVE-2018-3626","cve":"CVE-2018-3626","aliases":[],"title":"Intel SGX SDK (Edger8r generated code, side channel): Edger8r generated bridge code that was susceptible to a side channel, so a local user could recover…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (Edger8r generated code, side channel)","year":"2018","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Edger8r generated bridge code that was susceptible to a side channel, so a local user could recover information crossing the enclave boundary. Earliest member of the generated-bridge-code family.","attack_vector":"Local user interacting with a vulnerable enclave.","remediation":"Rebuild enclaves with SGX SDK 2.1.2 (Linux) / 1.9.6 (Windows) or later and re-attest.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-3626","https://security-center.intel.com/advisory.aspx?intelid=INTEL-SA-00117&languageid=en-fr"],"status":"curated"},{"id":"CVE-2020-5967","cve":"CVE-2020-5967","aliases":[],"title":"NVIDIA Linux GPU Display Driver, UVM driver (nvidia-uvm.ko): A race in the Unified Virtual Memory kernel module lets a local process wedge or crash the driver. UVM is on…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Linux GPU Display Driver, UVM driver (nvidia-uvm.ko)","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race in the Unified Virtual Memory kernel module lets a local process wedge or crash the driver. UVM is on the hot path for managed-memory CUDA workloads, so on a Linux training node this is a way for any job with GPU access to take out the whole host's GPUs. Ubuntu shipped it as a security update, so it is real on distro-packaged fleets.","attack_vector":"Any local user or container that has /dev/nvidia-uvm mapped in - which is every GPU container under the standard container toolkit configuration.","remediation":"Install the fixed Linux GPU Display Driver branch. nvidia.ko / nvidia-uvm.ko cannot be replaced while any process holds a GPU, so plan a node drain: cordon the node, stop every CUDA job and GPU container, unload the modules or reboot, install, reload. Container runtimes that bind-mount the driver libraries (nvidia-container-toolkit) need restarting so running pods pick up the new userspace. No firmware flash.","references":["https://usn.ubuntu.com/4404-1/","https://usn.ubuntu.com/4404-2/","https://nvd.nist.gov/vuln/detail/CVE-2020-5967"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-8563","cve":"CVE-2020-8563","aliases":[],"title":"Kubernetes (cloud-controller-manager): vSphere cloud credentials leaked into logs at verbosity 4+","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (cloud-controller-manager)","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"vSphere cloud credentials leaked into logs at verbosity 4+","attack_vector":"Anyone with log-pipeline read access","remediation":"Reduce log verbosity; rotate cloud credentials; scrub logs","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8563"],"status":"curated"},{"id":"CVE-2020-8564","cve":"CVE-2020-8564","aliases":[],"title":"Kubernetes: Malformed docker config leaks registry pull secrets into logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Malformed docker config leaks registry pull secrets into logs","attack_vector":"Anyone with log read access","remediation":"Reduce verbosity; rotate registry pull secrets","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8564"],"status":"curated"},{"id":"CVE-2020-8565","cve":"CVE-2020-8565","aliases":[],"title":"Kubernetes: Authorization and bearer tokens written to logs at verbosity 9","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Authorization and bearer tokens written to logs at verbosity 9","attack_vector":"Anyone with log read access","remediation":"Reduce verbosity; rotate tokens","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8565"],"status":"curated"},{"id":"CVE-2020-8566","cve":"CVE-2020-8566","aliases":[],"title":"Kubernetes (kube-controller-manager): Ceph RBD admin secrets written to controller-manager logs","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-controller-manager)","year":"2020","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Ceph RBD admin secrets written to controller-manager logs","attack_vector":"Anyone with log read access","remediation":"Reduce verbosity; rotate Ceph credentials","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8566"],"status":"curated"},{"id":"CVE-2021-1117","cve":"CVE-2021-1117","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): Improper input validation in the escape handler under specific configurations, giving an unprivileged local…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2021","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Improper input validation in the escape handler under specific configurations, giving an unprivileged local user a denial of service against the driver.","attack_vector":"Any local unprivileged user with GPU device access on a Windows host in the affected configuration.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1117"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-26318","cve":"CVE-2021-26318","aliases":[],"title":"AMD processors - PREFETCH instruction timing and power side channel: MULTI-TENANT ISOLATION: Timing and power measurements around the x86 PREFETCH instructions leak kernel…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - PREFETCH instruction timing and power side channel","year":"2021","cvss_score":4.7,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Timing and power measurements around the x86 PREFETCH instructions leak kernel address space layout on some AMD CPUs. Defeating KASLR is not itself a compromise, but it is the step that converts an unreliable kernel memory-corruption bug - and the amdgpu/amdkfd driver long tail is full of them - into a reliable exploit. Treat it as an exploitability multiplier for everything else in this database.","attack_vector":"Local, unprivileged. Reachable from inside a container.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Kernel-side mitigation also exists. Because the value of this bug is chaining, the practical defence is to keep the kernel memory-safety patches current rather than to treat KASLR as a real boundary.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26318","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-27672","cve":"CVE-2022-27672","aliases":["Cross-Thread Return Address Predictions"],"title":"AMD processors with SMT - speculative execution across SMT mode switch: MULTI-TENANT ISOLATION: With SMT enabled, certain AMD processors speculatively execute using a branch target…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors with SMT - speculative execution across SMT mode switch","year":"2022","cvss_score":4.7,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: With SMT enabled, certain AMD processors speculatively execute using a branch target supplied by the sibling thread after an SMT mode switch. One hardware thread therefore steers the other's speculation - a Spectre-v2-shaped attack that crosses the thread boundary rather than the process boundary, so it defeats isolation between two tenants sharing a physical core even when they are in different VMs.","attack_vector":"Local, requires SMT enabled and attacker/victim on sibling threads of the same core - the default arrangement on a bin-packed cluster.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. The zero-cost mitigation available today is scheduling policy rather than patching: these attacks need the attacker and victim co-resident on sibling SMT threads, so either disable SMT (costing roughly 10-25% throughput on most inference and training workloads) or enforce core isolation so no two tenants ever share a physical core. On a GPU fleet the CPU is rarely the bottleneck, which makes disabling SMT a cheaper trade than it looks on paper.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-27672","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-28693","cve":"CVE-2022-28693","aliases":["RSBA","Return Stack Buffer Alternate"],"title":"Intel processors (return stack buffer alternate prediction): When the return stack buffer underflows, the processor falls back to an alternate predictor that can be…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (return stack buffer alternate prediction)","year":"2022","cvss_score":4.7,"severity":"medium","kev":false,"impact":"When the return stack buffer underflows, the processor falls back to an alternate predictor that can be influenced by another context, leaking information. Same structural theme as PBRSB: the return predictor is not as well isolated as the ISA implies.","attack_vector":"Local authorised code on the host.","remediation":"Mitigated by an Intel microcode update plus OS/hypervisor changes. Microcode for this class is normally shipped by your distribution as an early-loadable image, so you can deploy it with a package update and a reboot without waiting for an OEM BIOS release - that distinction is the difference between a week and a quarter. Verify after reboot by reading /sys/devices/system/cpu/vulnerabilities/ rather than assuming the package took effect.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28693","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00707.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2022-49781","cve":"CVE-2022-49781","aliases":[],"title":"Linux perf/x86/amd - race between amd_pmu_enable_all, perf NMI and throttling: A race between AMD PMU enablement, the perf NMI handler and event throttling crashes the host. Performance…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux perf/x86/amd - race between amd_pmu_enable_all, perf NMI and throttling","year":"2022","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race between AMD PMU enablement, the perf NMI handler and event throttling crashes the host. Performance counters are enabled by every profiling and observability agent, and throttling kicks in precisely when counters are busy - so the crash window opens under exactly the monitoring load an AI cluster runs continuously.","attack_vector":"Local, through perf event enablement. Reachable by node observability agents and, where perf_event_paranoid is relaxed, by tenants.","remediation":"Distro kernel update plus reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-49781"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0192","cve":"CVE-2023-0192","aliases":[],"title":"GPU Display Driver: Improper privilege management","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Improper privilege management","attack_vector":"Local operator / tenant","remediation":"Driver upgrade at next maintenance window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0192","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:H/I:N/A:L","cwe":["CWE-269"]},{"id":"CVE-2023-20583","cve":"CVE-2023-20583","aliases":["Collide+Power"],"title":"AMD processors - power side channel on cache line data changes: MULTI-TENANT ISOLATION: An authenticated attacker who can read CPU power consumption can infer data as it…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - power side channel on cache line data changes","year":"2023","cvss_score":4.7,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An authenticated attacker who can read CPU power consumption can infer data as it moves through a shared cache line, because the energy cost of a write depends on how many bits flip. Collide+Power generalised this into a technique that works against any victim sharing the cache hierarchy, without needing the victim to have any specific code pattern. Leakage rates are low - this is not a fast exfiltration channel - but it is a channel that exists between tenants on the same socket regardless of what your isolation configuration says.","attack_vector":"Local, authenticated, needs access to power/energy telemetry (RAPL-style interfaces) and cache co-residency with the victim. Reachable from a container if you expose power telemetry into it, which some GPU and CPU monitoring stacks do.","remediation":"Mitigated primarily by restricting who can read fine-grained power telemetry: on Linux, the RAPL/energy interfaces should be root-only (this was tightened upstream), and you should not be passing power monitoring into tenant containers. Check what your DCGM-equivalent AMD telemetry stack exposes and to whom - operators often mount host telemetry paths into monitoring sidecars that tenants can reach. Kernel/driver-level fix plus a config review; no firmware flash and no reboot if you are only tightening permissions. The zero-cost mitigation available today is scheduling policy rather than patching: these attacks need the attacker and victim co-resident on sibling SMT threads, so either disable SMT (costing roughly 10-25% throughput on most inference and training workloads) or enforce core isolation so no two tenants ever share a physical core. On a GPU fleet the CPU is rarely the bottleneck, which makes disabling SMT a cheaper trade than it looks on paper.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20583","https://collidepower.com/","https://www.amd.com/en/resources/product-security.html"],"status":"curated"},{"id":"CVE-2023-4641","cve":"CVE-2023-4641","aliases":[],"title":"shadow-utils: Possible password leak during passwd(1) change (uninitialised memory)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"shadow-utils","year":"2023","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Possible password leak during passwd(1) change (uninitialised memory)","attack_vector":"Local user","remediation":"Package update; no reboot","references":["https://access.redhat.com/security/cve/CVE-2023-4641"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-21792","cve":"CVE-2024-21792","aliases":[],"title":"Intel Neural Compressor (TOCTOU): A time-of-check/time-of-use race in Neural Compressor lets an authenticated local user win a window and read…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel Neural Compressor (TOCTOU)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A time-of-check/time-of-use race in Neural Compressor lets an authenticated local user win a window and read information they should not. Low severity on its own; useful as a step in a chain on a shared optimisation node.","attack_vector":"Authenticated local user on the node running Neural Compressor.","remediation":"Upgrade to 2.5.0 or later. Package update, service restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21792","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01109.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-2201","cve":"CVE-2024-2201","aliases":[],"title":"Intel CPU (Native BHI): Native Branch History Injection - unprivileged user leaks kernel memory despite eIBRS","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Intel CPU (Native BHI)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Native Branch History Injection - unprivileged user leaks kernel memory despite eIBRS; also XSA-456 for Xen","attack_vector":"Any tenant process in a container; tenant VM guest","remediation":"Kernel mitigation (BHI_DIS_S or software sequence) + reboot; standing perf cost. Affects Intel Xeon hosts under GPU nodes","references":["https://access.redhat.com/security/cve/CVE-2024-2201"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-27040","cve":"CVE-2024-27040","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add 'replay' NULL check in 'edp_set_replay_allow_active()'","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27040","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-32473","cve":"CVE-2024-32473","aliases":[],"title":"Docker / moby: IPv6 not disabled on interfaces where it should be","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"IPv6 not disabled on interfaces where it should be; unexpected reachability of containers over IPv6","attack_vector":"Any pod on the cluster network","remediation":"Upgrade moby; audit IPv6 firewall posture","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-32473"],"status":"curated"},{"id":"CVE-2024-36024","cve":"CVE-2024-36024","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Disable idle reallow as part of command/gpint execution","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36024","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-42227","cve":"CVE-2024-42227","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Fix overlapping copy within dml_core_mode_programming","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42227","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46851","cve":"CVE-2024-46851","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Avoid race between dcn10_set_drr() and dc_state_destruct()","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46851","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-46870","cve":"CVE-2024-46870","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2024","cvss_score":4.7,"severity":"medium","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Disable DMCUB timeout for DCN35","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-46870","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23269","cve":"CVE-2025-23269","aliases":[],"title":"Jetson Xavier / Orin: Info disclosure via side channel","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson Xavier / Orin","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Info disclosure via side channel","attack_vector":"Local attacker on the device","remediation":"Flash JetPack; edge fleet","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23269","https://github.com/NVIDIA/product-security/tree/main/2025/5662"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-1423"]},{"id":"CVE-2025-37899","cve":"CVE-2025-37899","aliases":[],"title":"Linux kernel (ksmbd): Use-after-free in ksmbd session logoff (found by an LLM-assisted audit)","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel (ksmbd)","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Use-after-free in ksmbd session logoff (found by an LLM-assisted audit)","attack_vector":"Authenticated network to an exposed ksmbd share","remediation":"Livepatchable; otherwise drain + reboot. Best answer remains not shipping ksmbd","references":["https://access.redhat.com/security/cve/CVE-2025-37899"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-38104","cve":"CVE-2025-38104","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): Missing or insufficient validation of user-supplied parameters in the amdgpu power management…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu power management (SMU/powerplay). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Replace Mutex with Spinlock for RLCG register access to avoid Priority Inversion in SRIOV","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-38104","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-64432","cve":"CVE-2025-64432","aliases":[],"title":"KubeVirt: Flawed aggregation-layer authentication flow enables RBAC bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"KubeVirt","year":"2025","cvss_score":4.7,"severity":"medium","kev":false,"impact":"Flawed aggregation-layer authentication flow enables RBAC bypass","attack_vector":"Cluster user with namespace access","remediation":"Upgrade KubeVirt","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64432"],"status":"curated"},{"id":"CVE-2026-24199","cve":"CVE-2026-24199","aliases":[],"title":"GPU Display Driver: DoS (race in GPU resource allocation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2026","cvss_score":4.7,"severity":"medium","kev":false,"impact":"DoS (race in GPU resource allocation)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24199","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-362"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-33082","cve":"CVE-2021-33082","aliases":["INTEL-SA-00563","Solidigm SA-000563"],"title":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware - NVMe Sanitize (Block Erase) leaves prior data recoverable: Sensitive data is not removed before the media is reused. The drive accepts an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel / Solidigm SSD, SSD DC and Optane SSD firmware - NVMe Sanitize (Block Erase) leaves prior data recoverable","year":"2021","cvss_score":4.6,"severity":"medium","kev":false,"impact":"Sensitive data is not removed before the media is reused. The drive accepts an NVMe Sanitize with Block Erase, reports completion, and leaves the previous contents recoverable. BREAKS TENANT HANDOFF in the most direct way in this whole category: this IS the command most bare-metal reclaim pipelines issue between customers. Your automation records a successful sanitize, the node goes back into the pool, and the next tenant can carve the previous tenant's datasets, model checkpoints and any credentials that touched local scratch out of the flash. Worse, it fails clean - there is no error to alert on, so a fleet can run in this state for years and every audit log says the erase succeeded.","attack_vector":"The next tenant with root on the reclaimed bare-metal host, or anyone who obtains the physical drive later (RMA, decommission, resale). The attacker does not need to break anything - they read what the sanitize left behind.","remediation":"Two moves, do both. (1) Immediate policy/workaround, no downtime: stop using Sanitize with Block Erase (SANACT=04h in the Sanitize command's block-erase form) in your reclaim path and switch to Sanitize with Crypto Erase, or Format NVM with Crypto Erase / User Data Erase - this is the vendor's own prescribed workaround and it is a one-line change in most reclaim scripts, so ship it today. (2) Flash affected drives to fixed firmware; this needs the drive quiesced, usually a node drain, and Solidigm Storage Tool / Intel MAS with per-SKU firmware images, so plan it as a rolling maintenance campaign. Note that several older SKUs in the same advisory family are explicitly end-of-support with no fix planned unless a customer asks - for those, the workaround IS the remediation. Independently: layer LUKS/dm-crypt with an operator-held key on all tenant-visible local NVMe, so reclaim means destroying your key rather than trusting the drive's report. Verifying erase across a 10,000-drive fleet by actually reading back raw blocks is a weeks-long operation and essentially no operator does it - which is precisely why a firmware that lies about sanitize success goes undetected.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33082","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00563.html","https://www.solidigm.com/support-page/support-security.html","https://www.solidigm.com/content/dam/solidigm/en/site/support/support-community/cve-(security)/documents/public-security-advisory-v2.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-40633","cve":"CVE-2022-40633","aliases":["ICSA-23-061-03"],"title":"Rittal CMC III cabinet lock / access-card system: The access cards used to open control cabinets secured with Rittal CMC III locks can be cloned. In a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rittal CMC III cabinet lock / access-card system","year":"2022","cvss_score":4.6,"severity":"medium","kev":false,"impact":"The access cards used to open control cabinets secured with Rittal CMC III locks can be cloned. In a datacenter this is the rack-level and cabinet-level access boundary - the CMC III system is exactly what many operators use to electronically lock individual racks and cabinets and to log who opened them, and it is often the control that lets a provider tell a customer 'only your staff can open your rack'. Cloning a card defeats that, and it defeats it invisibly: the lock logs a legitimate card opening a legitimate cabinet. Someone at an opened rack can pull NVMe drives holding model weights and customer data, attach a console to a node's serial or VGA port, plug into the out-of-band management switch that fronts every BMC in the row, or insert a hardware implant on a management link. The CVSS of 4.6 reflects the physical-access precondition, not the consequence - for a bare-metal GPU provider whose entire isolation story is physical, this is a boundary failure, and one that also breaks tenant handoff because the same credentials and the same locks carry across tenancies.","attack_vector":"Physical proximity to a valid card, then physical presence at the cabinet. No network access is involved. The precondition that matters is that the attacker must already be inside the hall - so this is the second stage after tailgating, a compromised hall-door credential, or legitimate access as a contractor, landlord technician, or another tenant's staff in a shared hall.","remediation":"Not fixable by patching - it is the credential technology. Rittal's guidance is to move to a more secure card technology where the hardware supports it; in practice that means replacing readers and reissuing cards, a per-cabinet hardware cost. Compensating controls that work today: tamper alarms on cabinet doors wired into a system separate from the lock itself, camera coverage of aisles with retention long enough to review, and a policy that any cabinet-open event is reconciled against a work order rather than just logged. In a shared hall, treat the cabinet lock as a deterrent rather than a boundary, and put anything that genuinely requires isolation behind a full cage with a second access factor.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-061-03","https://nvd.nist.gov/vuln/detail/CVE-2022-40633"],"status":"curated"},{"id":"CVE-2023-6134","cve":"CVE-2023-6134","aliases":[],"title":"Keycloak: Redirect scheme filtering bypassed by appending a wildcard","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2023","cvss_score":4.6,"severity":"medium","kev":false,"impact":"Redirect scheme filtering bypassed by appending a wildcard -> XSS and follow-on attacks","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; tighten allowed redirect URIs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6134"],"status":"curated"},{"id":"CVE-2023-6927","cve":"CVE-2023-6927","aliases":[],"title":"Keycloak: Wildcard in the JARM form_post.jwt response mode","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Keycloak","year":"2023","cvss_score":4.6,"severity":"medium","kev":false,"impact":"Wildcard in the JARM form_post.jwt response mode -> steal authorization codes and tokens (bypasses the CVE-2023-6134 fix)","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade + remove wildcards from client redirect URIs","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-6927"],"status":"curated"},{"id":"CVE-2024-40635","cve":"CVE-2024-40635","aliases":[],"title":"containerd: UID:GID larger than 32-bit signed max wraps to 0, silently running the container as root","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2024","cvss_score":4.6,"severity":"medium","kev":false,"impact":"UID:GID larger than 32-bit signed max wraps to 0, silently running the container as root","attack_vector":"Malicious image setting a huge numeric USER","remediation":"Rolling containerd upgrade with node drain; add admission check on runAsUser","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-40635"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-0031","cve":"CVE-2025-0031","aliases":[],"title":"AMD SEV firmware - use-after-free allowing a SINGLE_SOCKET guest to activate on the wrong socket (AMD-SB-3023): MULTI-TENANT ISOLATION: A use-after-free in SEV firmware lets a migrated guest whose…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - use-after-free allowing a SINGLE_SOCKET guest to activate on the wrong socket (AMD-SB-3023)","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in SEV firmware lets a migrated guest whose policy specifies SINGLE_SOCKET be activated on a different socket than its migration agent. The guest chose that policy to constrain where its keys and memory live; violating it silently means a tenant's stated confidential-computing constraint is not being enforced, and they have no way to tell.","attack_vector":"Malicious hypervisor driving guest migration.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. Worth flagging to any tenant who sets socket-scoped SNP guest policies - if you sell policy enforcement as a feature, this is a period during which it was not enforced.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0031","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3023.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2025-23292","cve":"CVE-2025-23292","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): SQL injection in the DLS appliance reaching a high integrity impact and partial denial of service in the UI…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"SQL injection in the DLS appliance reaching a high integrity impact and partial denial of service in the UI - an authenticated attacker can alter licensing records.","attack_vector":"Adjacent network, high privileges, user interaction. An admin-level account on the licensing appliance.","remediation":"Update the DLS appliance per bulletin 5705 and review licensing records for tampering. Cost: appliance restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23292","https://github.com/NVIDIA/product-security/tree/main/2025/5705"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:H/UI:R/S:U/C:N/I:H/A:L","cwe":["CWE-943"]},{"id":"CVE-2025-29943","cve":"CVE-2025-29943","aliases":[],"title":"AMD CPU pipeline configuration - SEV-SNP guest stack pointer corruption: MULTI-TENANT ISOLATION: A write-what-where condition in CPU pipeline configuration lets an admin-privileged…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU pipeline configuration - SEV-SNP guest stack pointer corruption","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A write-what-where condition in CPU pipeline configuration lets an admin-privileged attacker corrupt the stack pointer inside an SEV-SNP guest. Corrupting a guest's stack pointer from outside is a control-flow attack on a VM whose memory the host is not supposed to be able to touch - a route to steering execution inside a confidential workload.","attack_vector":"Local, admin-privileged host attacker.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-29943","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-47291","cve":"CVE-2025-47291","aliases":[],"title":"containerd: User-namespaced containers not placed under the Kubernetes cgroup, defeating resource limits","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"containerd","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"User-namespaced containers not placed under the Kubernetes cgroup, defeating resource limits; noisy-neighbour / node DoS","attack_vector":"Any tenant workload using user namespaces","remediation":"Rolling containerd upgrade with node drain","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-47291"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-48517","cve":"CVE-2025-48517","aliases":[],"title":"AMD SEV firmware - ASID range enforcement between SEV-ES and SEV-SNP guests: MULTI-TENANT ISOLATION: A malicious hypervisor can launch a SEV-ES guest using an ASID from the range…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - ASID range enforcement between SEV-ES and SEV-SNP guests","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A malicious hypervisor can launch a SEV-ES guest using an ASID from the range reserved for SEV-SNP guests. ASIDs key the memory encryption, so overlapping the ranges lets a weaker-protected ES guest sit where an SNP guest's protections were assumed - a partial confidentiality loss for the SNP tenant. The interesting part is that the attack uses a legitimate hypervisor operation with an out-of-range parameter rather than any memory-safety bug.","attack_vector":"Requires hypervisor privilege and the ability to launch guests - i.e. the cloud operator or anyone who compromises the control plane.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48517","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-56548","cve":"CVE-2025-56548","aliases":["PT-2025-171"],"title":"Broadcom NetXtreme-E network adapter firmware: The lower-severity half of the same Positive Technologies NetXtreme-E firmware disclosure. Carried here…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Broadcom NetXtreme-E network adapter firmware","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"The lower-severity half of the same Positive Technologies NetXtreme-E firmware disclosure. Carried here because it is fixed by the same firmware image — an operator who patches for the 8.2 issue gets this for free, and one who tracks only high-severity CVEs will never see it listed.","attack_vector":"Adapter firmware interface; same exposure surface as the companion issue.","remediation":"Same NetXtreme-E firmware update; NIC flash and cold power cycle. No separate action needed once the fixed image is deployed.","references":["https://global.ptsecurity.com/en/about/news/pt-expert-helped-patch-vulnerabilities-broadcom-network-adapter-firmware/","https://nvd.nist.gov/vuln/detail/CVE-2025-56548"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-66664","cve":"CVE-2025-66664","aliases":[],"title":"AMD Secure Processor TEE SOC driver - SR-IOV GFX firmware load command: MULTI-TENANT ISOLATION: A malformed DRV_SOC_CMD_ID_LOAD_GFX_IP_FW SR-IOV command causes an out-of-bounds read…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor TEE SOC driver - SR-IOV GFX firmware load command","year":"2025","cvss_score":4.6,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A malformed DRV_SOC_CMD_ID_LOAD_GFX_IP_FW SR-IOV command causes an out-of-bounds read in the ASP's TEE SOC driver. This one matters specifically because it sits on the SR-IOV path - the mechanism by which a single GPU is carved up between virtual functions belonging to different tenants. A guest VF driver reaching a firmware-load command handler is precisely the boundary GPU virtualisation is supposed to hold.","attack_vector":"Reachable from an SR-IOV virtual function, i.e. from a guest VM that has been assigned a GPU VF - so this is guest-to-platform, not merely host-local. Only applies where GPU SR-IOV virtualisation is actually enabled.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. If you run GPU SR-IOV with untrusted tenants in the VFs, this is worth prioritising over its 4.6 score; if you do passthrough of whole physical GPUs instead, the path is not exercised.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-66664","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-3473","cve":"CVE-2021-3473","aliases":[],"title":"Lenovo XClarity Controller: Backup/restore password written to an internal XCC log buffer — credential leak to anyone who can pull BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller","year":"2021","cvss_score":4.5,"severity":"medium","kev":false,"impact":"Backup/restore password written to an internal XCC log buffer — credential leak to anyone who can pull BMC logs","attack_vector":"Local/network, authenticated","remediation":"XCC firmware update plus rotation of any XCC backup passwords used during fleet provisioning","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3473"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23252","cve":"CVE-2025-23252","aliases":[],"title":"NVDebug tool: Info disclosure / privesc via diagnostic tool","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVDebug tool","year":"2025","cvss_score":4.5,"severity":"medium","kev":false,"impact":"Info disclosure / privesc via diagnostic tool","attack_vector":"Local operator","remediation":"Upgrade NVDebug on nodes","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23252","https://github.com/NVIDIA/product-security/tree/main/2025/5651"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:H/UI:R/S:U/C:H/I:N/A:N","cwe":["CWE-1244"]},{"id":"CVE-2025-23274","cve":"CVE-2025-23274","aliases":[],"title":"CUDA Toolkit: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":4.5,"severity":"medium","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23274","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-125"]},{"id":"CVE-2025-4166","cve":"CVE-2025-4166","aliases":[],"title":"HashiCorp Vault: KV v2 leaks sensitive payload content into server and audit logs on malformed requests","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"HashiCorp Vault","year":"2025","cvss_score":4.5,"severity":"medium","kev":false,"impact":"KV v2 leaks sensitive payload content into server and audit logs on malformed requests","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; scrub and re-secure the audit log store","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-4166"],"status":"curated"},{"id":"CVE-2017-5698","cve":"CVE-2017-5698","aliases":["INTEL-SA-00082"],"title":"Intel AMT / ISM / SBT firmware anti-rollback, ME 11.0.25.3001 and 11.0.26.3000: The patched ME firmware does not enforce anti-rollback, so a local administrator can downgrade the ME back to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel AMT / ISM / SBT firmware anti-rollback, ME 11.0.25.3001 and 11.0.26.3000","year":"2017","cvss_score":4.4,"severity":"medium","kev":false,"impact":"The patched ME firmware does not enforce anti-rollback, so a local administrator can downgrade the ME back to an 11.6.x image that is vulnerable to the KEV-listed AMT auth bypass. Operationally this means a node you have already remediated can be silently un-remediated: your fleet inventory says 'patched', the ME says otherwise, and the auth bypass is live again. Because the downgrade lives in the ME region, it survives a host reimage and therefore survives tenant handoff.","attack_vector":"Local root or administrator on the host, using the normal Intel ME firmware update path (HECI/MEI device). No physical access, no ME exploit required - just the vendor's own update tool. Any tenant who has been given real root on a bare-metal node can do this before handing the node back.","remediation":"Flash to an ME build that enforces the rollback floor - again an OEM BIOS/ME bundle from Dell/HPE/Lenovo/Supermicro/Gigabyte/Quanta, with a reboot. The durable control is process, not firmware: read back and attest the actual ME firmware version at node reclaim time rather than trusting an inventory record, and block host-side ME update tooling (MEI device access, Intel MEInfo/FWUpdate binaries) inside tenant images.","references":["https://nvd.nist.gov/vuln/detail/CVE-2017-5698","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00082.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-0117","cve":"CVE-2019-0117","aliases":[],"title":"Intel SGX protected memory subsystem: Insufficient access control in the SGX protected-memory subsystem allows information disclosure from enclave…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX protected memory subsystem","year":"2019","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Insufficient access control in the SGX protected-memory subsystem allows information disclosure from enclave memory to a privileged local user. Part of the long tail of SGX hardware issues that each force a TCB recovery.","attack_vector":"Privileged local access on the host.","remediation":"Microcode/platform firmware update and re-attestation of all enclaves. Where the fix ships in microcode it can be late-loaded at boot; where it ships in the platform BIOS, expect an OEM release and a per-node drain and reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-0117","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00219.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2019-5698","cve":"CVE-2019-5698","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: an index value from the guest is not validated by the vGPU plugin, letting one tenant…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2019","cvss_score":4.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: an index value from the guest is not validated by the vGPU plugin, letting one tenant VM crash the host-side plugin and deny service to every co-tenant on that GPU.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-5698"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-24497","cve":"CVE-2020-24497","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13): Early-generation E810 firmware flaw (an access-control failure) resulting in denial of service on the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Early-generation E810 firmware flaw (an access-control failure) resulting in denial of service on the adapter. Relevant to any E810 adapter still on shipping-era NVM, which is common when NIC firmware was never brought into the patch pipeline.","attack_vector":"Varies by issue - the unauthenticated variant is reachable from the network, the others need privileged host access.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target NVM 1.4.1.13 or later, but jump straight to a current image rather than the minimum fixed version.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24497","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00456.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-24498","cve":"CVE-2020-24498","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13): Early-generation E810 firmware flaw (a buffer overflow reachable by a privileged user) resulting in denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Early-generation E810 firmware flaw (a buffer overflow reachable by a privileged user) resulting in denial of service on the adapter. Relevant to any E810 adapter still on shipping-era NVM, which is common when NIC firmware was never brought into the patch pipeline.","attack_vector":"Varies by issue - the unauthenticated variant is reachable from the network, the others need privileged host access.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target NVM 1.4.1.13 or later, but jump straight to a current image rather than the minimum fixed version.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24498","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00456.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-24500","cve":"CVE-2020-24500","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13): Early-generation E810 firmware flaw (a second buffer overflow reachable by a privileged user) resulting in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM < 1.4.1.13)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Early-generation E810 firmware flaw (a second buffer overflow reachable by a privileged user) resulting in denial of service on the adapter. Relevant to any E810 adapter still on shipping-era NVM, which is common when NIC firmware was never brought into the patch pipeline.","attack_vector":"Varies by issue - the unauthenticated variant is reachable from the network, the others need privileged host access.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules. Target NVM 1.4.1.13 or later, but jump straight to a current image rather than the minimum fixed version.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-24500","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00456.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5973","cve":"CVE-2020-5973","aliases":[],"title":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin): MULTI-TENANT ISOLATION: the vGPU plugin lets guest code reach privileged operations it should not have, which…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software (guest kernel-mode driver + vGPU plugin)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: the vGPU plugin lets guest code reach privileged operations it should not have, which a tenant uses to deny service on the shared GPU. vGPU 8.x before 8.4, 9.x before 9.4, 10.x before 10.3.","attack_vector":"A user inside a guest VM with a vGPU, using the guest kernel-mode driver.","remediation":"Upgrade the vGPU Manager on the host and the vGPU guest driver inside each tenant VM to the fixed release. Host side is a node drain plus reboot; guest side is a per-VM driver install and reboot. Because the guest driver is inside tenant-controlled VMs, in a multi-tenant estate you cannot fully remediate the guest half yourself - the host-side upgrade is the control you own. No VBIOS flash.","references":["https://usn.ubuntu.com/4404-1/","https://usn.ubuntu.com/4404-2/","https://nvd.nist.gov/vuln/detail/CVE-2020-5973"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2020-5982","cve":"CVE-2020-5982","aliases":[],"title":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys): No rate limiting in the kernel-mode scheduler: a local process floods it with requests and starves the GPU.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Windows GPU Display Driver (nvlddmkm.sys)","year":"2020","cvss_score":4.4,"severity":"medium","kev":false,"impact":"No rate limiting in the kernel-mode scheduler: a local process floods it with requests and starves the GPU. On a shared node, one noisy tenant is a denial-of-service tool against everything else on the card.","attack_vector":"Any local user or container with GPU access.","remediation":"Install the fixed Windows GPU Display Driver branch listed in the NVIDIA bulletin. nvlddmkm.sys is a kernel driver: the swap needs a host reboot, so on a Windows GPU node this is a drain-and-reboot, not a live driver reload. No VBIOS or BMC flash involved.","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-5982"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-0197","cve":"CVE-2021-0197","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0197","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00554.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-0198","cve":"CVE-2021-0198","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0198","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00554.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-0199","cve":"CVE-2021-0199","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-0199","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00554.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-1103","cve":"CVE-2021-1103","aliases":[],"title":"NVIDIA vGPU Manager (vGPU plugin): MULTI-TENANT ISOLATION: another guest-reachable NULL dereference in the vGPU plugin causing denial of service…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU Manager (vGPU plugin)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: another guest-reachable NULL dereference in the vGPU plugin causing denial of service to co-tenants. vGPU 12.x before 12.3, 11.x before 11.5, 8.x before 8.8.","attack_vector":"Any unprivileged user inside a guest VM with a vGPU.","remediation":"Upgrade the vGPU Manager on the hypervisor host to the fixed vGPU release. The host component is a kernel module inside the hypervisor, so this is a full node drain: evacuate or power off every tenant VM on the host, upgrade, reboot the host. Guest drivers must be kept within the supported version skew and updated per VM (guest reboot). No VBIOS flash, but expect a maintenance window per host and a matching hypervisor-vendor package (VMware/Citrix/KVM/Nutanix builds ship separately).","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1103"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-27363","cve":"CVE-2021-27363","aliases":[],"title":"Linux iSCSI: Kernel pointer leak - iscsi_transport handle exposed to unprivileged users via sysfs","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Linux iSCSI","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Kernel pointer leak - iscsi_transport handle exposed to unprivileged users via sysfs","attack_vector":"Local","remediation":"Data-plane: kernel patch, batch with a reboot window","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-27363"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2021-33128","cve":"CVE-2021-33128","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2021","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-33128","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00593.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-21894","cve":"CVE-2022-21894","aliases":["BlackLotus"],"title":"Windows Boot Manager: Secure Boot bypass exploited in the wild by the BlackLotus UEFI bootkit","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Windows Boot Manager","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Secure Boot bypass exploited in the wild by the BlackLotus UEFI bootkit; the bootkit survives OS reinstall and disk replacement because it lives in the ESP with a revoked-but-still-trusted bootloader","attack_vector":"Local, high privilege","remediation":"dbx revocation and the phased Microsoft boot-manager revocation rollout. Low CVSS badly understates it: the score reflects the local-privilege precondition, not the below-OS persistence that follows","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-21894"],"status":"curated"},{"id":"CVE-2022-28709","cve":"CVE-2022-28709","aliases":[],"title":"Intel E810 Ethernet controller firmware (NVM): Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel E810 Ethernet controller firmware (NVM)","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Firmware-level flaw in the E810 network controller allowing a privileged local user to cause denial of service on the adapter. Individually minor; collectively they are the reason NIC firmware needs to be in your patch pipeline rather than frozen at whatever the OEM shipped.","attack_vector":"Privileged local access on the host.","remediation":"Fixed in the E810 NVM (adapter firmware) image. Deploy with Intel's NVM Update Utility, which needs a driver reload and a power cycle - not just a warm reboot - for the new image to take effect. Drain the node. Distinct from the ice driver updates: you need both, and they ship on different schedules.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28709","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00593.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-34667","cve":"CVE-2022-34667","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A stack-based buffer overflow on a malformed input file yields limited denial of service and data-integrity…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"A stack-based buffer overflow on a malformed input file yields limited denial of service and data-integrity loss for the invoking user. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5373). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34667","https://github.com/NVIDIA/product-security/tree/main/2022/5373"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:L/A:L","cwe":["CWE-121"]},{"id":"CVE-2022-34673","cve":"CVE-2022-34673","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An out-of-bounds array access in nvidia.ko yields denial of service, information disclosure or data tampering…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An out-of-bounds array access in nvidia.ko yields denial of service, information disclosure or data tampering from an unprivileged local account. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-34673","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2022-42259","cve":"CVE-2022-42259","aliases":[],"title":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko): An integer overflow in nvidia.ko crashes the driver and takes the node's GPUs with it. Everything with a GPU…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Linux kernel module (nvidia.ko)","year":"2022","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An integer overflow in nvidia.ko crashes the driver and takes the node's GPUs with it. Everything with a GPU allocation on the node can reach this, because /dev/nvidia* is handed straight into every GPU container by the container toolkit - there is no additional gate between a tenant workload and the driver ioctl surface.","attack_vector":"Local and unprivileged. The attacker needs only to open /dev/nvidiactl and /dev/nvidia<N> and issue ioctls. On a shared node that is any scheduled tenant pod holding a GPU; no host shell, no root, no CAP_SYS_ADMIN.","remediation":"Move to the fixed datacenter driver branch named in NVIDIA bulletin 5415. Cost: nvidia.ko cannot be replaced while a process holds a GPU, so this is cordon + drain + module reload per node - stop persistence mode and nv-hostengine/DCGM first or the unload fails. With the GPU Operator it is a rolling node upgrade. No VBIOS, BMC or SBIOS flash needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42259","https://github.com/NVIDIA/product-security/tree/main/2022/5415"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-190"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-0193","cve":"CVE-2023-0193","aliases":[],"title":"CUDA Toolkit: DoS / info disclosure (buffer over-read in cuobjdump)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS / info disclosure (buffer over-read in cuobjdump)","attack_vector":"Malicious cubin/model artifact fed to build tooling","remediation":"Bump CUDA Toolkit in base images; rebuild and republish tenant base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0193","https://github.com/NVIDIA/product-security/tree/main/2023/5446"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2023-25520","cve":"CVE-2023-25520","aliases":[],"title":"Jetson TX2 / AGX Xavier (nvbootctrl): DoS via invalid boot config","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Jetson TX2 / AGX Xavier (nvbootctrl)","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS via invalid boot config","attack_vector":"Local privileged attacker","remediation":"Flash JetPack 32.7.4+","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5466/5466.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"]},{"id":"CVE-2023-27593","cve":"CVE-2023-27593","aliases":[],"title":"Cilium: Agent pod hostPath allows writing to /opt/cni/bin, replacing the CNI binary on the host","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Agent pod hostPath allows writing to /opt/cni/bin, replacing the CNI binary on the host","attack_vector":"Attacker with access to a Cilium agent pod","remediation":"Rolling Cilium upgrade; restrict who can exec into kube-system","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-27593"],"status":"curated"},{"id":"CVE-2023-31356","cve":"CVE-2023-31356","aliases":[],"title":"AMD SEV firmware - incomplete memory cleanup (AMD-SB-3003): MULTI-TENANT ISOLATION: Incomplete memory cleanup in the SEV firmware allows corruption of guest private…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - incomplete memory cleanup (AMD-SB-3003)","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Incomplete memory cleanup in the SEV firmware allows corruption of guest private memory. The recurring pattern in this database - firmware that does not scrub or fully release state between uses - applied to the memory of confidential guests, where the whole product claim is that nobody but the guest can touch it.","attack_vector":"Local, privileged, on a host running SEV guests.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31356","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-49100","cve":"CVE-2023-49100","aliases":["TFV-11","SDEI interrupt bind out-of-bounds read"],"title":"Arm Trusted Firmware-A before v2.10, SDEI service (sdei_interrupt_bind SMC handler): An SMC argument from the normal world reaches plat_ic_get_interrupt_type without adequate validation, giving…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-A before v2.10, SDEI service (sdei_interrupt_bind SMC handler)","year":"2023","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An SMC argument from the normal world reaches plat_ic_get_interrupt_type without adequate validation, giving an out-of-bounds read inside EL3. It is a small primitive on its own, but EL3 is the most privileged code on the part - it owns the secure world, the PSCI power state machine and the root of trust - so any memory-safety defect there is a stepping stone toward full platform compromise from a host kernel that is otherwise contained.","attack_vector":"Host kernel or hypervisor code (EL1/EL2) issuing SDEI SMCs. A guest cannot reach it directly unless the hypervisor forwards SDEI, so the realistic path is a tenant who has already got kernel on the host, or a compromised host agent.","remediation":"Upgrade to TF-A v2.10 or later with the TFV-11 fix, via an OEM platform firmware build. Flash + reboot + drain. If SDEI is not used by your platform, having it compiled out of BL31 is the cleaner answer - ask the OEM whether it is even enabled before assuming you are exposed.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-11.html","https://nvd.nist.gov/vuln/detail/CVE-2023-49100"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0110","cve":"CVE-2024-0110","aliases":[],"title":"CUDA Toolkit: OOB write","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"OOB write -> possible code exec in build tooling","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0110","https://github.com/NVIDIA/product-security/tree/main/2024/5564"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:N","cwe":["CWE-787"]},{"id":"CVE-2024-0131","cve":"CVE-2024-0131","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB kernel memory access)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Info disclosure (OOB kernel memory access)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0131","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-805"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0139","cve":"CVE-2024-0139","aliases":[],"title":"Base Command Manager (Linux): Local privesc via insecure temporary file handling","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Base Command Manager (Linux)","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Local privesc via insecure temporary file handling","attack_vector":"Local user on a managed node","remediation":"Patch Base Command Manager on head and compute nodes","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=Base%20Command%20Manager"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:N/I:N/A:H","cwe":["CWE-377"]},{"id":"CVE-2024-21970","cve":"CVE-2024-21970","aliases":[],"title":"AMD Power Management Firmware (SMU) - array index validation: An unvalidated array index in AMD's power management firmware lets a privileged attacker corrupt AGESA…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Power Management Firmware (SMU) - array index validation","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An unvalidated array index in AMD's power management firmware lets a privileged attacker corrupt AGESA memory. The SMU is a live, always-running microcontroller with broad platform reach, so memory corruption there is an integrity problem for the whole node's power and clock management - including, on GPU nodes, the mechanisms that keep accelerators inside their thermal envelope.","attack_vector":"Local, privileged. Reachable through the SMU mailbox interface, which requires root.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string. SMU firmware ships inside the SBIOS/AGESA bundle - there is no separate SMU update channel you can drive yourself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21970","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-28085","cve":"CVE-2024-28085","aliases":[],"title":"util-linux (wall): WallEscape: escape-sequence injection via wall(1) - can spoof a sudo prompt and steal a password on a shared…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"util-linux (wall)","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"WallEscape: escape-sequence injection via wall(1) - can spoof a sudo prompt and steal a password on a shared host","attack_vector":"Local user on a shared login/head node","remediation":"Package update; no reboot. Only matters where multiple tenants share a login node","references":["https://access.redhat.com/security/cve/CVE-2024-28085"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-51741","cve":"CVE-2024-51741","aliases":[],"title":"Redis: Malformed ACL selector triggers a server panic","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2024","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Malformed ACL selector triggers a server panic -> denial of service","attack_vector":"Local","remediation":"Control-plane: upgrade to 7.2.7/7.4.2","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-51741"],"status":"curated"},{"id":"CVE-2025-1118","cve":"CVE-2025-1118","aliases":["GRUB2 2025 batch"],"title":"GRUB2 (dump command lockdown): The dump command was not disabled under Secure Boot lockdown, letting a privileged user read arbitrary memory…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GRUB2 (dump command lockdown)","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"The dump command was not disabled under Secure Boot lockdown, letting a privileged user read arbitrary memory at boot. That includes anything the firmware left in RAM - most usefully, key material and Secure Boot state. Low CVSS, high value as a reconnaissance primitive before a real bypass.","attack_vector":"Local privileged user at the GRUB shell.","remediation":"grub2 package update + reboot. A GRUB password limits access to the shell in the meantime, but does not fix the lockdown gap.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-1118","https://www.openwall.com/lists/oss-security/2025/02/18/3"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-12896","cve":"CVE-2025-12896","aliases":["CVE-2025-12902","Solidigm locked-drive bypass"],"title":"Solidigm DC SSD firmware - unauthorized access to a LOCKED storage device via improper resource management: An attacker with local or physical access gains unauthorized access to a drive that is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Solidigm DC SSD firmware - unauthorized access to a LOCKED storage device via improper resource management","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"An attacker with local or physical access gains unauthorized access to a drive that is in the locked state; the paired issue additionally allows denial of service. 'Locked' is the state your decommission and reclaim runbooks rely on - the drive is supposed to be an inert brick until someone presents the credential. BREAKS TENANT HANDOFF: a drive locked at the end of a tenancy, or locked before shipping back on RMA, can be read anyway. The DoS variant additionally gives a tenant a way to take a drive out from under the host, which on a shared bare-metal node is a self-service outage. These are the most recent public confirmations that the locked-drive guarantee keeps failing on datacenter NVMe - four separate CVEs across two advisory cycles on the same product family.","attack_vector":"A tenant with local (host-level, elevated) access on the bare-metal machine, or anyone with physical access to the drive after it leaves the rack - RMA courier, decommission handler, resale buyer.","remediation":"Firmware update from Solidigm's security page for the affected DC SKUs; drive offline, node drained, Solidigm Storage Tool per SKU. Check your own inventory against the advisory rather than assuming coverage, because Solidigm's DC line spans many families with different firmware trains and older SKUs in this same advisory family have previously been marked no-fix. The structural lesson for an operator: this is now the fourth-plus distinct 'locked drive is not actually locked' finding on datacenter NVMe in two years, so budget for it as a recurring class rather than a one-off patch - assume drive-level locking will fail again and keep a software encryption layer (LUKS/dm-crypt, operator-held key) as the control that actually enforces tenant separation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-12896","https://nvd.nist.gov/vuln/detail/CVE-2025-12902","https://www.solidigm.com/support-page/support-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-21590","cve":"CVE-2025-21590","aliases":[],"title":"Juniper Junos OS kernel: **[KEV]** Improper isolation in the Junos kernel lets a local attacker with shell access inject arbitrary…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS kernel","year":"2025","cvss_score":4.4,"severity":"medium","kev":true,"impact":"**[KEV]** Improper isolation in the Junos kernel lets a local attacker with shell access inject arbitrary code — exploited in the wild to implant persistent backdoors on routers","attack_vector":"Local, shell access","remediation":"Junos upgrade fleet-wide with routing failover; the in-the-wild usage means affected devices need forensic verification, not just patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21590"],"status":"curated"},{"id":"CVE-2025-23247","cve":"CVE-2025-23247","aliases":[],"title":"CUDA Toolkit: DoS via malformed string parsing","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS via malformed string parsing","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23247","https://github.com/NVIDIA/product-security/tree/main/2025/5643"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:L/A:N","cwe":["CWE-130"]},{"id":"CVE-2025-23286","cve":"CVE-2025-23286","aliases":[],"title":"GPU Display Driver: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23286","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-23335","cve":"CVE-2025-23335","aliases":[],"title":"NVIDIA Triton Inference Server: A specific model configuration plus a specific input causes an underflow in the TensorRT backend and kills…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"A specific model configuration plus a specific input causes an underflow in the TensorRT backend and kills the server. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5687. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23335","https://github.com/NVIDIA/product-security/tree/main/2025/5687"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-191"]},{"id":"CVE-2025-23336","cve":"CVE-2025-23336","aliases":[],"title":"NVIDIA Triton Inference Server: Loading a misconfigured model causes a denial of service - relevant where tenants can register their own…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Triton Inference Server","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Loading a misconfigured model causes a denial of service - relevant where tenants can register their own models. On a shared inference tier this is a noisy-neighbour weapon: one tenant's request kills the server process and takes every co-resident model with it, and the GPU sits idle until the pod restarts.","attack_vector":"Network. Anyone who can reach the Triton HTTP or gRPC endpoint. In most clusters that is anything on the pod network; where ingress is loosely scoped it is the internet. No authentication step exists in Triton itself to stop it.","remediation":"Roll to the fixed Triton container image per bulletin 5691. Cost: an ordinary rolling deployment restart - no driver, firmware or node change. Worth pairing with an audit of Triton endpoint exposure, since almost every bug in this component is only interesting because the endpoint is reachable.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23336","https://github.com/NVIDIA/product-security/tree/main/2025/5691"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"]},{"id":"CVE-2025-23345","cve":"CVE-2025-23345","aliases":[],"title":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko): MULTI-TENANT ISOLATION: An out-of-bounds read in the driver's video decoder leaks memory or crashes the node.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - kernel mode layer (Windows nvlddmkm.sys and Linux nvidia.ko)","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds read in the driver's video decoder leaks memory or crashes the node. Relevant to transcode and video-analytics fleets that feed tenant-supplied streams straight into NVDEC. Both the Windows and Linux datacenter drivers are affected, so a mixed fleet needs two separate rollouts.","attack_vector":"Local and unprivileged on either OS. On Linux it is reachable from any GPU container via /dev/nvidia*; on Windows from any session holding a GPU handle.","remediation":"Upgrade both the Linux and the Windows datacenter driver branches listed in bulletin 5703. Cost: Linux needs a drain and nvidia.ko reload per node; Windows needs a reboot per node. Two change windows unless your fleet is homogeneous.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23345","https://github.com/NVIDIA/product-security/tree/main/2025/5703"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-33195","cve":"CVE-2025-33195","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: Unexpected memory buffer operations in SROOT firmware reach data tampering and privilege escalation. These…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Unexpected memory buffer operations in SROOT firmware reach data tampering and privilege escalation. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33195","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:L","cwe":["CWE-119"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33196","cve":"CVE-2025-33196","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: SROOT firmware reuses a resource without clearing it, leaking its previous contents to the next consumer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"SROOT firmware reuses a resource without clearing it, leaking its previous contents to the next consumer - residual-data exposure inside the root of trust. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33196","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:H/I:N/A:N","cwe":["CWE-226"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33221","cve":"CVE-2025-33221","aliases":[],"title":"GPU Display Driver: DoS (input validation)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2025","cvss_score":4.4,"severity":"medium","kev":false,"impact":"DoS (input validation)","attack_vector":"Privileged local user","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33221","https://github.com/NVIDIA/product-security/tree/main/2026/5821"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-20"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-47487","cve":"CVE-2026-47487","aliases":[],"title":"Triton Inference Server: Path traversal in model file operations","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Triton Inference Server","year":"2026","cvss_score":4.4,"severity":"medium","kev":false,"impact":"Path traversal in model file operations","attack_vector":"Authenticated client","remediation":"Upgrade Triton; redeploy","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-47487","https://github.com/NVIDIA/product-security/tree/main/2026/5860"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:L","cwe":["CWE-22"]},{"id":"CVE-2020-13788","cve":"CVE-2020-13788","aliases":[],"title":"Harbor: SSRF: a user who can edit projects scans the Harbor host's intranet","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Harbor","year":"2020","cvss_score":4.3,"severity":"medium","kev":false,"impact":"SSRF: a user who can edit projects scans the Harbor host's intranet","attack_vector":"Authenticated registry user","remediation":"Upgrade Harbor; egress-restrict the registry pod","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-13788"],"status":"curated"},{"id":"CVE-2020-8551","cve":"CVE-2020-8551","aliases":[],"title":"Kubernetes (kubelet): Kubelet API DoS, including via the unauthenticated read-only port","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2020","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Kubelet API DoS, including via the unauthenticated read-only port","attack_vector":"Any pod on the cluster network","remediation":"Rolling kubelet upgrade with node drain; disable the read-only port (10255)","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8551"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2021-3956","cve":"CVE-2021-3956","aliases":[],"title":"Lenovo XClarity Controller (LDAP mode): Read-only authentication bypass when XCC is in LDAP-only authentication mode","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (LDAP mode)","year":"2021","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Read-only authentication bypass when XCC is in LDAP-only authentication mode","attack_vector":"Network","remediation":"XCC firmware update; also affects the common neocloud pattern of centralising BMC auth in LDAP/AD","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-3956"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-39229","cve":"CVE-2022-39229","aliases":[],"title":"Grafana: A user can block another user's login by registering their email address as a username","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Grafana","year":"2022","cvss_score":4.3,"severity":"medium","kev":false,"impact":"A user can block another user's login by registering their email address as a username","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; low severity but a real ops denial vector","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-39229"],"status":"curated"},{"id":"CVE-2022-41354","cve":"CVE-2022-41354","aliases":[],"title":"Argo CD: Unauthenticated attackers can enumerate existing applications","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Argo CD","year":"2022","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Unauthenticated attackers can enumerate existing applications","attack_vector":"Unauthenticated network reaching the Argo CD API","remediation":"Rolling Argo CD upgrade; put the API behind auth","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-41354"],"status":"curated"},{"id":"CVE-2023-31203","cve":"CVE-2023-31203","aliases":[],"title":"OpenVINO Model Server: Input-validation flaw in OpenVINO Model Server reachable without authentication. Same exposure profile as the…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"OpenVINO Model Server","year":"2023","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Input-validation flaw in OpenVINO Model Server reachable without authentication. Same exposure profile as the later Model Server issues - it sits on the inference request path.","attack_vector":"Anything that can reach the serving endpoint.","remediation":"Upgrade to the 2022.3 or later Model Server build shipped with OpenVINO 2023. Container update and rolling restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31203","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00901.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-29040","cve":"CVE-2024-29040","aliases":[],"title":"tpm2-tss (FAPI quote verification): The JSON quote info returned by Fapi_Quote accepts an arbitrary TPM2_GENERATED magic value, so a malicious…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tss (FAPI quote verification)","year":"2024","cvss_score":4.3,"severity":"medium","kev":false,"impact":"The JSON quote info returned by Fapi_Quote accepts an arbitrary TPM2_GENERATED magic value, so a malicious device can hand back a quote that the library accepts but that the TPM never produced. That is attestation forgery: a compromised node convinces the verifier it booted a measured, clean image. For any operator selling verified bare metal or confidential GPU compute, this breaks the assertion the whole product rests on - and it breaks it silently, since a forged quote validates.","attack_vector":"A malicious or compromised endpoint being attested. The attacker is the machine claiming to be healthy, not a third party on the wire.","remediation":"Update tpm2-tss to 4.1.0 or later wherever your attestation verifier runs and restart the service - package-level, no firmware, no reboot. Then re-run attestation across the fleet, because any quote validated by the old library proves nothing. Worth auditing whether your verifier does its own magic-value check rather than trusting the library.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29040","https://github.com/tpm2-software/tpm2-tss/security/advisories"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-8059","cve":"CVE-2024-8059","aliases":["LEN-172051"],"title":"Lenovo XClarity Controller (XCC) - audit log: When an account username is exactly 16 characters, XCC writes the IPMI credentials into its own audit log…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo XClarity Controller (XCC) - audit log","year":"2024","cvss_score":4.3,"severity":"medium","kev":false,"impact":"When an account username is exactly 16 characters, XCC writes the IPMI credentials into its own audit log entries in the clear. The consequence is that anyone allowed to read BMC audit logs - a broader set than anyone allowed to administer the BMC, since logs get shipped to SIEMs, ticketing systems and shared dashboards - picks up working IPMI credentials for the node. Those credentials give power control and boot manipulation. It is a small bug with an awkward blast radius, because log data usually flows to systems with much weaker access control than the BMC itself. Affects a long list of ThinkSystem and ThinkAgile models.","attack_vector":"Anyone with read access to XCC audit logs, or to whatever downstream system those logs are forwarded into. Requires that at least one account on the node has a 16-character username, which is common where naming conventions produce fixed-length service account names.","remediation":"Flash XCC to the per-model version in LEN-172051 - out-of-band, per-node, no host reboot and no drain. Two things the flash does not do, and you must: purge or re-scope any already-collected XCC audit logs sitting in your log pipeline, and rotate the IPMI credentials that were exposed. A quick config-only check meanwhile: look for 16-character usernames across the fleet, since only those trigger the leak.","references":["https://support.lenovo.com/us/en/product_security/LEN-172051","https://nvd.nist.gov/vuln/detail/CVE-2024-8059"],"status":"curated"},{"id":"CVE-2025-25012","cve":"CVE-2025-25012","aliases":[],"title":"Kibana: Open redirect leading to SSRF via a specially crafted URL","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Kibana","year":"2025","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Open redirect leading to SSRF via a specially crafted URL","attack_vector":"Network (remote)","remediation":"Control-plane: Kibana upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25012"],"status":"curated"},{"id":"CVE-2025-33197","cve":"CVE-2025-33197","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A null-pointer dereference in SROOT firmware crashes the platform. These sit in the GB10 root-of-trust chain…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":4.3,"severity":"medium","kev":false,"impact":"A null-pointer dereference in SROOT firmware crashes the platform. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33197","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-476"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-24176","cve":"CVE-2026-24176","aliases":[],"title":"KAI Scheduler: Improper access control in resource allocation (cross-tenant quota abuse)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"KAI Scheduler","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Improper access control in resource allocation (cross-tenant quota abuse)","attack_vector":"Any tenant with cluster API access","remediation":"Upgrade the KAI Scheduler Helm chart; rolling control-plane update, no tenant eviction","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24176","https://github.com/NVIDIA/product-security/tree/main/2026/5818"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:L/A:N","cwe":["CWE-863"]},{"id":"CVE-2026-24232","cve":"CVE-2026-24232","aliases":[],"title":"Transformers4Rec: Code exec via insecure deserialization in the pipeline","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Transformers4Rec","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Code exec via insecure deserialization in the pipeline","attack_vector":"Malicious dataset/model","remediation":"Bump the package; rebuild recsys images","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24232","https://github.com/NVIDIA/product-security/tree/main/2026/5869"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:C/C:N/I:N/A:L","cwe":["CWE-502"]},{"id":"CVE-2026-24241","cve":"CVE-2026-24241","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): Improper authentication in the DLS lets an unauthenticated attacker on an adjacent network read information…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Improper authentication in the DLS lets an unauthenticated attacker on an adjacent network read information from the licensing service.","attack_vector":"Adjacent network, no privileges, no user interaction - the weakest precondition of the DLS set.","remediation":"Update the DLS appliance per bulletin 5789. Cost: appliance restart, no tenant impact.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24241","https://github.com/NVIDIA/product-security/tree/main/2026/5789"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-287"]},{"id":"CVE-2026-39395","cve":"CVE-2026-39395","aliases":[],"title":"cosign / sigstore: verify-blob-attestation reports \"Verified OK\" for malformed or mismatched payloads","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"verify-blob-attestation reports \"Verified OK\" for malformed or mismatched payloads","attack_vector":"Malicious artifact","remediation":"Upgrade cosign to 3.0.6/2.6.3+","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-39395"],"status":"curated"},{"id":"CVE-2026-65920","cve":"CVE-2026-65920","aliases":[],"title":"diffusers (shard file loader): Path traversal in `_get_checkpoint_shard_files`","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"diffusers (shard file loader)","year":"2026","cvss_score":4.3,"severity":"medium","kev":false,"impact":"Path traversal in `_get_checkpoint_shard_files`","attack_vector":"Customer-supplied sharded checkpoint (including safetensors shards)","remediation":"Upgrade past 0.39.0. Note safetensors' safety guarantee covers the tensor payload, not the shard-index filenames","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-65920"],"status":"curated"},{"id":"CVE-2015-7267","cve":"CVE-2015-7267","aliases":["CVE-2015-7268","CVE-2015-7269","Hot Plug attack","Forced Restart attack","Hot Unplug attack"],"title":"Self-encrypting drives in TCG Opal / eDrive mode - Samsung 850 Pro, Samsung PM851, Seagate ST500LT015, ST500LT025 on Dell and Lenovo platforms: Three related bypasses that all exploit the same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Self-encrypting drives in TCG Opal / eDrive mode - Samsung 850 Pro, Samsung PM851, Seagate ST500LT015, ST500LT025 on…","year":"2015","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Three related bypasses that all exploit the same design assumption: an SED stays UNLOCKED once the platform has authenticated it, and only re-locks on a true power cycle. If the drive never sees power drop - hot-plugging a SATA drive while the machine sleeps, forcing a soft reset and booting an alternative OS, or attaching a second SATA connector and an alternate power source before pulling the data cable - the DEK stays live and the attacker reads plaintext with no credential. Included here because this is the CLASS of failure operators keep re-encountering: the newer Solidigm 'locked drive is not locked' CVEs are the same assumption failing again a decade later. BREAKS TENANT HANDOFF for any workflow that assumes a drive is safe because it is locked - locked-while-powered is not locked, and a running or sleeping node is a readable node.","attack_vector":"A physically proximate attacker with access to the running or sleeping machine and its drive cabling - a colo neighbour, a rack tech, or anyone who reaches a node between tenancies while it is still powered. No password required.","remediation":"Mostly a policy and platform-configuration fix rather than a drive flash - the drives behaved as the Opal model specified, so there is no universal firmware patch. Concretely: disable sleep/suspend states on bare-metal hosts so the drive is never in the powered-but-unattended state that makes hot-plug work, require full power cycles rather than soft resets in your reclaim path, and verify that the platform re-authenticates the drive after any reset. Between tenants, do not hand over a node that has merely been rebooted - power it fully down as part of reclaim. The durable answer is the same as everywhere else in this category: software encryption with an operator-held key, so that a drive left unlocked still yields only ciphertext to whoever gets to it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-7267","https://nvd.nist.gov/vuln/detail/CVE-2015-7268","https://nvd.nist.gov/vuln/detail/CVE-2015-7269"],"status":"curated"},{"id":"CVE-2018-12038","cve":"CVE-2018-12038","aliases":["Self-Encrypting Deception","VU#395981"],"title":"Samsung 840 EVO SSD - disk encryption key exposed through wear-levelled NAND and vendor-specific commands: The drive stores key material in ordinary wear-levelled flash. Because the FTL never…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Samsung 840 EVO SSD - disk encryption key exposed through wear-levelled NAND and vendor-specific commands","year":"2018","cvss_score":4.2,"severity":"medium","kev":false,"impact":"The drive stores key material in ordinary wear-levelled flash. Because the FTL never overwrites in place, changing the password leaves the OLD key version sitting in a stale physical block that vendor-specific commands can still reach. So even a drive whose password was rotated - the exact hygiene step an operator performs between customers - still hands over a key that decrypts the previous tenant's data. BREAKS TENANT HANDOFF, and it breaks it in the specific way that defeats the mitigation most operators would reach for: rotating the credential is not enough, because the old key is physically still there.","attack_vector":"Anyone holding the physical drive who can issue vendor-specific (undocumented) ATA commands to the controller - the next bare-metal tenant, an RMA handler, or a buyer of decommissioned hardware. Requires no knowledge of any current or previous password.","remediation":"No firmware fix restores the guarantee, because the exposure is old key material already committed to NAND; flashing new firmware does not scrub blocks the FTL has retired. Physically destroy any 840 EVO that ever held tenant data. Going forward, do not use drive-managed encryption as the tenant boundary - use LUKS/dm-crypt with an operator-held key so that 'erase' is a key-deletion event in your KMS, not a request to the drive. This is also the canonical argument for why a wear-levelled device can never prove sanitization to you: the blocks that matter are the ones the drive has already hidden from the host address space.","references":["https://kb.cert.org/vuls/id/395981","https://nvd.nist.gov/vuln/detail/CVE-2018-12038","https://msrc.microsoft.com/update-guide/vulnerability/ADV180028","https://security.netapp.com/advisory/ntap-20181112-0001/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-0532","cve":"CVE-2022-0532","aliases":[],"title":"CRI-O: \"Safe\" sysctls applied to the host when a pod uses host IPC/network","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"CRI-O","year":"2022","cvss_score":4.2,"severity":"medium","kev":false,"impact":"\"Safe\" sysctls applied to the host when a pod uses host IPC/network","attack_vector":"Cluster user able to create a hostIPC pod","remediation":"Upgrade CRI-O; block hostIPC/hostNetwork for tenants","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-0532"],"status":"curated"},{"id":"CVE-2023-31014","cve":"CVE-2023-31014","aliases":[],"title":"GeForce NOW Android app: Info disclosure / code exec via implicit intent","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GeForce NOW Android app","year":"2023","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Info disclosure / code exec via implicit intent","attack_vector":"Malicious app on the same device","remediation":"Not applicable to server fleets","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5476/5476.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-927"]},{"id":"CVE-2023-31031","cve":"CVE-2023-31031","aliases":[],"title":"DGX A100 SBIOS: Buffer overflow in SBIOS","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"DGX A100 SBIOS","year":"2023","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Buffer overflow in SBIOS","attack_vector":"Local operator","remediation":"Flash SBIOS 1.25+; node power cycle","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31031","https://github.com/NVIDIA/product-security/tree/main/2024/5513"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:L/A:L","cwe":["CWE-122"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-0104","cve":"CVE-2024-0104","aliases":[],"title":"Mellanox OS / MetroX / Onyx / Skyway: Improper access control on switch mgmt","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Mellanox OS / MetroX / Onyx / Skyway","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Improper access control on switch mgmt","attack_vector":"Authenticated switch user","remediation":"Upgrade switch OS image","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0104","https://github.com/NVIDIA/product-security/tree/main/2024/5559"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:L/I:L/A:N","cwe":["CWE-284"]},{"id":"CVE-2024-29902","cve":"CVE-2024-29902","aliases":[],"title":"cosign / sigstore: Remote image with a malicious attachment DoSes the machine running cosign","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Remote image with a malicious attachment DoSes the machine running cosign","attack_vector":"Malicious image","remediation":"Upgrade cosign to 2.2.4+","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29902"],"status":"curated"},{"id":"CVE-2024-29903","cve":"CVE-2024-29903","aliases":[],"title":"cosign / sigstore: Crafted software artifacts DoS the cosign host, affecting all colocated services","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Crafted software artifacts DoS the cosign host, affecting all colocated services","attack_vector":"Malicious artifact","remediation":"Upgrade cosign to 2.2.4+; isolate the verifier","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-29903"],"status":"curated"},{"id":"CVE-2024-45678","cve":"CVE-2024-45678","aliases":["EUCLEAK"],"title":"Infineon cryptographic library (ECDSA) in security microcontrollers: Electromagnetic side channel in Infineon's ECDSA implementation allows secret key extraction from affected…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Infineon cryptographic library (ECDSA) in security microcontrollers","year":"2024","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Electromagnetic side channel in Infineon's ECDSA implementation allows secret key extraction from affected security chips. The library is used well beyond the headline YubiKey case, across Infineon security microcontrollers that also serve as TPMs and platform root-of-trust devices. Where such a chip anchors a fleet's attestation or admin authentication, a cloned credential is indistinguishable from the real one.","attack_vector":"Physical access plus specialised equipment and time with the device. In a datacenter this is not automatically out of scope: a colo cage, an RMA path, a decommissioning contractor, or remote-hands staff all supply that access, and hardware in transit is the classic exposure window.","remediation":"Chip firmware cannot be updated in the affected devices - the fix ships only in new hardware revisions. So the answer is inventory, then replacement or acceptance, plus rotating any credential the affected chip holds. For an operator, the practical control is chain-of-custody: tamper-evident sealing, tracked RMA handling, and never returning a root-of-trust device to a pool without re-provisioning.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45678","https://ninjalab.io/eucleak/"],"status":"curated"},{"id":"CVE-2025-23275","cve":"CVE-2025-23275","aliases":[],"title":"CUDA Toolkit: Code exec potential (buffer overflow)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"Code exec potential (buffer overflow)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23275","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:R/S:U/C:L/I:L/A:L","cwe":["CWE-787"]},{"id":"CVE-2025-23301","cve":"CVE-2025-23301","aliases":[],"title":"NVIDIA HGX / DGX (Hopper and Blackwell) - GPU VBIOS: MULTI-TENANT ISOLATION: A VBIOS misconfiguration lets an attacker set an unsafe GPU debug access level. Debug…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA HGX / DGX (Hopper and Blackwell) - GPU VBIOS","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A VBIOS misconfiguration lets an attacker set an unsafe GPU debug access level. Debug access on a datacenter GPU is the setting that gates whether host-side tooling can inspect GPU internal state - raising it on a shared or confidential-computing node undercuts the platform's own claim that tenant state is opaque. NVIDIA scores the direct impact as denial of service with a changed scope, but the interesting property is that the security posture of the GPU is settable from software.","attack_vector":"Local, low privileges, with high attack complexity. The attacker needs code on the host, not inside a guest - so this matters most where you run tenant containers on bare metal rather than behind a hypervisor.","remediation":"Apply the VBIOS update in NVIDIA bulletin 5674. Cost: a VBIOS flash on an HGX baseboard covers all eight GPUs on the board and requires a full node drain plus power cycle - you cannot flash per-GPU while jobs run. Fold it into your next planned firmware-bundle window rather than treating it as a hotfix; NVIDIA ships HGX firmware as a bundle (VBIOS + NVSwitch + ERoT) and mixing versions is unsupported.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23301","https://github.com/NVIDIA/product-security/tree/main/2025/5674"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:N/I:L/A:L","cwe":["CWE-1244"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23302","cve":"CVE-2025-23302","aliases":[],"title":"NVIDIA HGX / DGX (Hopper and Blackwell) - NVSwitch LS10 firmware: MULTI-TENANT ISOLATION: A misconfiguration of the LS10 NVSwitch lets an attacker set an unsafe debug access…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA HGX / DGX (Hopper and Blackwell) - NVSwitch LS10 firmware","year":"2025","cvss_score":4.2,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A misconfiguration of the LS10 NVSwitch lets an attacker set an unsafe debug access level on the fabric switch itself. The NVSwitch is the component that enforces which GPU can address which peer over NVLink; anything that loosens its debug posture is a question mark over the multi-GPU partitioning that a shared HGX node depends on. Direct scored impact is denial of service with a changed scope.","attack_vector":"Local, low privileges, high complexity, from the host that manages the fabric. Fabric Manager runs here, so the practical prerequisite is code on the baseboard host - not inside a tenant VM.","remediation":"Apply the NVSwitch firmware update from bulletin 5674 as part of the HGX firmware bundle. Cost: full node drain and power cycle across the whole baseboard; NVLink topology is re-established by Fabric Manager on restart, so verify fabric health and NVLink link counts after the flash before returning the node to the pool.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23302","https://github.com/NVIDIA/product-security/tree/main/2025/5674"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:C/C:N/I:L/A:L","cwe":["CWE-1244"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-8561","cve":"CVE-2020-8561","aliases":[],"title":"Kubernetes (kube-apiserver): Admission webhook responses redirect apiserver requests into private networks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Admission webhook responses redirect apiserver requests into private networks; SSRF","attack_vector":"Whoever controls a registered webhook backend","remediation":"Rolling control-plane upgrade; audit webhook configurations","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8561"],"status":"curated"},{"id":"CVE-2021-1088","cve":"CVE-2021-1088","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: debug mechanisms in the GPU's internal microcontroller are reachable with…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: debug mechanisms in the GPU's internal microcontroller are reachable with insufficient access control, leaking information from below the driver. The microcontroller sits underneath every context on the card, so what leaks is not bounded by process or VM. Applies across the datacenter line including DGX-1, DGX-2 and DGX Station A100.","attack_vector":"A user with elevated privileges on the host. In a bare-metal GPU-rental model that is the tenant themselves, which makes 'requires root' a much weaker precondition than it sounds.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1088"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-1105","cve":"CVE-2021-1105","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: debug registers on the GPU's internal microcontroller are readable at runtime…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: debug registers on the GPU's internal microcontroller are readable at runtime, leaking state from beneath the driver. Listed across the DGX-1, DGX-2 and DGX Station A100 lines as well as the GeForce/Quadro/Tesla range.","attack_vector":"A user with elevated privileges on the GPU host - which in bare-metal GPU rental is the tenant.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1105"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-1125","cve":"CVE-2021-1125","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: program data in the GPU's internal microcontroller can be corrupted by a privileged…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: program data in the GPU's internal microcontroller can be corrupted by a privileged user. Corrupting firmware-level state affects the whole card, not one context, and the corruption is not visible to anything running above the driver. Listed for DGX-1, DGX-2 and DGX Station A100 among others.","attack_vector":"A user with elevated privileges on the GPU host - the tenant themselves on rented bare metal.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-1125"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-23219","cve":"CVE-2021-23219","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: loading crafted microcode gives a privileged user access to information the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: loading crafted microcode gives a privileged user access to information the microcontroller is supposed to protect. Same family as the other 2021 microcontroller bugs and listed for DGX-1, DGX-2 and DGX Station A100.","attack_vector":"A user with elevated privileges on the GPU host.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-23219"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-34399","cve":"CVE-2021-34399","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: registers in the GPU's internal microcontroller are not scrubbed, so a privileged…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: registers in the GPU's internal microcontroller are not scrubbed, so a privileged user reads leftover state from whatever ran before them. On a shared or sequentially rented GPU this is residue from the previous tenant's workload.","attack_vector":"A user with elevated privileges on the GPU host - the incoming tenant on rented bare metal.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34399"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-34400","cve":"CVE-2021-34400","aliases":[],"title":"NVIDIA GPU firmware microcontroller (Falcon): MULTI-TENANT ISOLATION: unscrubbed microcontroller memory leaks data to a privileged user. Same residue…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU firmware microcontroller (Falcon)","year":"2021","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: unscrubbed microcontroller memory leaks data to a privileged user. Same residue problem as the unscrubbed-registers issue but with a bigger surface, and it is exactly the failure mode that makes bare-metal GPU handoff between tenants risky: memory that should have been wiped at teardown is still readable by whoever gets the card next.","attack_vector":"A user with elevated privileges on the GPU host, typically the next tenant to be scheduled onto the card.","remediation":"NVIDIA shipped the fix in GPU firmware/microcode delivered with the R470 and R450 driver branches and, on some SKUs, in an updated VBIOS. On most datacenter parts the microcontroller image is loaded by the driver at GPU init, so a driver upgrade plus a node reboot applies it; check the bulletin's product table, because a subset of boards also needs an out-of-band VBIOS/InfoROM update, which is an offline per-node flash with the GPU idle. Either way the node has to be drained. On any fleet that reassigns bare-metal GPU nodes between tenants, pair the driver/firmware update with an explicit GPU reset and memory-scrub step in the reprovisioning pipeline; do not rely on the card clearing itself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-34400"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-28192","cve":"CVE-2022-28192","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A use-after-free in the host vGPU Manager, reachable when host-side resources are…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2022","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the host vGPU Manager, reachable when host-side resources are freed out of sequence, crashes the hypervisor's GPU stack. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Rated hard to exploit because it needs elevated control over freeing host resources, so treat it as a chain component rather than a standalone break.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5353. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-28192","https://github.com/NVIDIA/product-security/tree/main/2022/5353"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:H/UI:N/S:U/C:N/I:N/A:H","cwe":["CWE-416"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2023-52862","cve":"CVE-2023-52862","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":4.1,"severity":"medium","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix null pointer dereference in error message","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52862","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0133","cve":"CVE-2024-0133","aliases":[],"title":"Container Toolkit: Unauthorized empty-file creation on the host","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit","year":"2024","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Unauthorized empty-file creation on the host","attack_vector":"Any tenant with a crafted container image","remediation":"Bump toolkit + restart runtime","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=NVIDIA%20Container%20Toolkit"],"status":"curated","fleet":{"ubiquity":"Universal - same package, same install base","remediation_pain":"`daemon-restart` (1.16.2)","pain_class":"daemon-restart","why_fleet_wide":"Sibling TOCTOU allowing creation of arbitrary files on the host from inside a container; low severity alone but a stepping stone to host compromise on shared nodes"},"cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:N/I:L/A:N","cwe":["CWE-367"]},{"id":"CVE-2024-0134","cve":"CVE-2024-0134","aliases":[],"title":"Container Toolkit / GPU Operator: Unauthorized file creation on the host (data tampering)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Container Toolkit / GPU Operator","year":"2024","cvss_score":4.1,"severity":"medium","kev":false,"impact":"Unauthorized file creation on the host (data tampering)","attack_vector":"Any tenant with a crafted container image","remediation":"Bump toolkit; upgrade GPU Operator Helm chart","references":["https://services.nvd.nist.gov/rest/json/cves/2.0?keywordSearch=NVIDIA%20GPU%20Operator"],"status":"curated","cvss_vector":"CVSS:3.1/AV:N/AC:L/PR:L/UI:R/S:C/C:N/I:L/A:N","cwe":["CWE-61"]},{"id":"CVE-2025-20044","cve":"CVE-2025-20044","aliases":[],"title":"Intel TDX module: MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module","year":"2025","cvss_score":4.1,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: The TDX module is the software that stands between the host/VMM and every confidential VM on the box; a privilege escalation inside it is a break of the boundary that separates a tenant's trust domain from the operator and from other TDs. Specific flaw: improper locking, letting a privileged host user escalate.","attack_vector":"A privileged user on the host - which in the TDX threat model is the adversary the whole design exists to exclude, so 'requires host privilege' is not a mitigating factor here.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20044","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01245.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2018-12037","cve":"CVE-2018-12037","aliases":["Self-Encrypting Deception","VU#395981","Radboud SED research"],"title":"Crucial/Micron MX100, MX200, MX300; Samsung 840 EVO and 850 EVO (ATA-high mode); Samsung T3 and T5 portable SSDs - ATA Security / TCG Opal self-encrypting drives: The drive's data-encryption key…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Crucial/Micron MX100, MX200, MX300; Samsung 840 EVO and 850 EVO (ATA-high mode); Samsung T3 and T5 portable SSDs…","year":"2018","cvss_score":4,"severity":"medium","kev":false,"impact":"The drive's data-encryption key is not derived from the password at all. Set an ATA password or an Opal credential, and the drive still encrypts with a key that sits in the controller independent of that secret - so anyone who can talk to the firmware recovers the plaintext without ever knowing the password. BREAKS TENANT HANDOFF: an operator who 'securely wipes' a node by setting/resetting a drive password, or who relies on the SED to make old data unreadable, has done nothing. The next tenant who gets that bare-metal box, or anyone who receives the drive on an RMA or decommission pallet, reads the prior tenant's training data, checkpoints, SSH keys and cloud credentials in the clear. It also means your 'crypto erase' is not a crypto erase - the key you thought you threw away was never the key protecting the data.","attack_vector":"Anyone who obtains the drive and can issue vendor/debug commands to its controller: the next tenant on the same bare-metal host, an RMA return path, a decommissioned node in a resale channel, or a rack tech with five minutes and a SATA cable. No password, no user credential, and no network access needed - physical or low-level bus access to the drive is the whole requirement.","remediation":"Treat as UNPATCHABLE in practice on the affected SKUs - these are consumer/prosumer drives, several vendors never shipped a firmware fix, and the flaw is in how the key hierarchy was designed rather than a bounds check. The real fix is a policy change: stop trusting hardware encryption as your tenant-separation boundary and layer software encryption you control (LUKS/dm-crypt with a key held in your KMS, or per-tenant filesystem encryption) on top of every local NVMe/SATA device. Then 'crypto erase between tenants' means destroying a key you own, which is verifiable, instead of asking the drive to forget something. Also disable Windows eDrive/BitLocker hardware offload fleet-wide (see the ADV180028 entry). For drives already in service: re-provision with software encryption requires a full re-image and re-encrypt of every affected node, and any drive that held sensitive data before the policy change should be physically destroyed rather than resold, because you cannot retroactively prove it was erased.","references":["https://kb.cert.org/vuls/id/395981","https://nvd.nist.gov/vuln/detail/CVE-2018-12037","https://www.ru.nl/en/research/research-news/radboud-university-researchers-discover-security-flaw-in-ssd-hard-drives","https://www.dell.com/support/kbdoc/en-us/000139235/self-encrypting-drives-vulnerabilities-cve-2018-12037-and-cve-2018-12038-impact-on-dell-emc-server-dell-storage-networking-and-dell-clients","https://msrc.microsoft.com/update-guide/vulnerability/ADV180028"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2021-26400","cve":"CVE-2021-26400","aliases":[],"title":"AMD processors - speculative reordering of loads on shared memory: MULTI-TENANT ISOLATION: AMD processors may speculatively reorder load instructions such that stale data is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - speculative reordering of loads on shared memory","year":"2021","cvss_score":4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: AMD processors may speculatively reorder load instructions such that stale data is observed when several processors operate on shared memory. Where that shared memory spans a trust boundary - a shared page between a guest and the host, or between containers - stale reads become a disclosure channel, and worse, code that relies on memory ordering for its own security checks can be made to see the wrong value.","attack_vector":"Local, requires shared memory between attacker and victim and multiple processors operating on it concurrently.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26400","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-25524","cve":"CVE-2023-25524","aliases":[],"title":"Omniverse Launcher: Access-token disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Omniverse Launcher","year":"2023","cvss_score":4,"severity":"medium","kev":false,"impact":"Access-token disclosure -> user impersonation","attack_vector":"Local user / browser","remediation":"Upgrade Launcher; low relevance to headless DC fleets","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5472/5472.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-598"]},{"id":"CVE-2024-47825","cve":"CVE-2024-47825","aliases":[],"title":"Cilium: Deny rules for prefixes broader than /32 can be ignored","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2024","cvss_score":4,"severity":"medium","kev":false,"impact":"Deny rules for prefixes broader than /32 can be ignored; egress restrictions silently fail","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-47825"],"status":"curated"},{"id":"CVE-2025-32793","cve":"CVE-2025-32793","aliases":[],"title":"Cilium: WireGuard encryption gap in a specific Cilium configuration","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"WireGuard encryption gap in a specific Cilium configuration","attack_vector":"Anyone on the underlay network","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32793"],"status":"curated"},{"id":"CVE-2025-48514","cve":"CVE-2025-48514","aliases":[],"title":"AMD SEV firmware - SEV-ES guest attacking an SNP guest: MULTI-TENANT ISOLATION: Coarse access-control granularity in SEV firmware lets a privileged attacker create a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - SEV-ES guest attacking an SNP guest","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"MULTI-TENANT ISOLATION: Coarse access-control granularity in SEV firmware lets a privileged attacker create a SEV-ES guest positioned to attack an SNP guest, costing the SNP guest confidentiality. The pattern is worth internalising: mixing SEV generations on one host means the weakest guest type in the mix can become the attack platform against the strongest.","attack_vector":"Requires the ability to launch guests with chosen SEV parameters - the operator or a compromised control plane.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands. Interim control: do not co-schedule legacy SEV/SEV-ES guests alongside SEV-SNP guests on the same host. Pinning confidential-tier workloads to SNP-only hosts removes the attack platform entirely and costs you scheduling flexibility rather than a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48514","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-54509","cve":"CVE-2025-54509","aliases":[],"title":"AMD IOMMU register interface - ASP coherency: Improper access control on the IOMMU register interface lets a privileged attacker force non-coherent…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU register interface - ASP coherency","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"Improper access control on the IOMMU register interface lets a privileged attacker force non-coherent accesses by the AMD Secure Processor. Incoherent reads by the security engine mean it can be shown stale or inconsistent data - a subtle way to make the ASP act on something other than what is actually in memory.","attack_vector":"Local, privileged, via the IOMMU register interface.","remediation":"Fixed in AMD reference firmware (AGESA / SEV firmware) and delivered to you only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo, Gigabyte and the ODMs each rebuild and requalify AMD's AGESA drop before it ships. **Expect months, not weeks**: AMD publishes the bulletin, the OEM ships BIOS somewhere between one and six months later, and for platforms past their support window it may never arrive at all. Applying it is a full node power cycle with the host drained - not a driver reload, not a live patch. Track it as a firmware campaign per server SKU, not per kernel version, and verify afterwards by reading back the SMU/PSP firmware version rather than trusting the BIOS revision string.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54509","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-64715","cve":"CVE-2025-64715","aliases":[],"title":"Cilium: Egress policies referencing AWS security group IDs are misapplied","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Cilium","year":"2025","cvss_score":4,"severity":"medium","kev":false,"impact":"Egress policies referencing AWS security group IDs are misapplied","attack_vector":"Any tenant workload","remediation":"Rolling Cilium upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-64715"],"status":"curated"},{"id":"CVE-2021-26387","cve":"CVE-2021-26387","aliases":[],"title":"AMD Secure Processor kernel - DRAM mapping into protected areas (AMD-SB-3003): An access-control gap in the ASP kernel allows DRAM to be mapped into areas the secure processor treats as…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor kernel - DRAM mapping into protected areas (AMD-SB-3003)","year":"2021","cvss_score":3.9,"severity":"low","kev":false,"impact":"An access-control gap in the ASP kernel allows DRAM to be mapped into areas the secure processor treats as protected. Mapping attacker-influenced DRAM into a protected region is how you get the secure processor to operate on data it believes is trustworthy - low score, but it is a building block rather than an endpoint.","attack_vector":"Local, privileged, through the ASP interface.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Marked 'no fix planned' on Naples (EPYC 7001) - on that generation the remediation is hardware retirement, which for 2017-era EPYC in an AI fleet is likely already overdue on performance grounds.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26387","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2021-46772","cve":"CVE-2021-46772","aliases":[],"title":"AGESA Boot Loader (ABL) - SPI ROM header input validation (AMD-SB-3003): The AGESA Boot Loader does not properly validate SPI ROM headers, so malformed header content is acted on…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AGESA Boot Loader (ABL) - SPI ROM header input validation (AMD-SB-3003)","year":"2021","cvss_score":3.9,"severity":"low","kev":false,"impact":"The AGESA Boot Loader does not properly validate SPI ROM headers, so malformed header content is acted on during early boot. Anything that runs before signature enforcement is fully established is disproportionately valuable to an attacker regardless of its CVSS.","attack_vector":"Local, requires SPI ROM write access - root plus flash, a compromised BMC, or supply-chain access.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Enable platform SPI write protection as the compensating control; boot-time parsers cannot be defended from the OS.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-46772","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2022-24735","cve":"CVE-2022-24735","aliases":[],"title":"Redis: Lua environment weakness lets a user inject code that runs with another Redis user's privileges","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2022","cvss_score":3.9,"severity":"low","kev":false,"impact":"Lua environment weakness lets a user inject code that runs with another Redis user's privileges","attack_vector":"Local","remediation":"Control-plane: upgrade + ACL review","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24735"],"status":"curated"},{"id":"CVE-2023-20867","cve":"CVE-2023-20867","aliases":[],"title":"VMware Tools: A fully compromised ESXi host can force VMware Tools to skip host-to-guest authentication - used by UNC3886…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"VMware Tools","year":"2023","cvss_score":3.9,"severity":"low","kev":true,"impact":"A fully compromised ESXi host can force VMware Tools to skip host-to-guest authentication - used by UNC3886 for stealthy guest access [KEV]","attack_vector":"Compromised hypervisor against tenant guests","remediation":"VMware Tools update inside every guest image - tenant-side action a neocloud can only mandate, not perform. Low CVSS, high real-world significance for post-escape persistence","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20867"],"status":"curated"},{"id":"CVE-2025-32004","cve":"CVE-2025-32004","aliases":[],"title":"Intel SGX SDK (Edger8r code generator): The Edger8r tool generates the trusted/untrusted bridge code for enclaves","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX SDK (Edger8r code generator)","year":"2025","cvss_score":3.9,"severity":"low","kev":false,"impact":"The Edger8r tool generates the trusted/untrusted bridge code for enclaves; an input-validation flaw here means the generated bridge itself can be unsafe. Every enclave built with the affected SDK inherits the problem.","attack_vector":"Local authenticated user against an enclave built with the affected generator.","remediation":"Rebuild enclaves with a fixed SGX SDK and re-attest. Vendor-side fix; no operator reboot but also nothing you can patch yourself.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-32004","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01383.html"],"status":"curated"},{"id":"CVE-2023-42776","cve":"CVE-2023-42776","aliases":[],"title":"Intel SGX DCAP for Windows: Input-validation flaw in the Windows DCAP components allowing local information disclosure. Low severity, but…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX DCAP for Windows","year":"2023","cvss_score":3.8,"severity":"low","kev":false,"impact":"Input-validation flaw in the Windows DCAP components allowing local information disclosure. Low severity, but DCAP is the attestation plumbing - anything that touches it deserves a look in a confidential-compute deployment.","attack_vector":"Local authenticated user on a Windows host running DCAP.","remediation":"Update SGX DCAP for Windows to 1.19.100.3 or later. Userspace, service restart.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-42776","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01014.html"],"status":"curated","fleet":{"pain_class":"daemon-restart"}},{"id":"CVE-2024-36348","cve":"CVE-2024-36348","aliases":["TSA","Transient Scheduler Attack"],"title":"AMD processors - speculative inference of control registers despite UMIP: MULTI-TENANT ISOLATION: Part of the Transient Scheduler Attacks batch AMD disclosed in July 2025. A user…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - speculative inference of control registers despite UMIP","year":"2024","cvss_score":3.8,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: Part of the Transient Scheduler Attacks batch AMD disclosed in July 2025. A user process can speculatively infer the contents of control registers even when UMIP - the feature specifically added to stop userspace reading them - is enabled. Control register contents leak kernel configuration and address-layout information, which is the reconnaissance step that makes a subsequent kernel exploit reliable. Low score, real utility to an attacker chaining it.","attack_vector":"Local, unprivileged user process. Part of the TSA family that also covers cross-thread and cross-privilege leakage on affected Zen parts.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. AMD shipped the TSA mitigations in microcode plus a kernel change (VERW-based clearing on transitions) in the July 2025 wave. Patch the whole TSA batch together - the siblings covering store-queue and L1 leakage carry the higher scores. Some TSA mitigations cost measurable performance on context-switch-heavy workloads, so benchmark before you assume the fix is free.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36348","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-36349","cve":"CVE-2024-36349","aliases":["TSA","Transient Scheduler Attack"],"title":"AMD processors - speculative inference of TSC_AUX when reads are disabled: MULTI-TENANT ISOLATION: Sibling of the other Transient Scheduler Attack disclosures. A user process can…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD processors - speculative inference of TSC_AUX when reads are disabled","year":"2024","cvss_score":3.8,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: Sibling of the other Transient Scheduler Attack disclosures. A user process can speculatively infer TSC_AUX even when the platform has disabled that read. TSC_AUX carries the CPU and NUMA node identity, so leaking it tells an attacker exactly where they are running - which is the prerequisite for arranging co-residency with a target tenant and then mounting a cross-core or cross-thread channel against them.","attack_vector":"Local, unprivileged user process on affected AMD parts.","remediation":"Mitigated by AMD microcode plus, on most of these, a kernel-side change - and the durable delivery vehicle is the OEM SBIOS/AGESA package, which carries **one to six months of OEM lag** and needs a drained node and a full power cycle. The linux-firmware amd-ucode blobs get you the microcode sooner via initramfs early-load and a reboot, but AMD does not support late-loading microcode on a running EPYC host, so either way this is reboot-required, not a live patch. Ships with the rest of the July 2025 TSA batch; do not cherry-pick individual CVEs out of it. Benchmark after applying - the TSA mitigations add work on privilege transitions.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36349","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-23134","cve":"CVE-2022-23134","aliases":[],"title":"Zabbix: Some setup.php steps reachable by unauthenticated users","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Zabbix","year":"2022","cvss_score":3.7,"severity":"low","kev":true,"impact":"[KEV] Some setup.php steps reachable by unauthenticated users -> configuration change","attack_vector":"Network (remote)","remediation":"Control-plane: upgrade; restrict frontend network access","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23134"],"status":"curated"},{"id":"CVE-2026-24122","cve":"CVE-2026-24122","aliases":[],"title":"cosign / sigstore: Expired issuing certificate treated as valid during verification","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2026","cvss_score":3.7,"severity":"low","kev":false,"impact":"Expired issuing certificate treated as valid during verification","attack_vector":"Malicious image","remediation":"Upgrade cosign","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-24122"],"status":"curated"},{"id":"CVE-2024-45310","cve":"CVE-2024-45310","aliases":[],"title":"runc: runc can be tricked into creating empty files/directories at arbitrary host locations","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2024","cvss_score":3.6,"severity":"low","kev":false,"impact":"runc can be tricked into creating empty files/directories at arbitrary host locations","attack_vector":"Any tenant workload with control over the pod's mount config","remediation":"Replace runc binary; drain node","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-45310"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2026-0995","cve":"CVE-2026-0995","aliases":["TFV-16","SME TLBI erratum"],"title":"Arm C1-Pro before r1p2; Trusted Firmware-A v2.10 and later on multi-core configurations with the CME complex enabled: SME / SVE / SIMD memory accesses on one core can outlive another core's TLB…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm C1-Pro before r1p2; Trusted Firmware-A v2.10 and later on multi-core configurations with the CME complex enabled","year":"2026","cvss_score":3.6,"severity":"low","kev":false,"impact":"SME / SVE / SIMD memory accesses on one core can outlive another core's TLB invalidate and barrier, so vector loads and stores complete against a translation that was supposed to be dead. The consequence is memory touched outside the expected translation or privilege boundary. Because SME and SVE are exactly what ML kernels use, this is a fault that fires most readily under the workload you actually run, not under a synthetic test.","attack_vector":"Requires code on multiple cores of an affected C1-Pro part - a guest or host process issuing wide vector memory operations while another core performs TLB maintenance. Local only.","remediation":"The TF-A mitigation is heavy: EL3 coordinates a secure-SGI rendezvous across cores using atomic counters, and the OS must call into that SMC interface during affected TLB maintenance. So you need both an OEM firmware build with WORKAROUND_CVE_2026_0995=1 and a patched kernel; neither alone is sufficient. Flash + reboot + drain, plus a kernel roll. Silicon revision r1p2 and later does not need it, so on a fleet refresh this is a spec item to demand from the vendor rather than a patch to carry forever.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-16.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-2295","cve":"CVE-2025-2295","aliases":["GHSA-8522-69fh-w74x"],"title":"EDK II NetworkPkg (IScsiDxe, Ready-To-Transfer PDU handling): A malicious iSCSI target sends a crafted R2T PDU with a bogus offset and length, and the booting firmware…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II NetworkPkg (IScsiDxe, Ready-To-Transfer PDU handling)","year":"2025","cvss_score":3.5,"severity":"low","kev":false,"impact":"A malicious iSCSI target sends a crafted R2T PDU with a bogus offset and length, and the booting firmware obligingly transmits chunks of its own memory back to the target. Low severity on paper, and it is a read-only leak, but what leaks is DXE-phase memory - boot secrets, variable contents, buffer addresses - which is exactly the reconnaissance an attacker needs before firing one of the higher-severity overflow bugs in the same stack.","attack_vector":"An attacker controlling or impersonating the iSCSI target a node boots from. Unauthenticated, pre-OS, from the storage network.","remediation":"OEM BIOS update, flash + reboot per node - but given the low score, do not expect OEMs to ship it urgently or to call it out prominently in release notes. The config workaround is the better first move: turn off the UEFI iSCSI initiator on nodes that boot locally, require mutual CHAP where you do boot from SAN, and segment the storage fabric.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2295","https://github.com/tianocore/edk2/security/advisories/GHSA-8522-69fh-w74x"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-70412","cve":"CVE-2026-70412","aliases":["DSA-2026-348"],"title":"Dell iDRAC9 / iDRAC10 (memory erase, data remanence): Data survives an iDRAC memory erase and stays readable afterwards. The CVSS is low, but the operational…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell iDRAC9 / iDRAC10 (memory erase, data remanence)","year":"2026","cvss_score":3.5,"severity":"low","kev":false,"impact":"Data survives an iDRAC memory erase and stays readable afterwards. The CVSS is low, but the operational meaning for a bare-metal GPU cloud is not: the erase step in your node-reprovisioning pipeline does not actually erase, so a low-privilege user on the next tenancy can read remnants left by the previous one. This is exactly the cross-tenant handoff failure that bare-metal operators promise does not happen, and it fails quietly - the wipe reports success. Affects both the iDRAC9 and iDRAC10 generations.","attack_vector":"A low-privilege account with remote access to the iDRAC - which on a bare-metal cloud can be the next tenant, if your product hands tenants any BMC-adjacent access at all, or anyone reaching the management VLAN.","remediation":"Flash iDRAC9 to 7.20.30.50 or iDRAC10 to 1.20.60.50 or later. Out-of-band, per-node, no host reboot and no job drain. Beyond the flash, treat this as a pipeline bug rather than a node bug: if your reprovisioning runbook relies on the iDRAC erase as the cross-tenant boundary, add an independent verification step, and consider re-checking nodes that were recycled between tenants on unpatched firmware.","references":["https://www.dell.com/support/kbdoc/en-us/000497902/dsa-2026-348-security-update-for-dell-idrac9-and-idrac10-vulnerability","https://nvd.nist.gov/vuln/detail/CVE-2026-70412"],"status":"curated"},{"id":"CVE-2023-2431","cve":"CVE-2023-2431","aliases":[],"title":"Kubernetes (kubelet): Pods with an empty localhost seccomp profile field silently bypass seccomp enforcement","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubelet)","year":"2023","cvss_score":3.4,"severity":"low","kev":false,"impact":"Pods with an empty localhost seccomp profile field silently bypass seccomp enforcement","attack_vector":"Any tenant workload","remediation":"Rolling kubelet upgrade with node drain; add an admission check on seccompProfile","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-2431"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"CVE-2018-20855","cve":"CVE-2018-20855","aliases":["RDMA/mlx5 uninitialized mlx5_ib_create_qp_resp"],"title":"Linux kernel mlx5_ib (create QP response): mlx5_ib_create_qp_resp is never initialized in create_qp_common, so creating a queue pair returns…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel mlx5_ib (create QP response)","year":"2018","cvss_score":3.3,"severity":"low","kev":false,"impact":"mlx5_ib_create_qp_resp is never initialized in create_qp_common, so creating a queue pair returns uninitialized kernel stack memory to the calling process. Low severity on its own, but it is a free kernel-stack read for any tenant with RDMA access, useful for defeating address-space layout randomization before a heavier exploit.","attack_vector":"Local, low-privileged - any user able to create an RDMA queue pair through libibverbs, which on a GPU cluster is every workload using RDMA collectives.","remediation":"Upgrade the host kernel past 4.18.7 or take the distro backport. Any modern kernel already carries this; the value here is checking that legacy long-lived nodes in the fleet are not still on pre-4.18 kernels. Host reboot to apply.","references":["https://ubuntu.com/security/CVE-2018-20855","https://nvd.nist.gov/vuln/detail/CVE-2018-20855"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2019-0174","cve":"CVE-2019-0174","aliases":["RAMBleed","INTEL-SA-00247"],"title":"DDR3 and DDR4 DRAM, including ECC modules; tracked by Intel as a partial-physical-address disclosure issue: Turns Rowhammer from a write primitive into a read primitive. The attacker does not…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DDR3 and DDR4 DRAM, including ECC modules; tracked by Intel as a partial-physical-address disclosure issue","year":"2019","cvss_score":3.3,"severity":"low","kev":false,"impact":"Turns Rowhammer from a write primitive into a read primitive. The attacker does not corrupt the victim's data - they observe whether their own bits flip, which is data-dependent on the neighbouring row's contents, and read the victim's memory out that way. The published result was extraction of a 2048-bit OpenSSH RSA host key from another process at about 0.3 bits per second. Because it is read-only, ECC does not stop it and nothing in the victim's process ever misbehaves, so there is no detection surface at all. The low CVSS is misleading for a multi-tenant operator: the practical claim is cross-tenant key theft with no artefact.","attack_vector":"Unprivileged local code sharing DRAM with the victim - a container, a VM, or a co-scheduled batch job. The attacker needs to get their pages physically adjacent to the victim's, which memory-massaging techniques make reliable on a busy host.","remediation":"No patch. ECC does not help - the attack never needs a flip to survive. Refresh-rate increases raise cost but do not close it. The controls that work: keep untrusted tenants off shared memory controllers, and at the application layer make the secrets worth less by rotating keys aggressively and using memory-hard or constantly-rekeyed representations for long-lived secrets on shared hosts. If you host customer inference on shared CPU memory, assume any long-lived key resident there is readable by a co-tenant given hours.","references":["https://rambleed.com/","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-00247.html","https://nvd.nist.gov/vuln/detail/CVE-2019-0174"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2021-26342","cve":"CVE-2021-26342","aliases":[],"title":"AMD SEV guest VMs - TLB flush after VMCB creation sequence: MULTI-TENANT ISOLATION: The CPU may fail to flush the TLB after a particular sequence involving creation of a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV guest VMs - TLB flush after VMCB creation sequence","year":"2021","cvss_score":3.3,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: The CPU may fail to flush the TLB after a particular sequence involving creation of a new VMCB in SEV guest VMs. Stale translations across a VM boundary mean one guest's memory can be reached through another's cached mapping - the low CVSS reflects the difficulty of arranging the sequence, not the severity of what happens if you do.","attack_vector":"Requires a host able to arrange a specific VMCB creation sequence - hypervisor-privileged.","remediation":"Fixed in AMD reference firmware (AGESA / PSP / SEV firmware) and delivered only as an OEM SBIOS/BIOS package - Dell, HPE, Supermicro, Lenovo and the ODMs each rebuild and requalify AMD's AGESA drop before shipping. **Expect one to six months of OEM lag**, and on end-of-support platforms expect nothing. Applying it is a drain plus full power cycle, not a driver reload. Verify by reading back the PSP/SMU firmware version afterwards rather than trusting the BIOS version string. This sits inside the SEV-SNP trust boundary, so the update moves the platform's reported TCB version: refresh VCEK certificates from AMD's KDS and update any attestation policy your tenants pin, or confidential guest launches will start failing right after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-26342","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2022-23649","cve":"CVE-2022-23649","aliases":[],"title":"cosign / sigstore: Cosign can be tricked into claiming a Rekor transparency-log entry exists when it does not","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2022","cvss_score":3.3,"severity":"low","kev":false,"impact":"Cosign can be tricked into claiming a Rekor transparency-log entry exists when it does not","attack_vector":"Malicious image with a crafted signature","remediation":"Upgrade cosign in the admission and CI paths","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-23649"],"status":"curated"},{"id":"CVE-2022-24736","cve":"CVE-2022-24736","aliases":[],"title":"Redis: Crafted Lua script triggers a NULL pointer dereference","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Redis","year":"2022","cvss_score":3.3,"severity":"low","kev":false,"impact":"Crafted Lua script triggers a NULL pointer dereference -> redis-server crash","attack_vector":"Local","remediation":"Control-plane: rolling Redis upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-24736"],"status":"curated"},{"id":"CVE-2022-42336","cve":"CVE-2022-42336","aliases":[],"title":"Xen on AMD Family 17h / Hygon Family 18h - guest SSBD selection: MULTI-TENANT ISOLATION: Setting Speculative Store Bypass Disable on AMD Family 17h and Hygon Family 18h has…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen on AMD Family 17h / Hygon Family 18h - guest SSBD selection","year":"2022","cvss_score":3.3,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: Setting Speculative Store Bypass Disable on AMD Family 17h and Hygon Family 18h has to be coordinated at the physical core level, and Xen's logic did not do that correctly. The consequence is that a guest which asked for SSBD protection may not actually get it - or may have it silently disabled by a sibling. A tenant hardening itself against Spectre-v4 gets a mitigation that is not in force, which is worse than knowing it is off.","attack_vector":"Cross-guest speculative execution between VMs sharing a physical core.","remediation":"Fixed in Xen (XSA-431). Hypervisor update plus host reboot. Verify per-guest that SSBD is genuinely active afterwards rather than trusting the requested setting.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-42336","https://xenbits.xen.org/xsa/advisory-431.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-0196","cve":"CVE-2023-0196","aliases":[],"title":"CUDA Toolkit: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref)","attack_vector":"Malicious ELF/cubin artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0196","https://github.com/NVIDIA/product-security/tree/main/2023/5446"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"]},{"id":"CVE-2023-20519","cve":"CVE-2023-20519","aliases":[],"title":"AMD SEV-SNP guest context page - use-after-free enabling migration-agent masquerade (AMD-SB-3002): MULTI-TENANT ISOLATION: A use-after-free in the SNP guest context page lets a malicious…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP guest context page - use-after-free enabling migration-agent masquerade (AMD-SB-3002)","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: A use-after-free in the SNP guest context page lets a malicious hypervisor masquerade as the guest's migration agent. The guest then negotiates its migration with the attacker instead of a legitimate MA - which is to say it hands over the state that memory encryption existed to protect, voluntarily, to the party it was protecting itself from.","attack_vector":"Malicious or compromised hypervisor. Only exercised where SEV-SNP live migration is enabled.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. If you do not offer live migration for confidential VMs, the path is unreachable and this can wait for the next firmware wave. If you do, disable it until the fleet is patched - that is a scheduler policy change, not a maintenance window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20519","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3002.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2023-25510","cve":"CVE-2023-25510","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A null-pointer dereference on a malformed binary crashes the tool. The realistic exposure is your build and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"A null-pointer dereference on a malformed binary crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25510","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"]},{"id":"CVE-2023-25511","cve":"CVE-2023-25511","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A division-by-zero on crafted input crashes the tool. The realistic exposure is your build and profiling…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"A division-by-zero on crafted input crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5456). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-25511","https://github.com/NVIDIA/product-security/tree/main/2023/5456"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-369"]},{"id":"CVE-2023-25523","cve":"CVE-2023-25523","aliases":[],"title":"CUDA Toolkit (nvdisasm): DoS (null deref via malformed ELF)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit (nvdisasm)","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref via malformed ELF)","attack_vector":"Malicious model/binary artifact","remediation":"Upgrade to CUDA Toolkit 12.2+; rebuild base images","references":["https://raw.githubusercontent.com/NVIDIA/product-security/main/2023/5469/5469.md"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"]},{"id":"CVE-2023-31306","cve":"CVE-2023-31306","aliases":[],"title":"AMD graphics driver - dynamic power management (DPM) array index validation: An unvalidated array index in the driver's dynamic power management functions produces an out-of-bounds…","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"AMD graphics driver - dynamic power management (DPM) array index validation","year":"2023","cvss_score":3.3,"severity":"low","kev":false,"impact":"An unvalidated array index in the driver's dynamic power management functions produces an out-of-bounds access. DPM controls clocks and power states; on Instinct parts that is the machinery keeping accelerators inside their power and thermal envelope, so corruption here is worth more attention than the 3.3 score suggests even though the direct security impact is limited.","attack_vector":"Local, requires the ability to pass malformed arguments to DPM functions.","remediation":"Update the AMD graphics driver and reload or reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31306","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-0072","cve":"CVE-2024-0072","aliases":[],"title":"CUDA Toolkit: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0072","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"]},{"id":"CVE-2024-0076","cve":"CVE-2024-0076","aliases":[],"title":"CUDA Toolkit: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0076","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-0102","cve":"CVE-2024-0102","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0102","https://github.com/NVIDIA/product-security/tree/main/2024/5548"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-0109","cve":"CVE-2024-0109","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious binary artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0109","https://github.com/NVIDIA/product-security/tree/main/2024/5564"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-0123","cve":"CVE-2024-0123","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: Improper input validation on a malicious ELF crashes the disassembler. The realistic exposure is your build…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Improper input validation on a malicious ELF crashes the disassembler. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5577). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0123","https://github.com/NVIDIA/product-security/tree/main/2024/5577"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-1285"]},{"id":"CVE-2024-0124","cve":"CVE-2024-0124","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: A use-after-free on a malformed ELF causes a crash and potentially worse depending on heap state. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"A use-after-free on a malformed ELF causes a crash and potentially worse depending on heap state. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5577). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0124","https://github.com/NVIDIA/product-security/tree/main/2024/5577"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-416"]},{"id":"CVE-2024-0125","cve":"CVE-2024-0125","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: A null-pointer dereference on a malformed ELF crashes the disassembler. The realistic exposure is your build…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"A null-pointer dereference on a malformed ELF crashes the disassembler. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5577). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0125","https://github.com/NVIDIA/product-security/tree/main/2024/5577"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"]},{"id":"CVE-2024-0149","cve":"CVE-2024-0149","aliases":[],"title":"GPU Display Driver: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Any tenant with a container","remediation":"Driver upgrade; rolling reboot","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0149","https://github.com/NVIDIA/product-security/tree/main/2025/5614"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-125"],"fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53870","cve":"CVE-2024-53870","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53870","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-53871","cve":"CVE-2024-53871","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53871","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-53872","cve":"CVE-2024-53872","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53872","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-53873","cve":"CVE-2024-53873","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53873","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-53874","cve":"CVE-2024-53874","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53874","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-53875","cve":"CVE-2024-53875","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53875","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-53876","cve":"CVE-2024-53876","aliases":[],"title":"CUDA Toolkit: Info disclosure (OOB read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (OOB read)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53876","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2024-53877","cve":"CVE-2024-53877","aliases":[],"title":"CUDA Toolkit: DoS (null deref)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":3.3,"severity":"low","kev":false,"impact":"DoS (null deref)","attack_vector":"Malicious cubin/model artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53877","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"]},{"id":"CVE-2025-20613","cve":"CVE-2025-20613","aliases":[],"title":"Intel TDX firmware (PRNG seeding): MULTI-TENANT ISOLATION: A predictable seed in the TDX firmware's pseudo-random number generator. Predictable…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX firmware (PRNG seeding)","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: A predictable seed in the TDX firmware's pseudo-random number generator. Predictable randomness inside a confidential-compute TCB undermines whatever the module derived from it - key material, nonces, address-space randomisation inside the boundary - so the low CVSS understates the structural concern.","attack_vector":"An authenticated user on the host.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-20613","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01312.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23248","cve":"CVE-2025-23248","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: An out-of-bounds read on a malformed ELF crashes the tool. The realistic exposure is your build and profiling…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An out-of-bounds read on a malformed ELF crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23248","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2025-23255","cve":"CVE-2025-23255","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: An out-of-bounds read on a malformed ELF crashes the tool. The realistic exposure is your build and profiling…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An out-of-bounds read on a malformed ELF crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23255","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2025-23271","cve":"CVE-2025-23271","aliases":[],"title":"CUDA Toolkit: Info disclosure (buffer over-read)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Info disclosure (buffer over-read)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23271","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2025-23287","cve":"CVE-2025-23287","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): An attacker with local access reads sensitive system-level information through the Windows display driver…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An attacker with local access reads sensitive system-level information through the Windows display driver - useful for fingerprinting the host before a heavier attack. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5670. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23287","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-497"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-23288","cve":"CVE-2025-23288","aliases":[],"title":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys): Local unprivileged access to the Windows display driver exposes sensitive system information. Only matters to…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPU Display Driver - Windows kernel mode layer (nvlddmkm.sys)","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Local unprivileged access to the Windows display driver exposes sensitive system information. Only matters to you if you run Windows GPU nodes - VDI/DaaS session hosts, cloud-gaming fleets, Windows render or CAE farms.","attack_vector":"Local and unprivileged on a Windows GPU node, through the driver's private IOCTL / DxgkDdiEscape path. Any interactive or RDP/Citrix session with a GPU handle can call it, so on a multi-session VDI host every logged-in user is in range.","remediation":"Install the fixed Windows display driver from bulletin 5670. Cost: a Windows display-driver replacement reboots the node, so drain sessions first. Linux-only fleets can skip this entirely.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23288","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-497"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-23308","cve":"CVE-2025-23308","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: A heap-based buffer overflow on a malicious ELF gives arbitrary code execution at the privilege level of…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A heap-based buffer overflow on a malicious ELF gives arbitrary code execution at the privilege level of whoever ran nvdisasm - the most serious of the CUDA parser bugs, and CI service accounts are usually not low-privilege. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23308","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:N/A:N","cwe":["CWE-122"]},{"id":"CVE-2025-23338","cve":"CVE-2025-23338","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: An out-of-bounds write on a malicious ELF crashes the tool and corrupts heap state. The realistic exposure is…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"An out-of-bounds write on a malicious ELF crashes the tool and corrupts heap state. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23338","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-129"]},{"id":"CVE-2025-23339","cve":"CVE-2025-23339","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A stack-based buffer overflow on a malicious ELF gives arbitrary code execution as the invoking user. The…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A stack-based buffer overflow on a malicious ELF gives arbitrary code execution as the invoking user. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23339","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:L/I:N/A:N","cwe":["CWE-121"]},{"id":"CVE-2025-23340","cve":"CVE-2025-23340","aliases":[],"title":"NVIDIA CUDA Toolkit - nvdisasm: Another out-of-bounds read on a malformed ELF, fixed alongside the rest. The realistic exposure is your build…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - nvdisasm","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Another out-of-bounds read on a malformed ELF, fixed alongside the rest. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run nvdisasm over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23340","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-125"]},{"id":"CVE-2025-23346","cve":"CVE-2025-23346","aliases":[],"title":"NVIDIA CUDA Toolkit - cuobjdump: A null-pointer dereference from an unprivileged user crashes the tool. The realistic exposure is your build…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA CUDA Toolkit - cuobjdump","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A null-pointer dereference from an unprivileged user crashes the tool. The realistic exposure is your build and profiling pipeline, not your runtime fleet: anything that automatically disassembles third-party fatbins, vendor kernels or model artifacts is running this parser on attacker-influenced input.","attack_vector":"Local, and requires a user or an automated job to run cuobjdump over an attacker-supplied file. CI jobs that inspect third-party CUDA binaries are the usual path.","remediation":"Update the CUDA Toolkit package (bulletin 5661). Cost: effectively zero - userspace SDK only, no driver reload, no node drain, no running-job impact. Rebuild build/CI images and move on.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23346","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-476"]},{"id":"CVE-2025-33198","cve":"CVE-2025-33198","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A second resource-reuse path in SROOT firmware leaks residual data. These sit in the GB10 root-of-trust…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"A second resource-reuse path in SROOT firmware leaks residual data. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33198","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-226"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-54410","cve":"CVE-2025-54410","aliases":[],"title":"Docker / moby: Related firewalld handling defect affecting Moby port exposure","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2025","cvss_score":3.3,"severity":"low","kev":false,"impact":"Related firewalld handling defect affecting Moby port exposure","attack_vector":"Unauthenticated network","remediation":"Upgrade Docker Engine","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-54410"],"status":"curated"},{"id":"CVE-2026-41579","cve":"CVE-2026-41579","aliases":[],"title":"runc: setupPtmx/setupDev rootfs setup flaw during container rootfs construction","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"runc","year":"2026","cvss_score":3.3,"severity":"low","kev":false,"impact":"setupPtmx/setupDev rootfs setup flaw during container rootfs construction","attack_vector":"Any tenant workload with crafted rootfs","remediation":"Replace runc binary at next maintenance window; low severity, batch with other node work","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-41579"],"status":"curated"},{"id":"CVE-2023-20573","cve":"CVE-2023-20573","aliases":[],"title":"AMD SEV-SNP - debug exception delivery to guests: A privileged attacker can suppress delivery of debug exceptions to SEV-SNP guests. The guest does not get…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - debug exception delivery to guests","year":"2023","cvss_score":3.2,"severity":"low","kev":false,"impact":"A privileged attacker can suppress delivery of debug exceptions to SEV-SNP guests. The guest does not get debug information it expects, which is mostly an availability and observability problem - but for a guest that relies on debug exceptions as part of a self-protection or integrity-checking scheme, silently swallowing them removes that check.","attack_vector":"Privileged host attacker against a confidential guest.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. Low priority relative to the RMP-bypass and microcode issues; batch it into the next BIOS wave.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20573","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-21977","cve":"CVE-2024-21977","aliases":[],"title":"AMD CPU microcode - RDRAND entropy after patch load: MULTI-TENANT ISOLATION: Incomplete cleanup after loading a microcode patch degrades the entropy of RDRAND, so…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU microcode - RDRAND entropy after patch load","year":"2024","cvss_score":3.2,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: Incomplete cleanup after loading a microcode patch degrades the entropy of RDRAND, so a privileged attacker can weaken the randomness that SEV-SNP guests draw on. Guests that seed keys or nonces from RDRAND get predictable material, which quietly breaks their crypto without breaking anything visible. The score is low; the failure mode - silent, undetectable, affects key generation - is not.","attack_vector":"Local, privileged attacker who can trigger microcode patch loading. Effect lands on SEV-SNP guests on that host.","remediation":"Fixed by an AMD microcode patch. Two delivery routes, and the difference matters: the linux-firmware amd-ucode blobs load early at boot (initramfs) and need only a reboot, while the durable fix is the microcode embedded in the OEM SBIOS/AGESA package, which carries the usual one-to-six-month OEM lag and a full power cycle. **For confidential computing you need the SBIOS route**: microcode late-loaded by the OS is not part of what SEV-SNP attests, so a guest checking the attestation report cannot tell the fix is present. AMD does not support late-loading microcode on a running EPYC host - treat this as reboot-required. After patching, expect the reported TCB version to change and plan the VCEK certificate refresh accordingly. Guests that generated long-lived keys on an affected host should rotate them; patching stops future weak output but does nothing about keys already derived from degraded entropy.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-21977","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2024-36331","cve":"CVE-2024-36331","aliases":[],"title":"AMD CPU cache initialization - SEV-SNP guest memory integrity: MULTI-TENANT ISOLATION: Improper initialization of CPU cache memory lets a hypervisor-privileged attacker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD CPU cache initialization - SEV-SNP guest memory integrity","year":"2024","cvss_score":3.2,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: Improper initialization of CPU cache memory lets a hypervisor-privileged attacker overwrite SEV-SNP guest memory, costing guest data integrity. Cache-state manipulation is a recurring theme in SEV attacks (the CacheWarp research works the same seam) because the encryption protects DRAM, not what the cache does on the way there.","attack_vector":"Hypervisor-privileged attacker.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36331","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33199","cve":"CVE-2025-33199","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: Incorrect control-flow behaviour in SROOT firmware permits data tampering. These sit in the GB10…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":3.2,"severity":"low","kev":false,"impact":"Incorrect control-flow behaviour in SROOT firmware permits data tampering. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33199","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:C/C:N/I:L/A:N","cwe":["CWE-670"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2021-25740","cve":"CVE-2021-25740","aliases":[],"title":"Kubernetes: Endpoint/EndpointSlice confused-deputy lets users reach networks they should not","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2021","cvss_score":3.1,"severity":"low","kev":false,"impact":"Endpoint/EndpointSlice confused-deputy lets users reach networks they should not","attack_vector":"Cluster user with namespace access","remediation":"No complete upstream fix; restrict Endpoint creation via admission policy","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25740"],"status":"curated"},{"id":"CVE-2021-32718","cve":"CVE-2021-32718","aliases":[],"title":"RabbitMQ: Unsanitized username rendered in the management UI","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"RabbitMQ","year":"2021","cvss_score":3.1,"severity":"low","kev":false,"impact":"Unsanitized username rendered in the management UI -> stored XSS against an admin","attack_vector":"Network (remote)","remediation":"Control-plane: management plugin upgrade; keep the UI off the public internet","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-32718"],"status":"curated"},{"id":"CVE-2023-32082","cve":"CVE-2023-32082","aliases":[],"title":"etcd: LeaseTimeToLive exposes key names to a user without read permission on those keys","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"etcd","year":"2023","cvss_score":3.1,"severity":"low","kev":false,"impact":"LeaseTimeToLive exposes key names to a user without read permission on those keys","attack_vector":"Network (remote)","remediation":"Control-plane: etcd upgrade; audit RBAC on the cluster store","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-32082"],"status":"curated"},{"id":"CVE-2023-46737","cve":"CVE-2023-46737","aliases":[],"title":"cosign / sigstore: Attacker-controlled registry returns unbounded attestations, DoSing the verifier","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"cosign / sigstore","year":"2023","cvss_score":3.1,"severity":"low","kev":false,"impact":"Attacker-controlled registry returns unbounded attestations, DoSing the verifier","attack_vector":"Malicious registry, e.g. a tenant-specified image source","remediation":"Upgrade cosign; restrict which registries the admission controller will contact","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-46737"],"status":"curated"},{"id":"CVE-2024-51744","cve":"CVE-2024-51744","aliases":[],"title":"Prometheus / Thanos (golang-jwt): Unclear ParseWithClaims error behavior","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"Prometheus / Thanos (golang-jwt)","year":"2024","cvss_score":3.1,"severity":"low","kev":false,"impact":"Unclear ParseWithClaims error behavior -> callers may accept an expired-and-invalid token","attack_vector":"Network (remote)","remediation":"Control-plane: dependency bump and rebuild of Go control-plane services","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-51744"],"status":"curated"},{"id":"CVE-2024-7598","cve":"CVE-2024-7598","aliases":[],"title":"Kubernetes (kube-apiserver): NetworkPolicy is not applied during a race in namespace termination, so pods briefly run unrestricted","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2024","cvss_score":3.1,"severity":"low","kev":false,"impact":"NetworkPolicy is not applied during a race in namespace termination, so pods briefly run unrestricted","attack_vector":"Cluster user who can create and delete namespaces","remediation":"Rolling control-plane upgrade; treat namespace churn as a policy-gap window","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2026-15605","cve":"CVE-2026-15605","aliases":[],"title":"wandb SDK (`ArtifactManifestEntry.download`): Hash-handling weakness in artifact download integrity","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"wandb SDK (`ArtifactManifestEntry.download`)","year":"2026","cvss_score":3.1,"severity":"low","kev":false,"impact":"Hash-handling weakness in artifact download integrity","attack_vector":"Poisoned artifact in the registry","remediation":"Upgrade the SDK in base images; weakens artifact-integrity guarantees for model supply chain","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15605"],"status":"curated"},{"id":"CVE-2026-24513","cve":"CVE-2026-24513","aliases":[],"title":"ingress-nginx: auth-url protection bypass","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"ingress-nginx","year":"2026","cvss_score":3.1,"severity":"low","kev":false,"impact":"auth-url protection bypass; authentication in front of a tenant service can be skipped","attack_vector":"Unauthenticated network","remediation":"Rolling controller upgrade","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2021-25743","cve":"CVE-2021-25743","aliases":[],"title":"Kubernetes (kubectl): kubectl does not neutralise ANSI escape sequences in output","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kubectl)","year":"2021","cvss_score":3,"severity":"low","kev":false,"impact":"kubectl does not neutralise ANSI escape sequences in output; terminal injection on the operator's machine","attack_vector":"Any tenant who can set an object field the operator will print","remediation":"Upgrade kubectl on operator machines","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25743"],"status":"curated"},{"id":"CVE-2021-41190","cve":"CVE-2021-41190","aliases":[],"title":"OCI Distribution Spec: Content-Type alone determines manifest type, so a manifest can be interpreted differently by different clients","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"OCI Distribution Spec","year":"2021","cvss_score":3,"severity":"low","kev":false,"impact":"Content-Type alone determines manifest type, so a manifest can be interpreted differently by different clients; signature and policy confusion","attack_vector":"Malicious image in any registry","remediation":"Upgrade registry and client tooling; pin by digest rather than tag","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41190"],"status":"curated"},{"id":"CVE-2021-41089","cve":"CVE-2021-41089","aliases":[],"title":"Docker / moby: `docker cp` into a crafted container changes Unix permissions of existing host files","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Docker / moby","year":"2021","cvss_score":2.8,"severity":"low","kev":false,"impact":"`docker cp` into a crafted container changes Unix permissions of existing host files","attack_vector":"Any tenant workload on a node where operators run docker cp","remediation":"Upgrade Docker Engine","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-41089"],"status":"curated"},{"id":"CVE-2023-31028","cve":"CVE-2023-31028","aliases":[],"title":"NVIDIA nvJPEG2000 library: Improper input validation on a crafted JPEG2000 file causes a partial denial of service in the decoding…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA nvJPEG2000 library","year":"2023","cvss_score":2.8,"severity":"low","kev":false,"impact":"Improper input validation on a crafted JPEG2000 file causes a partial denial of service in the decoding library. Low severity on its own, but nvJPEG2000 sits inside DALI and medical/geospatial imaging pipelines that ingest customer files by design, so the untrusted-input assumption is real.","attack_vector":"Local, requires the library to decode an attacker-supplied image. Any data-loading pipeline that accepts tenant or customer imagery is the delivery path.","remediation":"Update the nvJPEG2000 library per bulletin 5517 and rebuild the images that link it. Cost: package update and job restart only; no driver or firmware change.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-31028","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-20"]},{"id":"CVE-2024-0080","cve":"CVE-2024-0080","aliases":[],"title":"nvTIFF library: DoS via malformed image","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"nvTIFF library","year":"2024","cvss_score":2.8,"severity":"low","kev":false,"impact":"DoS via malformed image","attack_vector":"Malicious dataset input","remediation":"Bump nvTIFF in image-processing images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0080","https://github.com/NVIDIA/product-security/tree/main/2024/5517"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-20"]},{"id":"CVE-2024-53878","cve":"CVE-2024-53878","aliases":[],"title":"CUDA Toolkit: DoS (improper type checking)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":2.8,"severity":"low","kev":false,"impact":"DoS (improper type checking)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53878","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-1284"]},{"id":"CVE-2024-53879","cve":"CVE-2024-53879","aliases":[],"title":"CUDA Toolkit: DoS (improper type checking)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2024","cvss_score":2.8,"severity":"low","kev":false,"impact":"DoS (improper type checking)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild base images","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-53879","https://github.com/NVIDIA/product-security/tree/main/2025/5594"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:L/UI:R/S:U/C:N/I:N/A:L","cwe":["CWE-1284"]},{"id":"CVE-2021-25737","cve":"CVE-2021-25737","aliases":[],"title":"Kubernetes: Endpoint IPs can redirect pod traffic to private node networks","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes","year":"2021","cvss_score":2.7,"severity":"low","kev":false,"impact":"Endpoint IPs can redirect pod traffic to private node networks","attack_vector":"Cluster user able to create Endpoints","remediation":"Rolling control-plane upgrade","references":["https://nvd.nist.gov/vuln/detail/CVE-2021-25737"],"status":"curated"},{"id":"CVE-2024-3177","cve":"CVE-2024-3177","aliases":[],"title":"Kubernetes (kube-apiserver): Init/ephemeral container envFrom bypasses the ServiceAccount mountable-secrets policy","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2024","cvss_score":2.7,"severity":"low","kev":false,"impact":"Init/ephemeral container envFrom bypasses the ServiceAccount mountable-secrets policy","attack_vector":"Cluster user with namespace access","remediation":"Rolling control-plane upgrade; no GPU drain","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated","fleet":{"ubiquity":"Universal component, low impact - every cluster runs the ServiceAccount admission plugin","remediation_pain":"`daemon-restart` - control-plane-only upgrade, no GPU node drain needed","pain_class":"node-drain","why_fleet_wide":"`envFrom` bypasses the mountable-secrets restriction, leaking secrets across a namespace boundary; control-plane-scoped and low severity, so not a fleet emergency"}},{"id":"CVE-2025-4563","cve":"CVE-2025-4563","aliases":[],"title":"Kubernetes (kube-apiserver): Nodes can bypass DRA authorization checks. Directly relevant: DRA is how GPUs are allocated in modern clusters","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2025","cvss_score":2.7,"severity":"low","kev":false,"impact":"Nodes can bypass DRA authorization checks. Directly relevant: DRA is how GPUs are allocated in modern clusters","attack_vector":"A compromised node / kubelet credential","remediation":"Rolling control-plane upgrade; no GPU drain, but re-audit DRA ResourceClaim allocations","references":["https://kubernetes.io/docs/reference/issues-security/official-cve-feed/"],"status":"curated"},{"id":"CVE-2025-25183","cve":"CVE-2025-25183","aliases":[],"title":"vLLM (prefix cache hash collisions): Crafted prompts collide hashes","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (prefix cache hash collisions)","year":"2025","cvss_score":2.6,"severity":"low","kev":false,"impact":"Crafted prompts collide hashes → cache reuse across requests","attack_vector":"Co-tenant on a shared instance","remediation":"Same as above — architectural, not patchable away","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-25183"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2025-46570","cve":"CVE-2025-46570","aliases":[],"title":"vLLM (prefix cache): Prefix-cache timing side channel leaks other tenants' prompts","layer":"ai-serving","layer_name":"AI/ML frameworks & serving","component":"vLLM (prefix cache)","year":"2025","cvss_score":2.6,"severity":"low","kev":false,"impact":"Prefix-cache timing side channel leaks other tenants' prompts","attack_vector":"Co-tenant issuing timed prompts against a shared serving instance","remediation":"No clean fix while prefix caching is shared. Do not share a vLLM instance across tenants — the cache is a cross-tenant channel","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-46570"],"status":"curated"},{"id":"CVE-2023-20581","cve":"CVE-2023-20581","aliases":[],"title":"AMD IOMMU access control - SEV-SNP RMP check bypass (AMD-SB-3009): MULTI-TENANT ISOLATION: An IOMMU access-control flaw lets a privileged attacker bypass reverse-map table…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD IOMMU access control - SEV-SNP RMP check bypass (AMD-SB-3009)","year":"2023","cvss_score":2.5,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: An IOMMU access-control flaw lets a privileged attacker bypass reverse-map table checks, undermining SEV-SNP guest memory protection. One of a family of IOMMU-mediated RMP bypasses - the RMP guards CPU accesses well, and the recurring weakness is device accesses arriving through the IOMMU's error and edge-case paths.","attack_vector":"Privileged attacker with a compromised hypervisor, driving DMA.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands. Patch as a set with the other IOMMU/RMP bypasses rather than individually - they share firmware releases and any one of them reopens the class.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20581","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3009.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2024-27457","cve":"CVE-2024-27457","aliases":[],"title":"Intel TDX module firmware: Missing check for an exceptional condition in the TDX module allows a privileged user to reach information…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX module firmware","year":"2024","cvss_score":2.5,"severity":"low","kev":false,"impact":"Missing check for an exceptional condition in the TDX module allows a privileged user to reach information disclosure. Scored very low, but it is inside the TDX TCB, so it still triggers a module SVN bump and therefore a re-attestation cycle.","attack_vector":"Privileged host user.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-27457","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01099.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-23273","cve":"CVE-2025-23273","aliases":[],"title":"CUDA Toolkit: DoS (division by zero)","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"CUDA Toolkit","year":"2025","cvss_score":2.5,"severity":"low","kev":false,"impact":"DoS (division by zero)","attack_vector":"Malicious artifact","remediation":"Bump CUDA Toolkit; rebuild images","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23273","https://github.com/NVIDIA/product-security/tree/main/2025/5661"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-369"]},{"id":"CVE-2025-23290","cve":"CVE-2025-23290","aliases":[],"title":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko): MULTI-TENANT ISOLATION: A guest can read global GPU metrics that are influenced by work running in other…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA vGPU software - Virtual GPU Manager (host-side vGPU plugin / nvidia.ko)","year":"2025","cvss_score":2.5,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: A guest can read global GPU metrics that are influenced by work running in other tenants' VMs. The attacker sits inside a tenant guest VM and the blast radius is the hypervisor host, which is running every other tenant's vGPU on the same physical GPU. This is precisely the boundary a vGPU-based multi-tenant offering is sold on. Low CVSS (2.5) and no memory corruption, but this is a genuine cross-tenant side channel: utilisation, clock and memory-pressure telemetry that moves with a neighbour's workload leaks the shape of that workload. If you sell confidential or isolated GPU capacity, this is a claim you cannot make while it is unpatched.","attack_vector":"A tenant inside their own guest VM, driving the paravirtualised vGPU control interface. Several of these need only an unprivileged process in the guest; the rest need guest root, which a tenant already has on a VM they rented. No host credentials are involved at any point.","remediation":"Patch the vGPU Manager on the hypervisor host per bulletin 5670. Cost: the highest of any class here. The host driver cannot be reloaded while vGPUs are attached, so every tenant VM on that hypervisor must be live-migrated or powered off - a full host drain. NVIDIA also enforces a supported host/guest driver skew, so budget a matching guest-driver campaign in the same window or tenants lose their vGPU on next boot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23290","https://github.com/NVIDIA/product-security/tree/main/2025/5670"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:H/PR:L/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-200"],"fleet":{"pain_class":"node-drain"}},{"id":"CVE-2025-55703","cve":"CVE-2025-55703","aliases":[],"title":"Sunbird Power IQ 9.2.0 API: Error-based SQL injection through an outdated API endpoint with missing input validation. Low score, but…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Sunbird Power IQ 9.2.0 API","year":"2025","cvss_score":2.5,"severity":"low","kev":false,"impact":"Error-based SQL injection through an outdated API endpoint with missing input validation. Low score, but Power IQ is the power-monitoring layer that holds PDU credentials and outlet-level topology for the estate, so any read primitive into its database is worth closing.","attack_vector":"Access to the Power IQ API.","remediation":"Apply the Sunbird fix. Additionally, disable legacy API endpoints you do not use - the root cause here is an old endpoint left enabled.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-55703"],"status":"curated"},{"id":"CVE-2025-23291","cve":"CVE-2025-23291","aliases":[],"title":"NVIDIA License System - Delegated Licensing Service (DLS): An authorised-looking action leads to information disclosure from the licensing service. Low severity and…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA License System - Delegated Licensing Service (DLS)","year":"2025","cvss_score":2.4,"severity":"low","kev":false,"impact":"An authorised-looking action leads to information disclosure from the licensing service. Low severity and high complexity, but it is inventory data about your entire vGPU estate.","attack_vector":"Adjacent network, high privileges and user interaction required - realistically an insider or a compromised admin session.","remediation":"Update the DLS appliance per bulletin 5705. Cost: appliance restart, no tenant impact.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-23291","https://github.com/NVIDIA/product-security/tree/main/2025/5705"],"status":"curated","cvss_vector":"CVSS:3.1/AV:A/AC:H/PR:H/UI:R/S:C/C:L/I:N/A:N","cwe":["CWE-312"]},{"id":"CVE-2018-25103","cve":"CVE-2018-25103","aliases":["AMI-SA-2024002","originally CVE-2024-3708"],"title":"AMI MegaRAC SPx (embedded lighttpd web server): Use-after-free in the lighttpd request parser embedded in MegaRAC SPx. AMI rates the direct impact as minor…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMI MegaRAC SPx (embedded lighttpd web server)","year":"2018","cvss_score":2.3,"severity":"low","kev":false,"impact":"Use-after-free in the lighttpd request parser embedded in MegaRAC SPx. AMI rates the direct impact as minor - low confidentiality and availability effect via HTTP request smuggling. The reason it belongs on an operator's radar is not its score but what it proves: AMI was still shipping a 2018-era lighttpd in production BMC firmware in 2024, which tells you the embedded OSS stack in your BMCs (lighttpd, nginx, cURL, OpenSSL, busybox) is years behind and is not covered by whatever OS patching process you run on the host.","attack_vector":"Network access to the BMC's web server, unauthenticated but requiring a particular request shape and some user interaction. Reachable from anything that can hit the BMC's HTTP/HTTPS port.","remediation":"Firmware flash to SPx_12.7+ / SPx_13.6, out-of-band per node, ODM-gated. Do not schedule a fleet flash for this CVE alone - its real use is as an argument for building BMC firmware version inventory. The durable action is to start tracking the running BMC build per node and the OSS components inside it, so the next embedded-library CVE is an inventory query rather than a research project.","references":["https://9443417.fs1.hubspotusercontent-na1.net/hubfs/9443417/Security%20Advisories/2024/AMI-SA-2024002.pdf","https://www.runzero.com/blog/lighttpd/","https://nvd.nist.gov/vuln/detail/CVE-2018-25103"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-22853","cve":"CVE-2025-22853","aliases":[],"title":"Intel TDX firmware: Improper synchronisation in TDX firmware, exploitable by a privileged host user to escalate. Race conditions…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX firmware","year":"2025","cvss_score":2.3,"severity":"low","kev":false,"impact":"Improper synchronisation in TDX firmware, exploitable by a privileged host user to escalate. Race conditions in the module are hard to trigger but sit on the tenant boundary.","attack_vector":"Privileged host user, requires winning a race.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-22853","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01312.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-33200","cve":"CVE-2025-33200","aliases":[],"title":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware: A third resource-reuse path in SROOT firmware leaks information to a privileged local caller. These sit in…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA DGX Spark (GB10) - SROOT / OSROOT root-of-trust firmware","year":"2025","cvss_score":2.3,"severity":"low","kev":false,"impact":"A third resource-reuse path in SROOT firmware leaks information to a privileged local caller. These sit in the GB10 root-of-trust chain, so a successful exploit undermines the platform's own attestation and secure-boot story rather than just the OS above it.","attack_vector":"Local access to the DGX Spark. Several of the set need no privileges at all; the rest need host root. This is a desk-side developer box, so physical and local access assumptions are much weaker than for a racked DGX.","remediation":"Apply the DGX Spark firmware update from bulletin 5720. Cost: flash plus reboot, low drain cost given the form factor, but not live-patchable and root-of-trust firmware cannot be rolled back once applied.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-33200","https://github.com/NVIDIA/product-security/tree/main/2025/5720"],"status":"curated","cvss_vector":"CVSS:3.1/AV:L/AC:L/PR:H/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-226"],"fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2020-8562","cve":"CVE-2020-8562","aliases":[],"title":"Kubernetes (kube-apiserver): TOCTOU/DNS-rebinding bypass of the link-local and localhost proxy protections","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"Kubernetes (kube-apiserver)","year":"2020","cvss_score":2.2,"severity":"low","kev":false,"impact":"TOCTOU/DNS-rebinding bypass of the link-local and localhost proxy protections","attack_vector":"Cluster user with proxy rights","remediation":"Rolling control-plane upgrade; network-level egress controls on the control plane","references":["https://nvd.nist.gov/vuln/detail/CVE-2020-8562"],"status":"curated"},{"id":"CVE-2015-3456","cve":"CVE-2015-3456","aliases":[],"title":"QEMU / KVM / Xen (VENOM): VENOM: out-of-bounds write in the virtual Floppy Disk Controller - guest-to-host code execution","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"QEMU / KVM / Xen (VENOM)","year":"2015","cvss_score":2,"severity":"low","kev":false,"impact":"VENOM: out-of-bounds write in the virtual Floppy Disk Controller - guest-to-host code execution; present even when the FDC is disabled in the guest config","attack_vector":"Tenant VM guest","remediation":"QEMU update + VM restart. The origin of the \"unused emulated device is still attack surface\" lesson - audit and strip emulated devices from tenant VM templates","references":["https://nvd.nist.gov/vuln/detail/CVE-2015-3456"],"status":"curated"},{"id":"CVE-2023-0194","cve":"CVE-2023-0194","aliases":[],"title":"GPU Display Driver: Physical memory access","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":2,"severity":"low","kev":false,"impact":"Physical memory access","attack_vector":"Local operator with physical access","remediation":"Driver upgrade; no tenant eviction needed","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0194","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:U/C:N/I:N/A:L","cwe":["CWE-1284"]},{"id":"CVE-2023-0195","cve":"CVE-2023-0195","aliases":[],"title":"GPU Display Driver: Physical memory disclosure","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"GPU Display Driver","year":"2023","cvss_score":2,"severity":"low","kev":false,"impact":"Physical memory disclosure","attack_vector":"Local operator with physical access","remediation":"Driver upgrade at next window","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-0195","https://github.com/NVIDIA/product-security/tree/main/2023/5452"],"status":"curated","cvss_vector":"CVSS:3.1/AV:P/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:N","cwe":["CWE-1284"]},{"id":"CVE-2024-23591","cve":"CVE-2024-23591","aliases":["LEN-150020"],"title":"Lenovo ThinkSystem SR670 V2 (shipped in Manufacturing Mode): SR670 V2 servers built between roughly June 2021 and July 2023 left the factory still in Manufacturing Mode…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Lenovo ThinkSystem SR670 V2 (shipped in Manufacturing Mode)","year":"2024","cvss_score":2,"severity":"low","kev":false,"impact":"SR670 V2 servers built between roughly June 2021 and July 2023 left the factory still in Manufacturing Mode, which means Intel Boot Guard firmware-integrity enforcement and Intel SPS security settings can be modified or disabled. The CVSS is 2.0 and that number is misleading for a GPU operator: the SR670 V2 is a four-GPU A100/H100-class node, so the affected units are precisely the accelerator fleet, and what is broken is the hardware root of trust that is supposed to stop firmware tampering in the first place. The platform's firmware-resilience protections cannot do their job on a node in this state. This is also the rare entry where the defect ships with the hardware rather than accruing over time, so a node that has never been patched since delivery is affected by construction.","attack_vector":"An attacker with privileged logical access to the host, or physical access to the server internals - so an insider, a technician during an RMA or rack move, or a tenant with root on a bare-metal node during their tenancy. Nothing is reachable over the network.","remediation":"Flash UEFI to U8E126I-2.20 or later, which closes Manufacturing Mode. As a system firmware update it applies on the next reboot, so it costs a drain and a maintenance window on a GPU node. The operational action beyond the flash: check delivery dates against the June 2021 - July 2023 window and verify Boot Guard/Manufacturing Mode state per node, because a firmware version check alone will not tell you whether a given unit shipped in this condition.","references":["https://support.lenovo.com/us/en/product_security/LEN-150020","https://nvd.nist.gov/vuln/detail/CVE-2024-23591"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-21096","cve":"CVE-2025-21096","aliases":[],"title":"Intel TDX firmware: Improper buffer restrictions in TDX firmware reachable by a privileged host user for privilege escalation.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel TDX firmware","year":"2025","cvss_score":1.9,"severity":"low","kev":false,"impact":"Improper buffer restrictions in TDX firmware reachable by a privileged host user for privilege escalation. Bundled with the other 2025 TDX firmware fixes.","attack_vector":"Privileged host user.","remediation":"Update the Intel TDX module. The TDX module is loaded by the SEAM loader at boot, so the practical rollout is: stage the new module, drain every trust domain off the node, and reboot. It is not a live-patchable component and running TDs cannot be migrated through it. After the update, every TD must re-attest because the TDX module SVN is part of the attestation report - so anything that pinned the old measurement will fail until you update your attestation policy too. No OEM BIOS release needed for the module itself, which makes this materially faster than a platform firmware update.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-21096","https://www.intel.com/content/www/us/en/security-center/advisory/intel-sa-01312.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-0029","cve":"CVE-2025-0029","aliases":[],"title":"AMD SEV-SNP - selective DMA write drops on host-induced faults: MULTI-TENANT ISOLATION: By inducing faults, a high-privileged local attacker can selectively drop a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - selective DMA write drops on host-induced faults","year":"2025","cvss_score":1.8,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: By inducing faults, a high-privileged local attacker can selectively drop a confidential guest's DMA writes, costing SEV-SNP guest memory integrity. Selectivity is what makes this more than noise: an attacker who can choose *which* writes vanish can corrupt a computation in a targeted way - dropping a gradient update, a checkpoint write, or a log entry - rather than just breaking the VM.","attack_vector":"Local, high-privileged host attacker, against a confidential guest doing DMA.","remediation":"Fixed in AMD SEV firmware / AGESA and reaches you as an OEM SBIOS package - AMD hands AGESA to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before shipping BIOS. **Budget one to six months of OEM lag**, longer on older platforms and sometimes never on end-of-support SKUs. Applying it means draining the host and doing a full power cycle. Because the fix moves the platform's reported SEV-SNP TCB version, you must also pull fresh VCEK certificates from AMD's Key Distribution Service and update any attestation policy your tenants pin - otherwise guests will start failing launch validation the moment the BIOS lands. Some SEV firmware can alternatively be staged from linux-firmware (amd/amd_sev_*.sbin) and committed via the ccp driver at boot, which is faster than waiting on BIOS - check whether your platform supports firmware hot-load before assuming the OEM is the only route. CVSS 1.8 badly undersells this for anyone whose product claim is 'the host operator cannot tamper with your workload'; rate it against your own trust story, not the score.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-0029","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-48509","cve":"CVE-2025-48509","aliases":[],"title":"AMD SEV firmware - missing checks around RMP initialization (AMD-SB-3023): MULTI-TENANT ISOLATION: Missing checks around RMP initialization lead to I/O memory being misidentified…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - missing checks around RMP initialization (AMD-SB-3023)","year":"2025","cvss_score":1.8,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing checks around RMP initialization lead to I/O memory being misidentified, weakening the boundary the reverse-map table enforces. Lowest-severity member of the AMD-SB-3023 RMP batch; include it in the same firmware pass rather than tracking it separately.","attack_vector":"Local, privileged, during SNP platform initialization.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Because this touches the SEV-SNP trust boundary, the update moves the platform TCB version: refresh VCEK certificates from AMD's KDS and update tenant attestation policy, or confidential guest launches will fail immediately after the BIOS lands.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-48509","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3023.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2026-15791","cve":"CVE-2026-15791","aliases":[],"title":"BuildKit: Crafted low-level API message deletes the contents of the host /tmp","layer":"container-orchestration","layer_name":"Container, Kubernetes & orchestration","component":"BuildKit","year":"2026","cvss_score":1.8,"severity":"low","kev":false,"impact":"Crafted low-level API message deletes the contents of the host /tmp","attack_vector":"Anyone with build API access","remediation":"Upgrade BuildKit","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-15791"],"status":"curated"},{"id":"CVE-2023-20518","cve":"CVE-2023-20518","aliases":[],"title":"AMD Secure Processor - incomplete cleanup exposing the Master Encryption Key (AMD-SB-3003): MULTI-TENANT ISOLATION: Incomplete cleanup in the ASP exposes the platform Master Encryption Key. CVSS…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor - incomplete cleanup exposing the Master Encryption Key (AMD-SB-3003)","year":"2023","cvss_score":1.6,"severity":"low","kev":false,"impact":"MULTI-TENANT ISOLATION: Incomplete cleanup in the ASP exposes the platform Master Encryption Key. CVSS 1.6 is the lowest score in this database and the description is the most alarming sentence in it - the MEK is the key underneath the platform's memory encryption. Scores measure exploitability under a specific model; they do not measure what an attacker walks away with. If the MEK leaks, patching afterwards does not put it back, and the node's cryptographic identity is spent.","attack_vector":"Local, privileged, and per AMD's scoring hard to reach in practice - which is why the number is low.","remediation":"Fixed in AMD PI/AGESA firmware and delivered only as an OEM SBIOS package - AMD ships the PI drop to Dell, HPE, Supermicro, Lenovo and the ODMs, who each requalify before releasing BIOS. **Budget one to six months of OEM lag**, and note that several CVEs in this batch are marked 'no fix planned' on Naples (EPYC 7001) - for those the only remediation is retiring the hardware. Applying it means cordon, drain and a full power cycle per node; there is no driver reload, no live patch and no VBIOS step. Marked 'no fix planned' on some generations. Judge this on the asset at risk rather than the score: if you offer confidential computing, an unpatched-and-unpatchable MEK exposure path is something to know about when you write the SLA, not something to leave at the bottom of a CVSS-sorted queue.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-20518","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3003.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-012-qct-quanta-cloud-technology-serv","cve":null,"aliases":[],"title":"QCT (Quanta Cloud Technology) server security centre: QCT firmware is unmeasurable from public data despite the appearance of a PSIRT. This is worse than having no…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"QCT (Quanta Cloud Technology) server security centre","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"QCT firmware is unmeasurable from public data despite the appearance of a PSIRT. This is worse than having no page, because an operator doing vendor due diligence sees a security centre, ticks the box, and moves on - while in reality no QCT-originated BMC or BIOS vulnerability has ever been disclosed publicly, and the channel has been dormant for six years. Quanta is one of the largest ODM server manufacturers in the world and its boards are widely present in cloud and AI datacenter builds, so a substantial amount of deployed BMC firmware has no disclosure history whatsoever. An operator cannot distinguish 'this firmware has no known vulnerabilities' from 'nobody has ever looked, or looked and never told you'. A PSIRT page does exist at qct.io/Press-Releases/index/PR/Server/Security-Center and is readable, but every entry on it is an Intel advisory passed through - Intel-SA-00086, 00088, 00115, 00161, 00125/00131, 00233 - and the newest item dates to 2019. QCT has published no first-party BMC or BIOS advisory at all.","attack_vector":"Not an attack path - a disclosure gap that applies to any operator running QCT or Quanta-manufactured server and GPU chassis hardware.","remediation":"Nothing to flash and nothing to subscribe to. Practical steps: route firmware and security questions through your QCT account team in writing and keep the responses, since the public channel will not serve you; require a firmware support and vulnerability-notification commitment in the purchase agreement before the next order; and in the absence of vendor disclosure, apply the generic BMC controls that do not depend on knowing about specific CVEs - disable IPMI-over-LAN in favour of Redfish over TLS, use per-node unique BMC credentials, disable virtual media and SSH/SMASH on the BMC where unused, isolate the management VLAN, and baseline firmware hashes at turnup so change is at least detectable.","references":["https://www.qct.io/Press-Releases/index/PR/Server/Security-Center"],"status":"curated"},{"id":"NCVD-2026-013-supermicro-s-public-security-adv","cve":null,"aliases":[],"title":"Supermicro's public security advisory portal itself: An operator cannot programmatically track Supermicro firmware advisories. Supermicro publishes real, detailed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro's public security advisory portal itself","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"An operator cannot programmatically track Supermicro firmware advisories. Supermicro publishes real, detailed BMC and BIOS advisories on a roughly quarterly cadence, but the pages are unreadable to any scanner, SBOM pipeline, or vulnerability-management tool that fetches them without a browser session. The practical result is that Supermicro firmware CVEs enter operator awareness late and by hand, via NVD or a vendor account manager, and that fleet-wide 'are we patched' questions cannot be answered automatically. On a Supermicro-heavy GPU fleet this is a measurement gap, not a vulnerability - but it is the reason the vulnerability entries above have vendor advisory links that will not resolve for your tooling. (supermicro.com/en/support/security_center and the dated security_BMC_IPMI_* / security_BIOS_* advisory pages). Every one of them returns HTTP 403 from ordinary automated clients, including the site root and deliberately bogus paths, which means it is a blanket WAF block rather than a missing page.","attack_vector":"Not an attack - a visibility failure. It affects anyone trying to automate firmware advisory ingestion for a Supermicro fleet from a datacenter or CI egress IP rather than a human browser.","remediation":"There is no fix an operator can apply to the vendor's WAF. What works: subscribe to Supermicro's security notification mailing list through your reseller or account team so advisories arrive by email rather than by scraping; mirror each advisory's contents into your own internal tracker when it lands, since you cannot re-fetch it later; and drive automated detection off NVD and the CVE Program's cvelistV5 records, which do carry the Supermicro CNA entries and are freely fetchable. Budget a human in the loop for every Supermicro advisory cycle.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-36435","https://raw.githubusercontent.com/CVEProject/cvelistV5/main/cves/2025/12xxx/CVE-2025-12006.json"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-014-tyan-mitac-computing-psirt","cve":null,"aliases":[],"title":"Tyan / MiTAC Computing PSIRT: For Tyan, this vendor's firmware is unmeasurable from public data - the only substantive Tyan BMC…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tyan / MiTAC Computing PSIRT","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"For Tyan, this vendor's firmware is unmeasurable from public data - the only substantive Tyan BMC vulnerability in the public record, the S5552 TLS private key disclosure, was published by a third-party research lab rather than the vendor, and there is now no vendor channel at all through which fixed firmware or future advisories could be obtained. An operator with Tyan hardware in the fleet has a BMC with a known credential-grade disclosure bug and no path to a fix. For Advantech the situation is different but leads to the same operator conclusion: a working PSIRT exists, so subscribe to it, but do not expect it to cover server BMC or BIOS firmware, because it never has. Tyan has been absorbed into MiTAC Computing; www.tyan.com now serves a TLS certificate for *.mitaccomputing.com that does not match its own hostname, www.tyan.com.tw is a stub linking onward to MiTAC, and mitaccomputing.com returns HTTP 403 to automated clients. Advantech, checked alongside as an industrial-server vendor, does run a real and actively maintained PSIRT at advantech.com/en/security-advisory with ACIRT/AQIRT advisory IDs and an RSS feed - but every advisory on it covers IoT, wireless and software products, with no BMC or BIOS content.","attack_vector":"Not an attack path - a vendor-continuity and disclosure gap. It applies to any operator carrying Tyan-branded server boards, which persist in secondhand and budget capacity builds long after the brand's own support channel has gone.","remediation":"For Tyan hardware, assume no firmware fix will arrive. Approach MiTAC Computing directly through a sales channel if you need firmware, and otherwise treat these nodes as permanently unpatched: replace the BMC TLS certificate with one you control, isolate the management VLAN with an explicit allowlist, disable IPMI-over-LAN and virtual media, use per-node unique credentials, and plan the hardware out of the fleet on a defined timeline rather than indefinitely. For Advantech, subscribe to the ACIRT feed but keep server firmware tracking on a separate mechanism, since the feed does not cover it.","references":["https://www.advantech.com/en/security-advisory","https://www.nozominetworks.com/labs/vulnerability-advisories-cve-2023-2538/"],"status":"curated"},{"id":"NCVD-2026-015-wiwynn-celestica-ingrasys-foxcon","cve":null,"aliases":[],"title":"Wiwynn / Celestica / Ingrasys (Foxconn) / AIC BMC firmware: This vendor's firmware is unmeasurable from public data, and that is the finding. An operator running Wiwynn…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Wiwynn / Celestica / Ingrasys (Foxconn) / AIC BMC firmware","year":"2026","cvss_score":0,"severity":"unscored","kev":false,"impact":"This vendor's firmware is unmeasurable from public data, and that is the finding. An operator running Wiwynn, Celestica, Ingrasys or AIC hardware cannot answer basic questions: is there a known vulnerability in this BMC, has a fix shipped, what version am I supposed to be on. There is no feed to subscribe to and no advisory to correlate against a CVE. Since these builders ship BMC firmware derived from the same AMI MegaRAC and ASPEED lineage as everything else in this database, the realistic assumption is that they inherit the same vulnerability classes - unauthenticated REST handlers, weak firmware verification, credential exposure - without any of the disclosure that would let an operator act. A neocloud whose fleet is largely ODM whitebox is running an out-of-band management plane whose security posture it has no mechanism to assess. Four ODM whitebox and OCP chassis builders whose hardware carries a meaningful share of hyperscale and neocloud GPU capacity. Their corporate websites were fetched and read directly and none of them publishes a security advisory page, a PSIRT contact, or a CVE disclosure channel of any kind. Ingrasys serves HTTP 200 for every path including nonsense ones, and its /security and /psirt paths render the site's 404 message.","attack_vector":"Not an attack path - a disclosure gap. It applies to any operator whose GPU capacity sits on OCP or ODM whitebox chassis from builders who sell to hyperscalers under contract and have never built a public-facing security function.","remediation":"There is no patch, because there is no advisory. What an operator can actually do: make PSIRT existence a procurement requirement and get firmware-update commitments and a security contact written into the purchase contract, since these vendors will respond to a customer of size even without a public channel. Obtain firmware through the integrator or hyperscaler channel that sourced the hardware. Treat these BMCs as permanently unpatched and isolate them accordingly - dedicated management VLAN, no route from tenant networks, explicit management-host allowlist. Finally, measure independently: capture a firmware hash baseline at node turnup so you can at least detect change, since you will never be told about a vulnerability.","references":["https://www.wiwynn.com/","https://www.celestica.com/","https://www.ingrasys.com/","https://www.aicipc.com/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2012-2934","cve":"CVE-2012-2934","aliases":[],"title":"Xen on older AMD CPUs - 64-bit PV guest processor erratum: Xen 4.0 and 4.1 running a 64-bit PV guest on older AMD CPUs did not protect against a processor erratum…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Xen on older AMD CPUs - 64-bit PV guest processor erratum","year":"2012","cvss_score":null,"severity":"unscored","kev":false,"impact":"Xen 4.0 and 4.1 running a 64-bit PV guest on older AMD CPUs did not protect against a processor erratum, letting a guest OS user hang the host. Historical, but it is the earliest entry in a long pattern worth naming: AMD CPU errata that a hypervisor must actively work around, where forgetting the workaround hands guests a host-availability lever.","attack_vector":"From inside a 64-bit PV guest on affected legacy AMD silicon.","remediation":"Fixed in Xen (XSA-9). Hypervisor update plus reboot. Affected silicon is long retired; carried here for completeness of the AMD virtualisation history rather than as an action item.","references":["https://nvd.nist.gov/vuln/detail/CVE-2012-2934"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2013-6885","cve":"CVE-2013-6885","aliases":["Erratum 793"],"title":"AMD 16h processor microcode - locked instructions vs write-combined memory: Interaction between locked instructions and write-combined memory types hangs the system, reachable from an…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD 16h processor microcode - locked instructions vs write-combined memory","year":"2013","cvss_score":null,"severity":"unscored","kev":false,"impact":"Interaction between locked instructions and write-combined memory types hangs the system, reachable from an unprivileged local application. A tenant can wedge the whole machine with a small crafted program - no privilege needed, no recovery short of a hard reset. Included for completeness on legacy AMD hardware; it is the archetype of the 'any tenant can halt the node' class that keeps recurring in CPU errata.","attack_vector":"Local, unprivileged. Any process on the machine, including inside a container.","remediation":"Fixed by a BIOS/microcode update carrying the erratum 793 workaround. Affects AMD 16h family processors, long out of production - if you have any of these in a fleet, the realistic answer is that they are past firmware support and should be retired rather than patched.","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-6885"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2018-0021","cve":"CVE-2018-0021","aliases":[],"title":"Juniper Junos OS MACsec key configuration (CKN/CAK): TENANT ISOLATION: if you configure a MACsec connectivity-association name or key shorter than its full…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS MACsec key configuration (CKN/CAK)","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"TENANT ISOLATION: if you configure a MACsec connectivity-association name or key shorter than its full length, Junos silently zero-fills the remainder. A 16-character passphrase you believed was a 256-bit key is 16 characters followed by a long run of zeros, and it falls to dictionary and brute-force attack. MACsec is what protects inter-site and inter-pod links carrying every tenant's traffic, so a recoverable CAK means an attacker with a tap decrypts the lot — and nothing in the config output tells you the key is weak.","attack_vector":"An attacker with a passive tap on the MACsec-protected link who recovers the key offline. No access to the devices is needed.","remediation":"Config change, not a patch: reconfigure every MACsec association with the full 64-digit CKN and full 32-digit CAK, generated from a CSPRNG. Rekeying a MACsec link drops it briefly, so do redundant links one at a time. Then audit every MACsec key in the fabric for length — this is the kind of defect that survives for years because the config looks fine.","references":["https://nvd.nist.gov/vuln/detail/CVE-2018-0021"],"status":"curated"},{"id":"CVE-2020-10255","cve":"CVE-2020-10255","aliases":["TRRespass"],"title":"DDR4 / LPDDR4 DRAM - Target Row Refresh mitigation: Many-sided Rowhammer defeats the in-DRAM Target Row Refresh mitigation that vendors marketed as making…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DDR4 / LPDDR4 DRAM - Target Row Refresh mitigation","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"Many-sided Rowhammer defeats the in-DRAM Target Row Refresh mitigation that vendors marketed as making Rowhammer solved. On a GPU host node this is the substrate under your host memory: bit flips in page tables or in a hypervisor's structures are a classic path to cross-VM compromise, and on a bare-metal GPU rental it is a path from tenant code to host.","attack_vector":"Local code on the node with the ability to allocate and access memory at a controlled rate. Any tenant container or VM qualifies; no privileges needed.","remediation":"No universal fix. Practical controls, in order of value: use ECC DIMMs and actually monitor correctable-error rates (a Rowhammer campaign shows up as a correctable-error storm before it succeeds), enable the platform's refresh-rate and RFM settings in BIOS where offered, and prefer DDR5 modules with on-die ECC. Cost: a BIOS setting change is a drain and reboot; a DIMM refresh is a hardware refresh cycle. Effectively UNPATCHABLE in software.","references":["https://www.vusec.net/projects/trrespass/","https://download.vusec.net/papers/trrespass_sp20.pdf","https://nvd.nist.gov/vuln/detail/CVE-2020-10255"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"CVE-2021-0891","cve":"CVE-2021-0891","aliases":["PowerVR uninitialized heap disclosure"],"title":"Imagination PowerVR GPU driver - memory residue: MULTI-TENANT ISOLATION: an unprivileged application gets the GPU driver to hand back uninitialized heap…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Imagination PowerVR GPU driver - memory residue","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: an unprivileged application gets the GPU driver to hand back uninitialized heap memory, disclosing whatever the previous consumer left there. Same residual-memory class as LeftoverLocals and included here as evidence the class is a driver-design pattern rather than a one-off - if you evaluate a non-NVIDIA accelerator, ask the vendor specifically what their driver zeroes on allocation.","attack_vector":"An unprivileged local application with access to the GPU device node.","remediation":"Update to the fixed Imagination DDK / vendor driver. In the datacenter this matters as a due-diligence question rather than a fleet action: PowerVR is not a datacenter part. Cost: driver update and reload where it applies.","references":["https://source.android.com/security/bulletin/2022-08-01","https://nvd.nist.gov/vuln/detail/CVE-2021-0891"],"status":"curated"},{"id":"CVE-2022-23645","cve":"CVE-2022-23645","aliases":[],"title":"swtpm (state blob header parsing): An invalid hdrsize in swtpm's saved state header causes an out-of-bounds access, crashing swtpm or preventing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"swtpm (state blob header parsing)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"An invalid hdrsize in swtpm's saved state header causes an out-of-bounds access, crashing swtpm or preventing it from starting. The operationally sharp version: a VM whose vTPM will not start cannot unseal its disk key, so a corrupted or tampered state blob is a data-availability event, not just a crash. On a GPU cloud that migrates or restores VMs, a bad state file taken across a migration bricks the guest's boot.","attack_vector":"Whoever can write the swtpm state file - host-level access, a compromised migration path, or corruption in the storage holding VM state. Not reachable from inside a well-isolated guest.","remediation":"Package update to swtpm 0.5.3 / 0.6.2 / 0.7.1 or later on hypervisor hosts, then restart swtpm processes - package-level, no reboot of the host required. Worth pairing with an operational control: treat vTPM state blobs as data whose integrity you protect and back up, because losing one is equivalent to losing the guest's disk encryption key.","references":["https://github.com/stefanberger/swtpm/security/advisories/GHSA-2qgm-8xf4-3hqw","https://nvd.nist.gov/vuln/detail/CVE-2022-23645"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2022-42335","cve":"CVE-2022-42335","aliases":["XSA-430"],"title":"Xen (shadow paging): x86 shadow paging arbitrary pointer dereference - host crash or worse","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (shadow paging)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"x86 shadow paging arbitrary pointer dereference - host crash or worse","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-430.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50617","cve":"CVE-2022-50617","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/powerplay/psm): A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu/powerplay/psm)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu power management (SMU/powerplay). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/powerplay/psm: Fix memory leak in power state init","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50617","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50619","cve":"CVE-2022-50619","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdkfd (KFD compute driver, /dev/kfd). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdkfd: Fix memory leak in kfd_mem_dmamap_userptr()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50619","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50718","cve":"CVE-2022-50718","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu kernel driver core. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix pci device refcount leak","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50718","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50760","cve":"CVE-2022-50760","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Fix PCI device refcount leak in amdgpu_atrm_get_bios()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50760","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50781","cve":"CVE-2022-50781","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (amdgpu/pm): An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (amdgpu/pm)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu power management (SMU/powerplay) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: amdgpu/pm: prevent array underflow in vega20_odn_edit_dpm_table()","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50781","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2022-50844","cve":"CVE-2022-50844","aliases":[],"title":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu): Missing or insufficient validation of user-supplied parameters in the amdgpu power management…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu power management (SMU/powerplay) (drm/amdgpu)","year":"2022","cvss_score":null,"severity":"unscored","kev":false,"impact":"Missing or insufficient validation of user-supplied parameters in the amdgpu power management (SMU/powerplay). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Fix type of second parameter in odn_edit_dpm_table() callback","attack_vector":"Local. Reachable by a local user with root or with write access to the amdgpu sysfs/debugfs power interfaces; on most clusters that means a compromised node agent or a privileged DaemonSet, not a plain tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2022-50844","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-40286","cve":"CVE-2023-40286","aliases":[],"title":"Supermicro BMC (IPMI web interface): Part of the same 2023 Supermicro BMC web-interface batch - an injection flaw in the management UI that feeds…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Supermicro BMC (IPMI web interface)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"Part of the same 2023 Supermicro BMC web-interface batch - an injection flaw in the management UI that feeds the session-hijack-to-firmware-flash chain.","attack_vector":"Network reach to the BMC web interface; operator interaction for the injection to land.","remediation":"BMC firmware flash per board, bundled with the rest of the batch. Score not independently confirmed - treat it as equivalent to its siblings for prioritisation.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-40286"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2023-52482","cve":"CVE-2023-52482","aliases":[],"title":"Linux x86/srso - SRSO mitigation missing for Hygon processors: MULTI-TENANT ISOLATION: The kernel's Speculative Return Stack Overflow (Inception) mitigation was not applied…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/srso - SRSO mitigation missing for Hygon processors","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: The kernel's Speculative Return Stack Overflow (Inception) mitigation was not applied to Hygon processors, which are AMD Zen derivatives and carry the same defect. Any Hygon node in the fleet therefore ran with the SRSO mitigation silently inactive - a cross-privilege speculative disclosure channel that your vulnerability dashboard reported as mitigated.","attack_vector":"Local, cross-privilege speculative execution on Hygon silicon.","remediation":"Fixed in the Linux kernel by extending the SRSO mitigation to Hygon. Distro kernel update plus reboot; no firmware step. Verify afterwards by reading /sys/devices/system/cpu/vulnerabilities/spec_rstack_overflow on Hygon nodes rather than trusting the CPU vendor string to have been handled correctly - this bug existed precisely because it was not.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-52482"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53703","cve":"CVE-2023-53703","aliases":[],"title":"Linux HID/amd_sfh - shift out of bounds: A shift operation in the AMD Sensor Fusion Hub driver exceeds the maximum valid shift value, producing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux HID/amd_sfh - shift out of bounds","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A shift operation in the AMD Sensor Fusion Hub driver exceeds the maximum valid shift value, producing undefined behaviour in the kernel. Same story as the SFH use-after-free: low real exposure on a server, and a good reminder that unused autoloading drivers cost you something.","attack_vector":"Local, on hosts with amd_sfh loaded.","remediation":"Distro kernel update plus reboot, or blacklist the module on server images and carry neither the bug nor the next one.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53703"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53723","cve":"CVE-2023-53723","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A race condition or locking defect in the amdgpu RAS / GPU reset and recovery path. Concurrent paths touch…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdgpu RAS / GPU reset and recovery path. Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu: disable sdma ecc irq only when sdma RAS is enabled in suspend","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53723","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-53780","cve":"CVE-2023-53780","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd/display: fix FCLK pstate change underflow","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53780","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-54144","cve":"CVE-2023-54144","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Fix kernel warning during topology setup","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54144","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-54150","cve":"CVE-2023-54150","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd): An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu display core (DC/DM) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amd: Fix an out of bounds error in BIOS parser","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54150","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-54261","cve":"CVE-2023-54261","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Add missing gfx11 MQD manager callbacks","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54261","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2023-54296","cve":"CVE-2023-54296","aliases":[],"title":"Linux KVM/SVM - source vCPU selection in SEV-ES intra-host migration: KVM fetched source vCPUs from the wrong VM during SEV-ES intra-host migration. Operating on the wrong VM's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux KVM/SVM - source vCPU selection in SEV-ES intra-host migration","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"KVM fetched source vCPUs from the wrong VM during SEV-ES intra-host migration. Operating on the wrong VM's vCPU structures during a confidential-VM migration is a cross-VM state confusion bug - exactly the failure class you do not want on the code path that moves encrypted guest state around.","attack_vector":"Through the KVM migration ioctl path, from the VMM.","remediation":"Fixed in the Linux kernel. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and reboot the host - no firmware, VBIOS or AGESA step. On a GPU fleet this is a cordon, drain and rolling reboot; plan it as normal kernel maintenance.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-54296"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-29038","cve":"CVE-2024-29038","aliases":[],"title":"tpm2-tools (tpm2_checkquote TPM2_GENERATED magic validation): tpm2_checkquote does not verify that the structure it is checking was actually generated by a TPM, so a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tools (tpm2_checkquote TPM2_GENERATED magic validation)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"tpm2_checkquote does not verify that the structure it is checking was actually generated by a TPM, so a fabricated quote passes validation. The sibling of the tpm2-tss issue in the same disclosure, and it hits the command-line tool that most operators' attestation scripts actually shell out to. Result is the same: a machine with no TPM, or a machine whose TPM state is wrong, can present as attested.","attack_vector":"A malicious or compromised endpoint producing its own quote artefacts. The attacker is the thing claiming to be healthy.","remediation":"Package update to a fixed tpm2-tools wherever verification runs, then restart the verifying service and re-run attestation across the fleet - old results are not evidence. Package-level fix, no reboot. Check whether your attestation pipeline calls tpm2_checkquote, the tpm2-tss FAPI, or its own verifier, because all three had this same missing check and they patch through different channels.","references":["https://github.com/tpm2-software/tpm2-tools/security/advisories/GHSA-5495-c38w-gr6f","https://nvd.nist.gov/vuln/detail/CVE-2024-29038"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"CVE-2024-29039","cve":"CVE-2024-29039","aliases":[],"title":"tpm2-tools (tpm2_checkquote PCR selection handling): tpm2_checkquote does not validate the TPML_PCR_SELECTION in the supplied PCR input file, so an attacker who…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"tpm2-tools (tpm2_checkquote PCR selection handling)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"tpm2_checkquote does not validate the TPML_PCR_SELECTION in the supplied PCR input file, so an attacker who controls that file makes digests map to the wrong PCR slots and banks. The verifier then reports a perfectly valid signature over a misread picture of the machine's state - a node running a tampered boot chain can be made to look like it booted the golden image. Rated critical by the maintainers. For anyone whose product promise is 'verified clean bare metal' or whose scheduler gates on attestation, this is the tool that silently makes the check meaningless.","attack_vector":"The attacker is the attested machine, or anything that can influence the PCR input file the verifier reads. If your verifier consumes quote artefacts uploaded by the node being checked - which is the common design - the node controls its own grade.","remediation":"Package update to a fixed tpm2-tools on every host that verifies quotes, plus a restart of the verifying service. No firmware, no reboot, so it is cheap - but the second half is not: every attestation result produced by the old tool proves nothing, so re-attest the fleet after updating. Structurally, the verifier should pin the expected PCR selection itself rather than accepting one supplied alongside the quote.","references":["https://github.com/tpm2-software/tpm2-tools/security/advisories/GHSA-8rjm-5f5f-h4q6","https://nvd.nist.gov/vuln/detail/CVE-2024-29039"],"status":"curated"},{"id":"CVE-2024-35907","cve":"CVE-2024-35907","aliases":["mlxbf_gige call request_irq() after NAPI initialized"],"title":"Linux kernel mlxbf_gige (BlueField out-of-band management NIC): NULL function-pointer dereference when the DPU's oob_net0 interface comes up, reproduced on BlueField-3 under…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxbf_gige (BlueField out-of-band management NIC)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"NULL function-pointer dereference when the DPU's oob_net0 interface comes up, reproduced on BlueField-3 under kdump. oob_net0 is the out-of-band management interface on the DPU - the path you use to reach a DPU whose in-band networking is broken. Losing it during a crash-dump or recovery boot means losing the remote hands you were counting on.","attack_vector":"Local on the DPU - triggered by interface bring-up during kdump or an unusual boot sequence, not by an external attacker.","remediation":"Upgrade the DPU Arm-side kernel to 6.9 or a stable backport (5.15.154, 6.1.85, 6.6.26, 6.8.5), delivered as a DOCA/BFOS package upgrade or BFB re-image with a DPU reset per node. Bundle with other DPU-side fixes into one maintenance pass.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-35907","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-35907.json"],"status":"curated"},{"id":"CVE-2024-42142","cve":"CVE-2024-42142","aliases":["net/mlx5 E-switch: create ingress ACL when needed"],"title":"Linux kernel mlx5_core eswitch ingress ACL: TENANT ISOLATION: the eswitch ingress ACL - the table that enforces per-VF ingress policy - is only created…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core eswitch ingress ACL","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"TENANT ISOLATION: the eswitch ingress ACL - the table that enforces per-VF ingress policy - is only created when vport metadata match or prio tag is on. Turn vport metadata match off via devlink and bring up an active-backup LAG, and the driver panics on a missing ingress ACL. Beyond the crash, this is a config in which the VF ingress enforcement structure is simply absent when the driver expects it.","attack_vector":"Requires an administrator to have set esw_port_metadata=false via devlink and to be running active-backup LAG. Not tenant-triggered, but it is a realistic operator configuration on bonded ConnectX hosts.","remediation":"Upgrade the host kernel to 6.10 or a stable backport (6.1.98, 6.6.39, 6.9.9). Rolling reboot. Immediate config-level mitigation: keep esw_port_metadata at its default (true) on LAG hosts - a devlink change, no reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-42142","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2024/CVE-2024-42142.json"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2024-53241","cve":"CVE-2024-53241","aliases":["XSA-466"],"title":"Xen (x86 speculation): Xen hypercall page unsafe against speculative attacks - guest leaks hypervisor memory","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 speculation)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Xen hypercall page unsafe against speculative attacks - guest leaks hypervisor memory","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + host reboot with evacuation","references":["https://xenbits.xen.org/xsa/advisory-466.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-1713","cve":"CVE-2025-1713","aliases":["XSA-467"],"title":"Xen (VT-d passthrough): Deadlock potential with VT-d and legacy PCI device pass-through - host hang","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (VT-d passthrough)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Deadlock potential with VT-d and legacy PCI device pass-through - host hang","attack_vector":"Tenant VM guest with PCI passthrough","remediation":"Hypervisor patch + host reboot. Directly in the path of GPU passthrough, the default topology for a GPU cloud","references":["https://xenbits.xen.org/xsa/advisory-467.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21601","cve":"CVE-2025-21601","aliases":[],"title":"Juniper Junos OS (httpd / J-Web on QFX5120, EX, SRX, MX): FABRIC DOS: crafted HTTP requests to the web management process drive CPU consumption up until the device…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS (httpd / J-Web on QFX5120, EX, SRX, MX)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"FABRIC DOS: crafted HTTP requests to the web management process drive CPU consumption up until the device stops responding. QFX5120 is a common data-center leaf, and the vulnerability is in the management daemon, so an attacker who reaches the management interface can take the switch's control plane out without credentials.","attack_vector":"Unauthenticated, remote — reachability to the device's web management service. Only exposed if J-Web / HTTP management is enabled, which many operators leave on for convenience.","remediation":"Junos upgrade (24.2R2 or later on QFX5120) plus reboot. The far cheaper immediate fix is a config change: disable J-Web entirely (`delete system services web-management`) and manage the fabric through NETCONF or the CLI. Most data-center operators should have this off regardless.","references":["https://supportportal.juniper.net/s/article/2025-04-Security-Bulletin-Junos-OS-SRX-and-EX-Series-MX240-MX480-MX960-QFX5120-Series-When-web-management-is-enabled-for-specific-services-an-attacker-may-cause-a-CPU-spike-by-sending-genuine-packets-to-the-device-CVE-2025-21601","https://nvd.nist.gov/vuln/detail/CVE-2025-21601"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-21602","cve":"CVE-2025-21602","aliases":[],"title":"Juniper Junos OS / Junos OS Evolved (rpd, BGP UPDATE): FABRIC DOS: a crafted BGP UPDATE crashes the routing protocol daemon. In a BGP-underlay leaf/spine the rpd…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS / Junos OS Evolved (rpd, BGP UPDATE)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"FABRIC DOS: a crafted BGP UPDATE crashes the routing protocol daemon. In a BGP-underlay leaf/spine the rpd going down is the rack losing reachability; because BGP UPDATEs propagate, a single malicious or malformed announcement can ripple to every device that accepts it, which is how a one-switch bug becomes a fabric-wide outage.","attack_vector":"Unauthenticated attacker able to get a crafted UPDATE into the BGP mesh — either as a peer, or upstream of one, since the message is forwarded along.","remediation":"Junos/Junos Evolved upgrade plus reboot, staged so ECMP paths are never both down. Interim: BGP UPDATE filtering and strict inbound policy at the fabric edge, plus `bgp-error-tolerance` style handling where the release supports it — live config changes.","references":["https://supportportal.juniper.net/s/article/2025-01-Security-Bulletin-Junos-OS-and-Junos-OS-Evolved-Receipt-of-specially-crafted-BGP-update-packet-causes-RPD-crash-CVE-2025-21602","https://nvd.nist.gov/vuln/detail/CVE-2025-21602"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-2296","cve":"CVE-2025-2296","aliases":["GHSA-6pp6-cm5h-86g5"],"title":"EDK II OvmfPkg (X86QemuLoadImageLib, QemuLoadKernelImage direct-boot path): With Secure Boot on, a kernel whose signature is not in the allowed database is correctly rejected by image…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"EDK II OvmfPkg (X86QemuLoadImageLib, QemuLoadKernelImage direct-boot path)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"With Secure Boot on, a kernel whose signature is not in the allowed database is correctly rejected by image verification - and then loaded anyway, because the code falls back to the legacy loader. The guest boots an unsigned kernel while reporting that Secure Boot is enforcing. Any tenant-isolation or attestation story that rests on guest Secure Boot in a QEMU/KVM GPU environment is simply not true on affected OVMF builds.","attack_vector":"Whoever controls the kernel image or the direct-boot command line for the VM - a tenant with access to their own instance's boot configuration, or an attacker who has compromised the image pipeline. Requires the VM to use QEMU direct kernel boot rather than a normal bootloader.","remediation":"Hypervisor-side firmware package update (edk2/OVMF), not a server BIOS flash - update the ovmf/edk2 build on your hosts and restart guests onto the new firmware; no host reboot needed. Config workaround: stop using QEMU direct kernel boot for guests that rely on Secure Boot, and boot through a signed shim/bootloader from a virtual disk instead.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-2296","https://github.com/tianocore/edk2/security/advisories/GHSA-6pp6-cm5h-86g5"],"status":"curated"},{"id":"CVE-2025-23374","cve":"CVE-2025-23374","aliases":[],"title":"Dell Enterprise SONiC (sensitive information in log files): Sensitive information is written into log files on the switch. Switch logs are routinely shipped wholesale to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Dell Enterprise SONiC (sensitive information in log files)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Sensitive information is written into log files on the switch. Switch logs are routinely shipped wholesale to a central syslog or observability stack that far more people can read than can log into the switch — so secrets written to the log leak to a much wider audience than the device's own access control implies.","attack_vector":"Anyone with read access to the switch's logs or to the log aggregation pipeline they are shipped to.","remediation":"Upgrade to Enterprise SONiC 4.4.1 or 4.2.3 or later — NOS image upgrade plus reboot. Also purge historical logs from your aggregator and rotate anything that appeared in them; that cleanup is the part people skip.","references":["https://www.dell.com/support/kbdoc/en-us/000340083/dsa-2025-275-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities","https://nvd.nist.gov/vuln/detail/CVE-2025-23374"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-2826","cve":"CVE-2025-2826","aliases":["Arista Security Advisory 0120"],"title":"Arista EOS (ingress ACL enforcement on ethernet/LAG): TENANT ISOLATION: with IPv4 ingress, MAC ingress, or IPv6 standard ingress ACLs applied to one or more…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arista EOS (ingress ACL enforcement on ethernet/LAG)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"TENANT ISOLATION: with IPv4 ingress, MAC ingress, or IPv6 standard ingress ACLs applied to one or more ethernet or LAG interfaces, the policies may not be enforced at all. The third ACL-enforcement defect in the same family — worth treating Arista ingress ACLs as a control that needs periodic active verification rather than a set-and-forget boundary.","attack_vector":"Any traffic arriving on an affected interface. No attacker capability required.","remediation":"EOS upgrade plus reload on affected platforms. Because ACL enforcement is the thing that fails, the only trustworthy verification is sending traffic that should be dropped and confirming it is — do that as a standing test in your fabric CI, not just after this patch.","references":["https://www.arista.com/en/support/advisories-notices/security-advisory/21414-security-advisory-0120"],"status":"curated"},{"id":"CVE-2025-37866","cve":"CVE-2025-37866","aliases":[],"title":"Linux kernel mlxbf-bootctl (BlueField secure boot fuse state): The BlueField boot-control driver mishandles the sysfs buffer when reporting secure-boot fuse state…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlxbf-bootctl (BlueField secure boot fuse state)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The BlueField boot-control driver mishandles the sysfs buffer when reporting secure-boot fuse state, producing a kernel warning. Low direct impact, but it sits on the interface an operator uses to verify that a DPU's secure boot fuses are actually burned - if you attest DPU secure-boot state by reading this sysfs node, the reporting path itself was not sound.","attack_vector":"Local on the DPU - triggered by reading the secure_boot_fuse_state sysfs attribute on BlueField.","remediation":"Upgrade the DPU Arm-side kernel via a DOCA/BFOS package upgrade or BFB re-image; requires a DPU reset per node. Low urgency on its own - fold it into the next scheduled BFB refresh rather than taking a dedicated window.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-37866","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-37866.json"],"status":"curated"},{"id":"CVE-2025-40148","cve":"CVE-2025-40148","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Add NULL pointer checks in dc_stream cursor attribute functions","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40148","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40191","cve":"CVE-2025-40191","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: Fix kfd process ref leaking when userptr unmapping","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40191","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40288","cve":"CVE-2025-40288","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdgpu GEM/VM/command-submission ioctl surface. Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdgpu: Fix NULL pointer dereference in VRAM logic for APU devices","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40288","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40289","cve":"CVE-2025-40289","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: hide VRAM sysfs attributes on GPUs without VRAM","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40289","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40300","cve":"CVE-2025-40300","aliases":["VMSCAPE"],"title":"AMD Zen 1-Zen 5 - branch predictor isolation between guest and userspace hypervisor (AMD-SB-7046): MULTI-TENANT ISOLATION: Insufficient branch-predictor isolation between a guest VM and the…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"AMD Zen 1-Zen 5 - branch predictor isolation between guest and userspace hypervisor (AMD-SB-7046)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Insufficient branch-predictor isolation between a guest VM and the **userspace** hypervisor process - QEMU - lets a malicious guest train the predictor and then steer speculation inside the VMM that manages it. The VMM has the guest's memory mapped, so this is a Spectre-class read of confidential guest state from a process the guest can influence. AMD rates it High and, unusually, it spans **Zen 1 through Zen 5** - there is no 'we are on new silicon' escape from this one.","attack_vector":"From inside a guest VM. Tenant-reachable, no host privilege required. Affects every AMD generation currently in datacenter service.","remediation":"Fixed in the Linux kernel by issuing a conditional IBPB after every VMexit before returning to userspace. Take the distro kernel update and reboot the host - no firmware, BIOS or microcode step, which makes it one of the cheaper fixes to deploy. **Expect a real performance cost**: the kernel commit itself notes the IBPB duplicates the context-switch IBPB and that workloads switching frequently between hypervisor and userspace absorb the most overhead. On a virtualised GPU fleet with heavy device emulation, benchmark before and after rather than assuming it is free.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40300","https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7046.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40310","cve":"CVE-2025-40310","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (amd/amdkfd)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: amd/amdkfd: resolve a race in amdgpu_amdkfd_device_fini_sw","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40310","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40332","cve":"CVE-2025-40332","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdkfd (KFD compute driver, /dev/kfd). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdkfd: Fix mmap write lock not release","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40332","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40335","cve":"CVE-2025-40335","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu user-mode queues (doorbell submission path). A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate userq input args","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40335","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-40339","cve":"CVE-2025-40339","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix nullptr err of vm_handle_moved","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-40339","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-49133","cve":"CVE-2025-49133","aliases":[],"title":"libtpms (CryptHmacSign, vTPM): Out-of-bounds read when signKey and signScheme are mismatched, aborting the vTPM. libtpms is the TPM behind…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"libtpms (CryptHmacSign, vTPM)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Out-of-bounds read when signKey and signScheme are mismatched, aborting the vTPM. libtpms is the TPM behind QEMU/KVM guests, so on a GPU cloud that rents VMs this is a guest reaching out and killing its own virtual TPM - which takes down measured boot and any disk unlock or key sealing that depended on it, and on some stacks takes the guest with it. The libtpms instance of the same defect as the TCG reference implementation issue, which is worth noting because they patch through completely different channels.","attack_vector":"A guest able to issue TPM commands to its vTPM - i.e. any tenant in a VM you provisioned with a virtual TPM. No escape required, no host access.","remediation":"libtpms package update on the hypervisor hosts plus a restart of affected guests' swtpm processes - so a rolling VM restart rather than a firmware flash, which is the cheap end of this database. Do not assume patching the host TPM stack covers your physical TPMs or your platform firmware TPM; those are separate code paths with separate fixes.","references":["https://github.com/stefanberger/libtpms/security/advisories/GHSA-25w5-6fjj-hf8g","https://nvd.nist.gov/vuln/detail/CVE-2025-49133"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-52953","cve":"CVE-2025-52953","aliases":["JSA100059"],"title":"Juniper Junos OS / Junos OS Evolved (rpd BGP session handling): FABRIC DOS: a genuine, valid BGP UPDATE message resets a live BGP session between peers. Because the message…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS / Junos OS Evolved (rpd BGP session handling)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"FABRIC DOS: a genuine, valid BGP UPDATE message resets a live BGP session between peers. Because the message is valid, no filter catches it — the fabric tears down and rebuilds sessions on demand from an adjacent attacker. Repeated, it keeps the underlay in permanent reconvergence, which for a synchronous training job is indistinguishable from an outage.","attack_vector":"Unauthenticated attacker with adjacent network access able to originate a BGP UPDATE into the fabric.","remediation":"Junos upgrade plus reboot, staged across ECMP pairs. There is no clean config workaround because the triggering message is legitimate; tighten which peers you accept sessions from and monitor for session-reset storms in the meantime.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-52953"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-52989","cve":"CVE-2025-52989","aliases":[],"title":"Juniper Junos OS / Junos OS Evolved (annotate configuration command): The `annotate` configuration command can be used to escalate privileges. A user with limited configuration…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS / Junos OS Evolved (annotate configuration command)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The `annotate` configuration command can be used to escalate privileges. A user with limited configuration rights — the sort of account you give an automation system or a junior operator — gains more than their class allows on a device that may be a spine. Junos login classes are the mechanism most Juniper shops use for internal separation of duties, and this puts a hole in it.","attack_vector":"Authenticated low-privileged user with configuration access to the device.","remediation":"Junos upgrade plus reboot. Interim: restrict the `annotate` command in affected login classes via the allow-/deny-commands regex — a live config change that closes it without a maintenance window.","references":["https://supportportal.juniper.net/s/article/2025-07-Security-Bulletin-Junos-OS-and-Junos-OS-Evolved-Annotate-configuration-command-can-be-used-for-privilege-escalation-CVE-2025-52989"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-54505","cve":"CVE-2025-54505","aliases":["XSA-488"],"title":"Xen / x86 CPU: Floating Point Divider State Sampling - transient-execution leak of FP divider state across domains","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen / x86 CPU","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Floating Point Divider State Sampling - transient-execution leak of FP divider state across domains","attack_vector":"Tenant VM guest; any tenant process in a container","remediation":"Microcode + hypervisor/kernel mitigation + reboot; standing perf cost on FP-heavy workloads","references":["https://xenbits.xen.org/xsa/advisory-488.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"CVE-2025-58147","cve":"CVE-2025-58147","aliases":["XSA-475"],"title":"Xen (Viridian): Incorrect input sanitisation in Viridian (Hyper-V enlightenment) hypercalls - guest attacks the hypervisor","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (Viridian)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Incorrect input sanitisation in Viridian (Hyper-V enlightenment) hypercalls - guest attacks the hypervisor","attack_vector":"Tenant VM guest (Windows guests using Viridian)","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-475.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-59957","cve":"CVE-2025-59957","aliases":[],"title":"Juniper Junos OS (QFX5000-Series, EX4600-Series): A physical-access path into affected QFX5000 and EX4600 switches. QFX5000 is a mainstream data-center leaf…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS (QFX5000-Series, EX4600-Series)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A physical-access path into affected QFX5000 and EX4600 switches. QFX5000 is a mainstream data-center leaf platform. Physical-access bugs matter more in this market than they used to: GPU capacity is increasingly resold, subleased and colocated, so 'someone was alone in the rack' is a routine event rather than an exotic threat model.","attack_vector":"Physical access to the switch. Realistic in colocation, shared cages, subleased capacity, and during rack-and-stack by contractors.","remediation":"Junos upgrade plus reboot on affected QFX5000/EX4600 platforms. Pair with physical controls — locked cabinets, console-port discipline, and tamper-evident seals — because a firmware patch does not address the underlying access.","references":["https://supportportal.juniper.net/s/article/2025-10-Security-Bulletin-Junos-OS"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-59969","cve":"CVE-2025-59969","aliases":[],"title":"Juniper Junos OS Evolved (QFX5000 / PTX, multicast packet handling): FABRIC DOS: crafted multicast packets crash and restart evo-aftmand / evo-pfemand — the forwarding-plane…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Juniper Junos OS Evolved (QFX5000 / PTX, multicast packet handling)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"FABRIC DOS: crafted multicast packets crash and restart evo-aftmand / evo-pfemand — the forwarding-plane management daemons on Junos Evolved. Affects QFX5000-series and PTX, both of which appear in AI-cluster spine and super-spine roles. Multicast is reachable from any tenant that can source it, and the crash takes the forwarding plane with it.","attack_vector":"An attacker able to send crafted multicast packets into the fabric — any tenant workload on a VLAN the device serves.","remediation":"Junos Evolved upgrade plus reboot. Interim: rate-limit or filter unexpected multicast at the fabric edge with a control-plane policer — live config. If your cluster does not use multicast (most GPU fabrics do not, outside of some MPI transports), block it outright at the edge.","references":["https://supportportal.juniper.net/s/article/2026-04-Security-Bulletin-Junos-OS-Evolved-QFX5000-Series-and-PTX-Series-An-attacker-sending-crafted-multicast-packets-will-cause-evo-aftmand-evo-pfemand-to-crash-and-restart-CVE-2025-59969"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68172","cve":"CVE-2025-68172","aliases":[],"title":"ASPEED crypto/ACRY accelerator driver (drivers/crypto/aspeed): The ACRY driver's probe error path and its remove path both hand-free a clock that the device-managed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED crypto/ACRY accelerator driver (drivers/crypto/aspeed)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The ACRY driver's probe error path and its remove path both hand-free a clock that the device-managed allocator will free again, giving a double free in the BMC kernel. It sits in the crypto accelerator the BMC uses for TLS and firmware signature work, which is the wrong place to have allocator corruption: a double free in a driver that touches signing and TLS paths is the kind of primitive that turns into controlled kernel memory reuse rather than just a crash. Realistic near-term impact is BMC kernel instability during driver load failures or module removal.","attack_vector":"BMC-local. Reached through driver probe failure or an explicit driver removal, so it needs root on the BMC or a boot-time condition that makes probe fail. Not host- or network-reachable directly.","remediation":"Kernel fix, backported into 6.6.117, 6.12.58 and 6.17.8 and later. In practice: BMC firmware flash per node, out-of-band, whenever your ODM rebases - which for a fix this recent will realistically be a full release cycle away. No config mitigation; the crypto driver is loaded because bmcweb's TLS wants it. Track it, do not run a special campaign for it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68172","https://git.kernel.org/stable/c/e8407dfd267018f4647ffb061a9bd4a6d7ebacc6"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-68180","cve":"CVE-2025-68180","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu display core (DC/DM). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amd/display: Fix NULL deref in debugfs odm_combine_segments","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68180","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68190","cve":"CVE-2025-68190","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/atom): A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/atom)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu/atom: Check kcalloc() for WS buffer in amdgpu_atom_execute_table_locked()","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68190","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68196","cve":"CVE-2025-68196","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: Cache streams targeting link when performing LT automation","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68196","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68201","cve":"CVE-2025-68201","aliases":[],"title":"Linux kernel amdgpu kernel driver core (drm/amdgpu): A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu kernel driver core (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu kernel driver core reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu: remove two invalid BUG_ON()s","attack_vector":"Local. Reachable by a local user with a render node open, i.e. reachable from inside a GPU tenant container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68201","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68230","cve":"CVE-2025-68230","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu): MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdgpu)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: An out-of-bounds access in the amdkfd (KFD compute driver, /dev/kfd) - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu: fix gpu page fault after hibernation on PF passthrough","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68230","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68798","cve":"CVE-2025-68798","aliases":[],"title":"Linux perf/x86/amd - general protection fault from a NULL event on enable: A subtle race lets cpuc","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux perf/x86/amd - general protection fault from a NULL event on enable","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"A subtle race lets cpuc->events[idx] become NULL on AMD machines, and enabling the event then takes a general protection fault, panicking the host. Same family as the other AMD perf races: the monitoring stack crashes the machine it is monitoring.","attack_vector":"Local, through perf event scheduling. Reachable by observability agents and by tenants where perf is permitted.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68798"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2025-68816","cve":"CVE-2025-68816","aliases":["net/mlx5 fw_tracer validate format string parameters"],"title":"Linux kernel mlx5_core firmware tracer (diag/fw_tracer): The firmware tracer took format strings directly from device firmware and passed them to kernel formatting…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel mlx5_core firmware tracer (diag/fw_tracer)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The firmware tracer took format strings directly from device firmware and passed them to kernel formatting with no validation. Malicious or compromised NIC/DPU firmware supplying %s, %p or %n reads arbitrary kernel memory. This is the clearest firmware-as-attacker case in the mlx5 stack, and it matters specifically for operators who accept hardware from third parties, run rented bare metal, or cannot fully attest NIC firmware provenance - the host kernel was trusting the device.","attack_vector":"Requires control of the NIC or DPU firmware image - a supply-chain or prior-tenant-persistence scenario on bare metal, or an attacker who already flashed the adapter. Not reachable from ordinary network traffic.","remediation":"Upgrade the host kernel to 6.19 or a stable backport (5.10.248, 5.15.198, 6.1.160, 6.6.120, 6.12.64, 6.18.3) - the fix restricts the tracer to integer and hex specifiers. Rolling reboot. Pair it with the real control: enforce signed firmware and re-flash adapters to a known-good version between bare-metal tenants.","references":["https://nvd.nist.gov/vuln/detail/CVE-2025-68816","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2025/CVE-2025-68816.json"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-7026","cve":"CVE-2025-7026","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, unchecked RBX pointer): An attacker-controlled register is used as an unchecked pointer inside System Management Mode, giving…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, unchecked RBX pointer)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"An attacker-controlled register is used as an unchecked pointer inside System Management Mode, giving arbitrary write in ring -2. SMM sits beneath the hypervisor and beneath the OS kernel, so code that lands here can disable Secure Boot, tamper with firmware, and persist through a full disk wipe and OS reinstall. On a rented GPU node this is the canonical 'tenant leaves something behind for the next tenant' primitive, and no host-based tooling can detect it.","attack_vector":"Local privileged code on the host - kernel-level or a driver, so a tenant with root on bare metal, or an attacker who already got kernel execution. Not remote, but on bare-metal rental the precondition is exactly what you sell.","remediation":"UEFI/BIOS firmware update from Gigabyte, per board model, requiring a host reboot - which on GPU nodes means draining running training jobs. Gigabyte shipped fixed firmware; the practical problem is coverage, because affected models span consumer and server lines and not every SKU gets an image. There is no config workaround for an SMM callout. If you buy Gigabyte boards, make the firmware version part of your node-acceptance check, not a post-hoc audit.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7026"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-7027","cve":"CVE-2025-7027","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, NVRAM double pointer dereference): An unvalidated NVRAM variable is dereferenced twice inside SMM, letting a local attacker write to SMRAM and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, NVRAM double pointer dereference)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"An unvalidated NVRAM variable is dereferenced twice inside SMM, letting a local attacker write to SMRAM and gain ring -2 execution. NVRAM variables are writable from the OS on many platforms, which makes this a comparatively short path from host root to firmware-level persistence.","attack_vector":"Local privileged code on the host, able to set the UEFI variable. Any tenant with root on a bare-metal node qualifies.","remediation":"Gigabyte BIOS update per board model plus reboot. No config workaround - locking down NVRAM variable writes is not generally available to operators. Bundle with the other three SMM CVEs in the same advisory; they ship in the same image.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7027"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-7028","cve":"CVE-2025-7028","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, unvalidated flash function pointers): Function pointer structures governing SPI flash operations are not validated, so an attacker in SMM context…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, unvalidated flash function pointers)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Function pointer structures governing SPI flash operations are not validated, so an attacker in SMM context can redirect the routines that read, write and erase the platform firmware. This is the worst of the four: it hands the attacker the flash-write primitive directly, meaning a permanent bootkit rather than a runtime compromise.","attack_vector":"Local privileged code on the host.","remediation":"Gigabyte BIOS update per board plus reboot. Because the payoff is a flash write, assume any node you believe was compromised needs firmware re-flashed from a known-good image and its integrity independently verified - a BIOS update applied by a compromised system does not prove anything. For high-value nodes, external SPI verification is the only real assurance.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7028"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2025-7029","cve":"CVE-2025-7029","aliases":["VU#746790"],"title":"Gigabyte UEFI firmware (SMM, OcHeader/OcData pointer control): Unchecked register use lets the attacker control the OcHeader and OcData pointers handled in SMM, producing…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (SMM, OcHeader/OcData pointer control)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"Unchecked register use lets the attacker control the OcHeader and OcData pointers handled in SMM, producing arbitrary read and write at ring -2 - the same firmware-persistence outcome as the rest of the batch.","attack_vector":"Local privileged code on the host.","remediation":"Gigabyte BIOS update per board plus reboot; delivered in the same firmware image as CVE-2025-7026 through 7028, so treat all four as one change.","references":["https://kb.cert.org/vuls/id/746790","https://nvd.nist.gov/vuln/detail/CVE-2025-7029"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-21444","cve":"CVE-2026-21444","aliases":[],"title":"libtpms (OpenSSL 3.x symmetric cipher IV handling): libtpms 0.10.0/0.10.1 built against OpenSSL 3.x returned the initial IV instead of the last IV for certain…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"libtpms (OpenSSL 3.x symmetric cipher IV handling)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"libtpms 0.10.0/0.10.1 built against OpenSSL 3.x returned the initial IV instead of the last IV for certain symmetric ciphers, weakening every subsequent encrypt and decrypt step in the chain. Anything a tenant's vTPM encrypted - sealed keys, protected blobs - is protected less than the cryptography claims. Quiet failure mode: nothing errors, the data is just weaker than the threat model assumes, and you only find out when someone attacks it.","attack_vector":"No active attacker needed to introduce the weakness - it is present in every affected operation. Exploiting it requires an attacker who obtains the ciphertext, which for vTPM state means host-level access or a leaked VM state file.","remediation":"libtpms package update on hypervisor hosts and restart the swtpm processes. Because the weakness is in data already produced, rotate anything a vulnerable vTPM sealed rather than assuming the update is retroactive. Package-level, no firmware flash - but the re-sealing step is the part that takes planning on a fleet with long-lived guests.","references":["https://github.com/stefanberger/libtpms/security/advisories/GHSA-7jxr-4j3g-p34f","https://nvd.nist.gov/vuln/detail/CVE-2026-21444"],"status":"curated"},{"id":"CVE-2026-23034","cve":"CVE-2026-23034","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A memory or reference-count leak in the amdgpu user-mode queues (doorbell submission path). Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu user-mode queues (doorbell submission path). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu/userq: Fix fence reference leak on queue teardown v2","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23034","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23051","cve":"CVE-2026-23051","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu firmware, ACPI and IP-block initialisation. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: fix drm panic null pointer when driver not support atomic","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-23051","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-23554","cve":"CVE-2026-23554","aliases":["XSA-480"],"title":"Xen (EPT): Use-after-free of EPT paging structures - HVM guest to host compromise","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (EPT)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Use-after-free of EPT paging structures - HVM guest to host compromise","attack_vector":"Tenant VM guest (HVM)","remediation":"Hypervisor patch + host reboot with guest evacuation. Highest-severity recent Xen item for a multi-tenant HVM fleet","references":["https://xenbits.xen.org/xsa/advisory-480.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-34878","cve":"CVE-2026-34878","aliases":["TFV-15","FIP ToC offset validation"],"title":"Arm Trusted Firmware-A BL1/BL2 boot stages on platforms that load firmware from a Firmware Image Package (FIP) container: FIP headers are not signed, so BL1/BL2 parse attacker-influenceable…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Arm Trusted Firmware-A BL1/BL2 boot stages on platforms that load firmware from a Firmware Image Package (FIP) container","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"FIP headers are not signed, so BL1/BL2 parse attacker-influenceable metadata before authentication happens. Bad offsets and lengths make the loader read from secure memory mapped in the EL3 translation regime and land that data in non-secure memory. That is a pre-authentication leak of secure-world contents at boot - keys, secure heap, whatever is mapped. For a bare-metal provider this is a tenant-handoff problem: a tenant who can write the boot FIP (firmware update path, A/B slot, provisioning share) gets to exfiltrate secure-world state on the next boot for the next tenant.","attack_vector":"An attacker who can modify or supply the FIP image the platform boots from - anyone with the firmware-update path, a writable boot partition, or control of the provisioning pipeline. On bare-metal GPU rental that is the previous tenant if you do not verify firmware between handoffs.","remediation":"Update TF-A past commits 48351 and 49485 (strict ToC bounds validation, overflow-safe arithmetic, short-read detection) and rebuild BL1/BL2 for the platform. That is an OEM-shipped firmware build, flashed to each node, reboot and drain. Independently: treat the FIP as tenant-writable until proven otherwise, and re-verify or re-flash boot firmware from a known-good image at every tenant handoff rather than relying on the parser being safe.","references":["https://trustedfirmware-a.readthedocs.io/en/latest/security_advisories/security-advisory-tfv-15.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"CVE-2026-42487","cve":"CVE-2026-42487","aliases":["XSA-491"],"title":"Xen (x86 HVM): x86 HVM I/O port list traversal flaw","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (x86 HVM)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"x86 HVM I/O port list traversal flaw","attack_vector":"Tenant VM guest (HVM)","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-491.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-61810","cve":"CVE-2026-61810","aliases":["DMTF-2026-0003"],"title":"DMTF SPDM specification DSP0274 1.4 (FINISH transcript definition): A specification-level defect rather than an implementation bug. SPDM 1.4 added OpaqueDataLength and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DMTF SPDM specification DSP0274 1.4 (FINISH transcript definition)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A specification-level defect rather than an implementation bug. SPDM 1.4 added OpaqueDataLength and OpaqueData fields to FINISH and FINISH_RSP, but the transcript definitions used to compute the FINISH signature and the verify-data HMACs were never updated to cover them. A literal implementation therefore authenticates a message while leaving part of that message unauthenticated - so an attacker in the path can alter the opaque fields without breaking the signature. This is exactly the class of flaw that undermines device attestation quietly: everything validates, and the validated thing is not what was sent.","attack_vector":"An attacker able to modify SPDM traffic between requester and responder - a PCIe interposer, a compromised switch or retimer in the path, or a malicious intermediary in a disaggregated fabric where SPDM crosses a network rather than a board trace.","remediation":"Requires a specification erratum plus updated implementations on both ends, so the fix arrives as vendor firmware updates once DMTF publishes and vendors rebase - expect a long tail and no OS-level patch. Nothing to configure. The operator action today is to know that SPDM 1.4 opaque data is not integrity-protected in affected implementations, and not to build any policy decision on values carried there.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-chjj-xvqx-c8w4","https://nvd.nist.gov/vuln/detail/CVE-2026-61810"],"status":"curated"},{"id":"CVE-2026-62428","cve":"CVE-2026-62428","aliases":["XSA-500"],"title":"Xen (grant tables): Type confusion in grant-copy - guest corrupts hypervisor state","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (grant tables)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Type confusion in grant-copy - guest corrupts hypervisor state","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation. Grant tables are on the hot path for every PV driver, so this is unavoidable exposure","references":["https://xenbits.xen.org/xsa/advisory-500.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-62435","cve":"CVE-2026-62435","aliases":["XSA-501"],"title":"Xen (grant tables): Grant-table version change racing with other operations","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Xen (grant tables)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Grant-table version change racing with other operations","attack_vector":"Tenant VM guest","remediation":"Hypervisor patch + reboot/evacuation","references":["https://xenbits.xen.org/xsa/advisory-501.html"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63878","cve":"CVE-2026-63878","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: check num_entries in GEM_OP GET_MAPPING_INFO","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63878","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63880","cve":"CVE-2026-63880","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: fix lock leak on ENOMEM in AMDGPU_GEM_OP_GET_MAPPING_INFO","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63880","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-63882","cve":"CVE-2026-63882","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdkfd (KFD compute driver, /dev/kfd). An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdkfd: fix NULL pointer bug in svm_range_set_attr","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-63882","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64229","cve":"CVE-2026-64229","aliases":[],"title":"Linux x86/mm - broadcast TLB flush with PCID disabled: Booting with nopcid clears the PCID feature but broadcast TLB flushing stayed enabled, leaving TLB…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux x86/mm - broadcast TLB flush with PCID disabled","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Booting with nopcid clears the PCID feature but broadcast TLB flushing stayed enabled, leaving TLB invalidation in an inconsistent configuration. Stale TLB entries are a memory-isolation problem: a translation that should have been invalidated but was not means one address space can still reach a mapping that was revoked.","attack_vector":"Local, on hosts booted with nopcid. Not attacker-selected unless the attacker controls boot parameters - but plenty of fleets set nopcid for debugging or for old mitigation workarounds and forget it.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot. Also audit your boot parameters: nopcid is a performance and now correctness liability that is often left in place long after the reason for it is gone.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64229"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-64309","cve":"CVE-2026-64309","aliases":[],"title":"Linux crypto/ccp - SNP initialization on ioctl(SNP_COMMIT): The ccp driver initialised SNP from the SNP_COMMIT ioctl path, so a userspace process holding /dev/sev could…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux crypto/ccp - SNP initialization on ioctl(SNP_COMMIT)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The ccp driver initialised SNP from the SNP_COMMIT ioctl path, so a userspace process holding /dev/sev could drive SNP platform initialisation at a time the host was not expecting it - including while ordinary VMs are running. Triggering SNP platform state transitions underneath live guests is a route to destabilising the host and the confidential-computing state machine on it.","attack_vector":"Local, from a process with access to /dev/sev. On a well-run host that is the VMM or a management daemon, so the realistic path is a compromised control-plane component rather than a tenant.","remediation":"Fixed in the Linux kernel - KVM/x86 SEV code or the ccp/PSP driver. Take the distro kernel update (RHEL/Rocky, Ubuntu, SLES) and **reboot the host**; SEV/SNP hypervisor paths cannot be live-patched in any meaningful way, and SNP platform init/shutdown is not safe to cycle under running guests. Drain confidential-VM tenants, reboot, then re-admit. No firmware, VBIOS or AGESA step needed, which makes this one of the cheaper classes of SEV fix to roll out. Worth checking who actually has /dev/sev open on your hosts - the permissions on that node are the difference between 'root only' and 'any service account that got a bit too much'.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-64309"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68102","cve":"CVE-2026-68102","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix aperture mapping leak","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68102","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68105","cve":"CVE-2026-68105","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: Fix kernel panic during driver load failure","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68105","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68109","cve":"CVE-2026-68109","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.1): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.1)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma7.1: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68109","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68110","cve":"CVE-2026-68110","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma4.4.2): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma4.4.2)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma4.4.2: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68110","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68111","cve":"CVE-2026-68111","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx9: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68111","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68112","cve":"CVE-2026-68112","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9.4.3): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx9.4.3)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx9.4.3: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68112","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68113","cve":"CVE-2026-68113","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx12: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68113","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68114","cve":"CVE-2026-68114","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12.1): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx12.1)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx12.1: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68114","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68115","cve":"CVE-2026-68115","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx10): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx10)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx10: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68115","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68234","cve":"CVE-2026-68234","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu GEM/VM/command-submission ioctl surface. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: fix bo->pin leaking in amdgpu_bo_create_reserved","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68234","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68235","cve":"CVE-2026-68235","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: dce100: skip non-DP stream encoders for DP MST","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68235","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68237","cve":"CVE-2026-68237","aliases":[],"title":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq): A race condition or locking defect in the amdgpu user-mode queues (doorbell submission path). Concurrent…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu user-mode queues (doorbell submission path) (drm/amdgpu/userq)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdgpu user-mode queues (doorbell submission path). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amdgpu/userq: fix indefinite fence wait during GPU reset","attack_vector":"Local. Reachable by any process with a render node open that can create user-mode queues - the normal ROCm submission path, reachable from an unprivileged container. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68237","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68238","cve":"CVE-2026-68238","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu): A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu firmware, ACPI and IP-block initialisation. Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amdgpu: Release VFCT ACPI table reference","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68238","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68243","cve":"CVE-2026-68243","aliases":[],"title":"Linux i915 GPU kernel driver (context SSEU parameter): NULL dereference reachable by setting a context engine slot to an invalid engine class and then querying the…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux i915 GPU kernel driver (context SSEU parameter)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"NULL dereference reachable by setting a context engine slot to an invalid engine class and then querying the SSEU context parameter. A tenant can panic the GPU node from an ordinary unprivileged ioctl sequence.","attack_vector":"Any local user or container with a DRM render node - i.e. any tenant that was scheduled a GPU. No privileged capability needed.","remediation":"Fix ships in the Linux kernel. Update the kernel and reboot the node - in practice this is a drain plus reboot because the accelerator driver cannot be unloaded while jobs hold device file descriptors. No BIOS or firmware update needed.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68243","https://git.kernel.org/stable/c/2b56757a9a7456825eb668fde92299e01c5e2721"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68246","cve":"CVE-2026-68246","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx11): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx11)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx11: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68246","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68249","cve":"CVE-2026-68249","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.0): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.0)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma5.0: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68249","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68250","cve":"CVE-2026-68250","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.2): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma5.2)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma5.2: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68250","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68251","cve":"CVE-2026-68251","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma6.0): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma6.0)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma6.0: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68251","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68252","cve":"CVE-2026-68252","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.0): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/sdma7.0)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/sdma7.0: replace BUG_ON() with WARN_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68252","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68256","cve":"CVE-2026-68256","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A memory or reference-count leak in the amdgpu display core (DC/DM). Each pass through the affected path drops an allocation or a refcount on the floor. A tenant that loops the operation drives the node into memory exhaustion or pins objects that can never be freed, which on a long-lived GPU host shows up as creeping unreclaimable memory, failed allocations for other tenants, and eventually an OOM kill or a driver that will not unbind. Refcount leaks that wrap can also degrade into use-after-free. Upstream fix: drm/amd/display: detect_link_and_local_sink: DP alt mode timeout path leaks prev_sink reference","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68256","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68259","cve":"CVE-2026-68259","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: Check bounds in allocate_event_notification_slot","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68259","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68272","cve":"CVE-2026-68272","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Missing or insufficient validation of user-supplied parameters in the amdgpu GEM/VM/command-submission ioctl surface. A value that crosses the ioctl boundary - a size, a count, an offset, a buffer-object mapping range - is trusted rather than checked, so a tenant can drive the driver outside the range its authors assumed. Where the unchecked value indexes or sizes a kernel allocation this is a memory-corruption primitive and therefore a host-compromise route out of a GPU container; where it only reaches a sanity check further down it costs the node a crash. Upstream fix: drm/amdgpu: validate CP_GFX_SHADOW chunk size in CS pass1","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68272","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68275","cve":"CVE-2026-68275","aliases":[],"title":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu): A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu GEM/VM/command-submission ioctl surface (drm/amdgpu)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A NULL pointer dereference in the amdgpu GEM/VM/command-submission ioctl surface. An unchecked pointer - typically an optional IP block, an absent connector, or a failed allocation - is dereferenced on an error or corner-case path, panicking the kernel. There is no data disclosure here, but on a shared GPU node the blast radius is the whole machine: the panic kills every tenant's job on that host, not just the one that triggered it, and long training runs lose everything since the last checkpoint. Upstream fix: drm/amdgpu: check amdgpu_vm_bo_find() result in GET_MAPPING_INFO","attack_vector":"Local. Reachable by any process with /dev/dri/renderD* open - the render node is handed to tenant containers by every GPU device plugin, so this is unprivileged-tenant reachable. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68275","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68276","cve":"CVE-2026-68276","aliases":[],"title":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx): An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu firmware, ACPI and IP-block initialisation (drm/amdgpu/gfx)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An out-of-bounds access in the amdgpu firmware, ACPI and IP-block initialisation - a length, index or size supplied through the driver interface is used without a proper bound check. Depending on the path this reads kernel memory back to the caller (leaking whatever happens to sit past the buffer, including other tenants' data still resident in the slab) or writes past the allocation, which is a kernel heap-corruption and privilege-escalation primitive. Upstream fix: drm/amdgpu/gfx: fix cleaner shader IB buffer overflow","attack_vector":"Local. Reachable by a local user with driver-load influence, or an attacker who controls platform ACPI tables / VBIOS content - usually root or firmware-level access rather than a tenant. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68276","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68347","cve":"CVE-2026-68347","aliases":[],"title":"Linux iommu/amd - IRQ-unsafe locking in guest domain allocation: An IRQ-unsafe lock taken during AMD IOMMU guest domain allocation can deadlock the host. Guest domain…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux iommu/amd - IRQ-unsafe locking in guest domain allocation","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An IRQ-unsafe lock taken during AMD IOMMU guest domain allocation can deadlock the host. Guest domain allocation happens when you attach a device to a VM - so this fires on the GPU passthrough path, wedging the host at exactly the moment you are provisioning a tenant's accelerator.","attack_vector":"Local, on the IOMMU guest-domain allocation path - reachable by whatever attaches devices to guests, i.e. the VMM or the orchestrator.","remediation":"Fixed in the Linux kernel. Distro kernel update plus reboot; no firmware step.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68347"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68364","cve":"CVE-2026-68364","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A race condition or locking defect in the amdgpu display core (DC/DM). Concurrent paths touch shared state without the right serialisation, so the outcome depends on timing an attacker can influence by hammering the interface from several threads. The visible symptom is a deadlock or hang that wedges the GPU and any job on it; the worse outcome, when the race lands on an object lifetime, is memory corruption. Upstream fix: drm/amd/display: Fix ISM dc_lock deadlock during suspend","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68364","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68430","cve":"CVE-2026-68430","aliases":[],"title":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx8): A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu RAS / GPU reset and recovery path (drm/amdgpu/gfx8)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu RAS / GPU reset and recovery path reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdgpu/gfx8: drop unecessary BUG_ON()","attack_vector":"Local. Reachable by a local user who can trigger or observe a GPU reset, plus anything that can induce ECC/RAS events; some paths are only reachable by the node's own error handling. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68430","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-68436","cve":"CVE-2026-68436","aliases":[],"title":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display): A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Linux kernel amdgpu display core (DC/DM) (drm/amd/display)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdgpu display core (DC/DM) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amd/display: use kvzalloc to allocate struct dc","attack_vector":"Local. Reachable by a local user with access to the DRM primary node, or an attacker who controls the attached display's EDID/DisplayPort topology. On headless Instinct nodes the display block is largely unused, which cuts real exposure sharply. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-68436","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72140","cve":"CVE-2026-72140","aliases":["i2c: mlxbf use-after-free in mlxbf_i2c_init_resource()"],"title":"Linux kernel i2c-mlxbf (BlueField DPU I2C controller): mlxbf_i2c_init_resource() frees a resource struct and then reads a field out of it to build the error code…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux kernel i2c-mlxbf (BlueField DPU I2C controller)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"mlxbf_i2c_init_resource() frees a resource struct and then reads a field out of it to build the error code - a use-after-free on the DPU during I2C controller initialisation. Same subsystem as the earlier BlueField I2C stack overflow, which is the real signal: the DPU's platform glue has repeatedly shipped memory-safety bugs, and that is the layer your infrastructure services sit on.","attack_vector":"Local on the DPU, reached on the I2C init error path during driver probe.","remediation":"Upgrade the DPU Arm-side kernel to 7.2 or one of the wide stable backports (5.10.261 through 7.1.5), delivered as a DOCA/BFOS package upgrade or BFB re-image plus DPU reset. Low urgency alone - bundle into the next BFB refresh alongside the other BlueField platform fixes.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72140","https://git.kernel.org/pub/scm/linux/security/vulns.git/tree/cve/published/2026/CVE-2026-72140.json"],"status":"curated"},{"id":"CVE-2026-72237","cve":"CVE-2026-72237","aliases":[],"title":"Linux perf/x86/amd/brs - kernel address leakage through Branch Sampling: MULTI-TENANT ISOLATION: A user-only branch stack collected via AMD Branch Sampling can contain branches that…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux perf/x86/amd/brs - kernel address leakage through Branch Sampling","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: A user-only branch stack collected via AMD Branch Sampling can contain branches that originated in the kernel, so unprivileged profiling leaks kernel addresses. That is a KASLR break handed to any tenant allowed to profile their own code - and on AI clusters, letting tenants profile their own GPU and CPU kernels is a feature people ask for. Combine it with any of the amdgpu memory-safety bugs in this database and you have a reliable local privilege escalation.","attack_vector":"Local, from a process permitted to use perf branch sampling. How reachable this is depends entirely on your perf_event_paranoid setting - if you relaxed it so tenants can profile, you granted this.","remediation":"Fixed in the Linux kernel by filtering kernel branches out of user-only branch stacks. Distro kernel update plus reboot; no firmware step. In the meantime, review perf_event_paranoid on GPU nodes: the value that makes tenant profiling work is the same value that exposes this.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72237"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-72325","cve":"CVE-2026-72325","aliases":[],"title":"Linux perf/x86/amd/core - Branch Sampling enabled from the SVM reload path: Branch Sampling and Last Branch Record are mutually exclusive, and the KVM/SVM reload path could enable BRS…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux perf/x86/amd/core - Branch Sampling enabled from the SVM reload path","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Branch Sampling and Last Branch Record are mutually exclusive, and the KVM/SVM reload path could enable BRS anyway. Conflicting branch-tracing state on a virtualisation host means corrupted profiling data and, on the SVM path, incorrect state restored around guest entry - a correctness problem sitting on the boundary between host and guest execution.","attack_vector":"Local, on hosts running KVM guests with branch sampling in use.","remediation":"Distro kernel update plus reboot.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-72325"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74353","cve":"CVE-2026-74353","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"A correctness defect in the amdkfd (KFD compute driver, /dev/kfd) reachable through the driver's user-facing interface. The practical effect is an unstable or crashing GPU driver on a shared node. Upstream fix: drm/amdkfd: always resume_all after suspend_all","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74353","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"CVE-2026-74448","cve":"CVE-2026-74448","aliases":[],"title":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd): MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdkfd…","layer":"kernel-hypervisor","layer_name":"Kernel, userspace & hypervisor","component":"Linux kernel amdkfd (KFD compute driver, /dev/kfd) (drm/amdkfd)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Memory is handed to a consumer without being initialised or cleared in the amdkfd (KFD compute driver, /dev/kfd). Whatever the previous owner left behind is readable - and on a GPU node the previous owner is very often a different tenant's job. This is the classic residual-data leak between workloads sharing a card: model weights, activations, keys or tokens from the prior tenant can surface in a fresh allocation. Upstream fix: drm/amdkfd: fix QID bit leak in pqm_create_queue()","attack_vector":"Local. Reachable by any process or container with /dev/kfd and /dev/dri/renderD* mapped in - which is every ROCm workload, including an unprivileged tenant pod. Not reachable over the network and not reachable from a container that has no GPU device node mapped in.","remediation":"Kernel-side fix: this lands in mainline Linux and flows into distro kernels (RHEL/Rocky, Ubuntu HWE, SLES) and into AMD's out-of-tree DKMS amdgpu package shipped with ROCm. Patch the kernel or the DKMS module, then **reload the amdgpu module or reboot the node** - you cannot fix a running driver in place. Reloading amdgpu requires no process holding /dev/kfd or a render node, so in practice this is a cordon + drain + reboot per node. Plan it as a rolling maintenance across the fleet; there is no VBIOS flash, no SBIOS/AGESA step and no firmware update involved. Nodes running the ROCm DKMS stack often lag mainline by a release or two, so confirm the fix is actually present in the AMD driver version you deploy rather than assuming a new distro kernel covers it. Until the reboot window, the only real mitigation is to stop handing the render node to untrusted workloads - the device plugin has to be mapping /dev/dri/renderD* and /dev/kfd into the container for a tenant to reach this at all.","references":["https://nvd.nist.gov/vuln/detail/CVE-2026-74448","https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-0000-001-aspeed-bmc-host-to-bmc-bridges-g","cve":null,"aliases":[],"title":"ASPEED BMC (host-to-BMC bridges generally): The ASPEED LPC/PCIe bridge architecture exists to let the host talk to the BMC","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED BMC (host-to-BMC bridges generally)","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The ASPEED LPC/PCIe bridge architecture exists to let the host talk to the BMC; disabling it breaks legitimate in-band management (ipmitool, firmware update tooling). Operators frequently leave it on","attack_vector":"Local, host CPU","remediation":"Explicit per-fleet decision: lock the AHB bridges and lose in-band management tooling, or accept a host→BMC escalation path. There is no configuration that gives both","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-6260"],"status":"curated"},{"id":"NCVD-0000-002-ipmi-over-lan-as-a-protocol","cve":null,"aliases":[],"title":"IPMI over LAN as a protocol: IPMI has no transport confidentiality guarantees worth relying on, weak session handling, and per-vendor…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"IPMI over LAN as a protocol","year":"2013-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"IPMI has no transport confidentiality guarantees worth relying on, weak session handling, and per-vendor implementations that diverge from spec. Every BMC on the fleet speaks it by default","attack_vector":"Network, management VLAN","remediation":"Fleet policy decision to disable IPMI-over-LAN and force Redfish-only; costs a rewrite of provisioning/monitoring tooling and loses compatibility with older ODM chassis","references":["https://nvd.nist.gov/vuln/detail/CVE-2013-4786"],"status":"curated"},{"id":"NCVD-0000-003-internet-exposed-bmc","cve":null,"aliases":[],"title":"Internet-exposed BMC: Shodan-visible BMCs are a recurring finding at colo/neocloud buildouts — a management interface reachable…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Internet-exposed BMC","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Shodan-visible BMCs are a recurring finding at colo/neocloud buildouts — a management interface reachable from the internet turns every BMC CVE above into a directly exploitable, pre-auth, below-OS foothold","attack_vector":"Network, from the public internet","remediation":"Continuous external attack-surface scanning of the management ranges plus an enforced out-of-band network design; the operational cost is that remote-hands and vendor support workflows often depend on the exposure","references":["https://www.shodan.io/search?query=ipmi"],"status":"curated"},{"id":"NCVD-0000-004-infiniband-subnet-manager-opensm","cve":null,"aliases":[],"title":"InfiniBand subnet manager (OpenSM / UFM): The IB subnet manager has unilateral authority over LID assignment, routing and partitioning for the entire…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"InfiniBand subnet manager (OpenSM / UFM)","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The IB subnet manager has unilateral authority over LID assignment, routing and partitioning for the entire fabric, and the management datagram (MAD) path has historically had weak authentication. There is no per-tenant trust boundary in the SM","attack_vector":"Fabric-local","remediation":"Requires an architectural control: pinned partition keys (pkeys) per tenant, SM redundancy, and treating the SM host as tier-0 infrastructure. No patch exists because it is the protocol model","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-0130"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-0000-005-rdma-roce","cve":null,"aliases":["ReDMArk"],"title":"RDMA / RoCE: RoCE and IB RDMA have no cryptographic authentication of the QP connection setup or of subsequent RDMA…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RDMA / RoCE","year":"2021-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"RoCE and IB RDMA have no cryptographic authentication of the QP connection setup or of subsequent RDMA reads/writes; an on-fabric attacker can inject and impersonate, and remote memory reads bypass the target CPU entirely. No CVE — it is the RDMA specification","attack_vector":"Fabric-local, tenant-to-tenant","remediation":"Requires fabric-level isolation (per-tenant pkeys / VXLAN-isolated RoCE domains) or PSP/IPsec offload on the NIC. Shared-fabric multi-tenancy without this is an unmitigated tenant-to-tenant read primitive","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/rothenberger"],"status":"curated"},{"id":"NCVD-0000-006-nvme-of-over-rdma","cve":null,"aliases":["NeVerMore"],"title":"NVMe-oF over RDMA: NVMe-over-Fabrics inherits RDMA's lack of authentication","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe-oF over RDMA","year":"2022-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"NVMe-over-Fabrics inherits RDMA's lack of authentication; storage targets are addressable by any host on the fabric, so a compromised tenant node can reach other tenants' namespaces. No CVE — design","attack_vector":"Fabric-local","remediation":"Enforce NVMe-oF host NQN allowlisting plus DH-HMAC-CHAP authentication and separate storage fabric; costs throughput and adds provisioning complexity","references":["https://arxiv.org/abs/2202.08080"],"status":"curated"},{"id":"NCVD-0000-007-facility-power-dcim-as-a-class","cve":null,"aliases":[],"title":"Facility power / DCIM as a class: PDUs, CRAC controllers, BMS and DCIM platforms run long-lived embedded firmware, sit on flat facility…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Facility power / DCIM as a class","year":"2023-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"PDUs, CRAC controllers, BMS and DCIM platforms run long-lived embedded firmware, sit on flat facility networks, and are frequently outside the compute operator's change-control. An attacker with PDU control has a physical-availability weapon against a GPU cluster","attack_vector":"Network, facility LAN","remediation":"Contractual and architectural: demand facility-network segmentation and a firmware SLA from the colo provider, and monitor the power plane independently. For a tenant in someone else's datacenter there is no patch you can apply yourself","references":["https://thehackernews.com/2023/08/multiple-flaws-in-cyberpower-and.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-0000-008-firmware-signing-key-compromise","cve":null,"aliases":[],"title":"Firmware signing-key compromise as a class: Firmware trust anchors (Boot Guard KM/BPM, UEFI PK/KEK, BMC image-signing keys) are held by ODMs with weaker…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Firmware signing-key compromise as a class","year":"2023-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Firmware trust anchors (Boot Guard KM/BPM, UEFI PK/KEK, BMC image-signing keys) are held by ODMs with weaker security postures than the clouds that deploy their hardware, and several have been breached. A neocloud inherits the ODM's key hygiene","attack_vector":"Supply chain","remediation":"Requires diligence on ODM key custody at purchase time and an independent measured-boot / firmware-integrity baseline so a signed-but-malicious image is still detectable. Cannot be remediated after the fact","references":["https://www.binarly.io/blog/pkfail-untrusted-platform-keys-undermine-secure-boot-on-uefi-ecosystem"],"status":"curated"},{"id":"NCVD-0000-009-redfish-implementations-all-vend","cve":null,"aliases":[],"title":"Redfish implementations (all vendors): Redfish replaced IPMI but reintroduced the same class of flaws at the HTTP layer — CVE-2024-54085…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Redfish implementations (all vendors)","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Redfish replaced IPMI but reintroduced the same class of flaws at the HTTP layer — CVE-2024-54085, CVE-2023-34330, CVE-2023-25191/25192, CVE-2018-15774 are all Redfish-surface bugs. The spec mandates no minimum implementation assurance, and every ODM ships its own stack","attack_vector":"Network / management VLAN","remediation":"Treat Redfish as an untrusted-by-default surface: mTLS or a management-plane proxy in front of every BMC, per-node unique credentials, and no direct operator access. This is architecture work, not patching","references":["https://nvd.nist.gov/vuln/detail/CVE-2024-54085"],"status":"curated"},{"id":"NCVD-0000-010-kvm-over-ip-virtual-media","cve":null,"aliases":[],"title":"KVM-over-IP / virtual media: The BMC's virtual-media function can mount an arbitrary ISO as the host's boot device. Any BMC compromise…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"KVM-over-IP / virtual media","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"The BMC's virtual-media function can mount an arbitrary ISO as the host's boot device. Any BMC compromise therefore converts directly into arbitrary host boot, bypassing disk encryption and OS controls (CVE-2019-16649 is the concrete instance)","attack_vector":"Network / BMC","remediation":"Disable virtual media in the BMC baseline except during provisioning windows, and gate the KVM/vmedia ports at the management-network edge. Costs the remote-hands workflow that most operations teams rely on","references":["https://nvd.nist.gov/vuln/detail/CVE-2019-16649"],"status":"curated"},{"id":"NCVD-0000-011-serial-console-servers-out-of-ba","cve":null,"aliases":[],"title":"Serial console servers / out-of-band access appliances: Console servers (Opengear, Lantronix, Digi and similar) hold plaintext-equivalent access to every device's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Serial console servers / out-of-band access appliances","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Console servers (Opengear, Lantronix, Digi and similar) hold plaintext-equivalent access to every device's console including switches and BMCs, run long-lived embedded Linux, and are the deliberate bypass path around all network segmentation","attack_vector":"Network / OOB LAN","remediation":"Bring the console-server fleet into the same patch and credential-rotation cadence as the compute; require per-device authentication rather than a shared console password. Frequently owned by the facility, not the operator","references":["https://nvd.nist.gov/vuln/detail/CVE-2011-3997"],"status":"curated"},{"id":"NCVD-0000-012-gpu-accelerator-firmware-vbios-g","cve":null,"aliases":[],"title":"GPU / accelerator firmware (VBIOS, GSP, NVSwitch): GPU-resident firmware sits below the host OS and is not covered by host EDR or reimaging. NVIDIA ships GPU…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"GPU / accelerator firmware (VBIOS, GSP, NVSwitch)","year":"2023-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"GPU-resident firmware sits below the host OS and is not covered by host EDR or reimaging. NVIDIA ships GPU firmware fixes through opaque bundled updates, so a neocloud cannot independently verify what changed or attest the running image between tenants","attack_vector":"Local, from a tenant with device access","remediation":"Enforce a firmware re-flash and attestation step at every tenant handoff rather than trusting a host reimage. NVIDIA's bundle model means the operator cannot audit the fix, only apply it","references":["https://www.nvidia.com/en-us/security/"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-0000-013-tenant-handoff-on-bare-metal","cve":null,"aliases":[],"title":"Tenant handoff on bare metal: Reimaging the host disk clears nothing in the BMC, UEFI/SPI flash, NIC/DPU firmware, GPU VBIOS or the ESP.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Tenant handoff on bare metal","year":"2019-2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Reimaging the host disk clears nothing in the BMC, UEFI/SPI flash, NIC/DPU firmware, GPU VBIOS or the ESP. Every below-OS CVE in this table is therefore also a cross-tenant persistence primitive on any bare-metal GPU offering","attack_vector":"Local, from a prior tenant","remediation":"Requires a full firmware re-provision and remote attestation between tenants — measured boot with verified PCRs, BMC reflash, NIC firmware verification. This is the single largest unpriced operational cost in bare-metal GPU rental","references":["https://www.binarly.io/reports/pkfail"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2011-001-ata-secure-erase-nvme-sanitize-f","cve":null,"aliases":["NIST SP 800-88","Reliably Erasing Data From Flash-Based Solid State Drives","sanitize verification gap"],"title":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors - drives that report sanitization success while data remains: CLASS ENTRY and the empirical foundation under every specific CVE in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors - drives that report sanitization success while…","year":"2011","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY and the empirical foundation under every specific CVE in this file. UCSD researchers physically desoldered NAND from twelve SSDs after running the built-in sanitize commands and read the raw flash back. Findings, verbatim in substance: built-in commands are effective but manufacturers sometimes implement them INCORRECTLY; of eight drives claiming ATA SECURITY support only four executed ERASE UNIT reliably; two more silently erased just the first LBA unless the firmware had recently been reset; and one drive REPORTED SANITIZATION SUCCESS WHILE ALL DATA REMAINED INTACT - the filesystem was still mountable afterwards. They also found that overwriting the full visible address space twice is usually but not always sufficient, and that single-file sanitization techniques consistently FAIL on SSDs, because the FTL keeps copies at physical addresses the host cannot address. BREAKS TENANT HANDOFF generically: your 'secure erase between customers' step may silently do nothing, and the drive's success code is not evidence. The next tenant carves the previous tenant's checkpoints, datasets and cloud credentials out of blocks your wipe never reached.","attack_vector":"The next tenant on the reclaimed bare-metal host, reading unallocated or remapped blocks; or anyone who obtains the physical drive later and is willing to read the NAND directly, which is the only method that sees over-provisioned, retired and bad blocks at all.","remediation":"There is no patch - this is a verification and architecture problem. (1) Follow NIST SP 800-88 Rev 1 and pick the sanitization tier that matches the data: for anything that held tenant data, Purge (cryptographic erase or a verified device sanitize) at minimum, and Destroy for high-sensitivity media. (2) Stop trusting return codes. Sample-verify: after sanitize, read back a statistical sample of LBAs and confirm they are zeroed or random, and keep the evidence. This does NOT cover over-provisioned or retired blocks, which no host-side read can reach - be honest in your compliance story about that limit rather than claiming a guarantee you cannot make. (3) The only sanitization that is both fast and verifiable at fleet scale is throwing away a key YOU control: run LUKS/dm-crypt per tenant with the key in your KMS, so reclaim is a key-destruction event you can log and audit in milliseconds, and the drive's own erase behaviour stops mattering. (4) Qualify each SKU's sanitize implementation once, in a lab, before it enters the fleet - the researchers' own conclusion was that every implementation must be individually tested before it can be trusted. Verifying erase across a 10,000-drive fleet is a weeks-long operation that most operators never perform at all, which is exactly why a drive that lies about it survives undetected for years.","references":["https://www.usenix.org/legacy/events/fast11/tech/full_papers/Wei.pdf","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-33082"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2011-002-ata-secure-erase-nvme-sanitize-f","cve":null,"aliases":["NIST SP 800-88","Reliably Erasing Data From Flash-Based Solid State Drives","sanitize verification gap"],"title":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors - drives that report sanitization success while data remains: CLASS ENTRY and the empirical foundation under every specific CVE in…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ATA Secure Erase / NVMe Sanitize / Format NVM across SSD vendors - drives that report sanitization success while…","year":"2011","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY and the empirical foundation under every specific CVE in this file. UCSD researchers physically desoldered NAND from twelve SSDs after running the built-in sanitize commands and read the raw flash back. Findings, verbatim in substance: built-in commands are effective but manufacturers sometimes implement them INCORRECTLY; of eight drives claiming ATA SECURITY support only four executed ERASE UNIT reliably; two more silently erased just the first LBA unless the firmware had recently been reset; and one drive REPORTED SANITIZATION SUCCESS WHILE ALL DATA REMAINED INTACT - the filesystem was still mountable afterwards. They also found that overwriting the full visible address space twice is usually but not always sufficient, and that single-file sanitization techniques consistently FAIL on SSDs, because the FTL keeps copies at physical addresses the host cannot address. BREAKS TENANT HANDOFF generically: your 'secure erase between customers' step may silently do nothing, and the drive's success code is not evidence. The next tenant carves the previous tenant's checkpoints, datasets and cloud credentials out of blocks your wipe never reached.","attack_vector":"The next tenant on the reclaimed bare-metal host, reading unallocated or remapped blocks; or anyone who obtains the physical drive later and is willing to read the NAND directly, which is the only method that sees over-provisioned, retired and bad blocks at all.","remediation":"There is no patch - this is a verification and architecture problem. (1) Follow NIST SP 800-88 Rev 1 and pick the sanitization tier that matches the data: for anything that held tenant data, Purge (cryptographic erase or a verified device sanitize) at minimum, and Destroy for high-sensitivity media. (2) Stop trusting return codes. Sample-verify: after sanitize, read back a statistical sample of LBAs and confirm they are zeroed or random, and keep the evidence. This does NOT cover over-provisioned or retired blocks, which no host-side read can reach - be honest in your compliance story about that limit rather than claiming a guarantee you cannot make. (3) The only sanitization that is both fast and verifiable at fleet scale is throwing away a key YOU control: run LUKS/dm-crypt per tenant with the key in your KMS, so reclaim is a key-destruction event you can log and audit in milliseconds, and the drive's own erase behaviour stops mattering. (4) Qualify each SKU's sanitize implementation once, in a lab, before it enters the fleet - the researchers' own conclusion was that every implementation must be individually tested before it can be trusted. Verifying erase across a 10,000-drive fleet is a weeks-long operation that most operators never perform at all, which is exactly why a drive that lies about it survives undetected for years.","references":["https://www.usenix.org/legacy/events/fast11/tech/full_papers/Wei.pdf","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf","https://nvd.nist.gov/vuln/detail/CVE-2021-33082"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2015-001-hdd-and-ssd-controller-firmware","cve":null,"aliases":["Equation Group","nls_933w.dll","GrayFish","drive firmware implant"],"title":"HDD and SSD controller firmware as a persistence surface - demonstrated against Seagate, Western Digital, Toshiba, Maxtor and IBM drives: CLASS ENTRY. Kaspersky documented a module that reprograms…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HDD and SSD controller firmware as a persistence surface - demonstrated against Seagate, Western Digital, Toshiba…","year":"2015","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY. Kaspersky documented a module that reprograms the microcontroller firmware of drives from at least five major vendors - described as the most powerful tool in that actor's arsenal. Once the implant is in controller firmware it runs below the operating system, survives reformatting and reimaging because those operations only rewrite the data area and never touch the microcontroller, and even reflashing does not reliably remove it: some firmware regions are not covered by an update, and the drive may report it is already on the latest version and refuse. TOP TENANT-HANDOFF RISK IN THIS CATEGORY. A bare-metal tenant with root has, by definition, the host-side access needed to attempt a firmware write over the standard command set. If it lands, your entire reclaim process - wipe, reimage, revalidate, re-rent - is theatre, because the malicious code is in a place none of those steps inspect. The implant sees every block the next tenant writes, and can lie about sanitize, lock state and firmware version to every tool you own.","attack_vector":"A tenant with root on the bare-metal host issuing firmware-write commands over the standard storage admin command set - no physical access needed. Also reachable by anyone in the supply chain or the RMA/decommission path who has the drive in hand. Detection is the hard part: a compromised controller is the thing reporting its own firmware version to you, so host-side attestation of drive firmware is self-referential and unreliable.","remediation":"Assume UNPATCHABLE and UNVERIFIABLE once suspected - you cannot trust a controller's self-report of its own integrity. Controls are preventative and procedural. (1) Deny tenants the ability to write drive firmware at all: block firmware-download/commit and vendor-specific pass-through commands at the hypervisor, IOMMU/VFIO policy or storage-controller layer, and do not hand tenants unfiltered raw block devices where the workload does not require it. (2) Prefer drives that enforce signed firmware, and confirm the vendor's signing story in procurement rather than assuming it. (3) Record expected firmware version per drive serial in a system the host cannot edit, and alert on any change - especially a version that goes backwards. (4) For high-sensitivity tenancies, retire local media at end of tenancy rather than recycling it into the pool - physical destruction is the only remediation with a guarantee attached, and it costs one drive versus an undetectable persistent foothold across every future customer on that node. (5) Push the risk off the local drive entirely: keep tenant data on network storage with operator-held encryption so that a compromised local controller sees only ciphertext and transient scratch.","references":["https://securelist.com/equation-the-death-star-of-malware-galaxy/68750/","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2015-002-hdd-and-ssd-controller-firmware","cve":null,"aliases":["Equation Group","nls_933w.dll","GrayFish","drive firmware implant"],"title":"HDD and SSD controller firmware as a persistence surface - demonstrated against Seagate, Western Digital, Toshiba, Maxtor and IBM drives: CLASS ENTRY. Kaspersky documented a module that reprograms…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HDD and SSD controller firmware as a persistence surface - demonstrated against Seagate, Western Digital, Toshiba…","year":"2015","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY. Kaspersky documented a module that reprograms the microcontroller firmware of drives from at least five major vendors - described as the most powerful tool in that actor's arsenal. Once the implant is in controller firmware it runs below the operating system, survives reformatting and reimaging because those operations only rewrite the data area and never touch the microcontroller, and even reflashing does not reliably remove it: some firmware regions are not covered by an update, and the drive may report it is already on the latest version and refuse. TOP TENANT-HANDOFF RISK IN THIS CATEGORY. A bare-metal tenant with root has, by definition, the host-side access needed to attempt a firmware write over the standard command set. If it lands, your entire reclaim process - wipe, reimage, revalidate, re-rent - is theatre, because the malicious code is in a place none of those steps inspect. The implant sees every block the next tenant writes, and can lie about sanitize, lock state and firmware version to every tool you own.","attack_vector":"A tenant with root on the bare-metal host issuing firmware-write commands over the standard storage admin command set - no physical access needed. Also reachable by anyone in the supply chain or the RMA/decommission path who has the drive in hand. Detection is the hard part: a compromised controller is the thing reporting its own firmware version to you, so host-side attestation of drive firmware is self-referential and unreliable.","remediation":"Assume UNPATCHABLE and UNVERIFIABLE once suspected - you cannot trust a controller's self-report of its own integrity. Controls are preventative and procedural. (1) Deny tenants the ability to write drive firmware at all: block firmware-download/commit and vendor-specific pass-through commands at the hypervisor, IOMMU/VFIO policy or storage-controller layer, and do not hand tenants unfiltered raw block devices where the workload does not require it. (2) Prefer drives that enforce signed firmware, and confirm the vendor's signing story in procurement rather than assuming it. (3) Record expected firmware version per drive serial in a system the host cannot edit, and alert on any change - especially a version that goes backwards. (4) For high-sensitivity tenancies, retire local media at end of tenancy rather than recycling it into the pool - physical destruction is the only remediation with a guarantee attached, and it costs one drive versus an undetectable persistent foothold across every future customer on that node. (5) Push the risk off the local drive entirely: keep tenant data on network storage with operator-held encryption so that a compromised local controller sees only ciphertext and transient scratch.","references":["https://securelist.com/equation-the-death-star-of-malware-galaxy/68750/","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2018-001-amd-sev-memory-encryption-host-c","cve":null,"aliases":["SEVered"],"title":"AMD SEV memory encryption - host-controlled guest physical to host physical mapping: The original demonstration that SEV does not protect against the host: the hypervisor remaps a guest's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV memory encryption - host-controlled guest physical to host physical mapping","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"The original demonstration that SEV does not protect against the host: the hypervisor remaps a guest's physical pages while the guest is servicing a network request, and reads the plaintext of the entire guest memory out of the responses the guest itself sends back. No CVE was ever assigned - AMD's position was that SEV was not designed to resist this - which is why it is easy to miss when auditing. It matters historically and commercially: it is the reason SEV-ES and then SEV-SNP exist, and it means any 'confidential computing' claim made on plain SEV hardware was never true against the operator.","attack_vector":"Malicious or compromised hypervisor against a SEV guest, using ordinary VMM control over second-level page tables plus any network service running in the guest. No exploit primitive needed beyond normal host capabilities.","remediation":"UNPATCHABLE on SEV. The architectural fix is SEV-SNP with the Reverse Map Table (Milan and later) and it is a hardware generation, not a firmware update. Practical guidance for an operator: do not market plain SEV or SEV-ES as protection from yourself, check that SNP is actually enabled rather than merely supported on the SKU, and make sure your attestation flow proves SNP is on rather than proving only that the CPU could do it.","references":["https://www.amd.com/en/developer/sev.html","https://arxiv.org/abs/1805.09604"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2018-002-ecc-ddr3-server-memory-on-intel","cve":null,"aliases":["ECCploit"],"title":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms; the technique generalises to later ECC DRAM: ECC is the answer most operators give when asked about…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms; the technique generalises to…","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"ECC is the answer most operators give when asked about Rowhammer, and ECCploit is why that answer is wrong. Correcting a flip takes measurably longer than a clean read, so the attacker gets a timing side channel that tells them exactly which bits they flipped - turning ECC from a defence into a feedback oracle for template building. With that feedback they place three flips in one word, which ECC neither corrects nor detects, producing silent corruption. For an AI datacenter this is the nastiest variant: the corruption is by construction invisible to the machine-check path, so your fleet health dashboard shows green while a tenant's memory is being rewritten.","attack_vector":"Unprivileged local code on an ECC server sharing DRAM with the victim. Attack time was about 32 minutes when corrections were directly observable and up to a week in noisy production-like conditions - slow, but a long-running batch tenant has that time.","remediation":"Do not treat ECC as a Rowhammer mitigation; treat it as error reporting. Make sure the reporting is actually wired up - EDAC or the equivalent collecting correctable-error counts per DIMM, exported to your monitoring, with alerting on rate rather than absolute count, since a burst of corrections on one rank is the strongest hammering signal you will get. Confirm firmware and OS handle uncorrectable errors by isolating rather than silently continuing. Retire DIMM SKUs that show elevated correctable-error rates. The structural fix is the same as every other entry here: do not share a memory controller between untrusted tenants.","references":["https://www.vusec.net/projects/eccploit/","https://download.vusec.net/papers/eccploit_sp19.pdf"],"status":"curated"},{"id":"NCVD-2018-004-ecc-ddr3-server-memory-on-intel","cve":null,"aliases":["ECCploit"],"title":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms; the technique generalises to later ECC DRAM: ECC is the answer most operators give when asked about…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ECC DDR3 server memory on Intel Xeon (Haswell, Sandy Bridge) and AMD Opteron platforms; the technique generalises to…","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"ECC is the answer most operators give when asked about Rowhammer, and ECCploit is why that answer is wrong. Correcting a flip takes measurably longer than a clean read, so the attacker gets a timing side channel that tells them exactly which bits they flipped - turning ECC from a defence into a feedback oracle for template building. With that feedback they place three flips in one word, which ECC neither corrects nor detects, producing silent corruption. For an AI datacenter this is the nastiest variant: the corruption is by construction invisible to the machine-check path, so your fleet health dashboard shows green while a tenant's memory is being rewritten.","attack_vector":"Unprivileged local code on an ECC server sharing DRAM with the victim. Attack time was about 32 minutes when corrections were directly observable and up to a week in noisy production-like conditions - slow, but a long-running batch tenant has that time.","remediation":"Do not treat ECC as a Rowhammer mitigation; treat it as error reporting. Make sure the reporting is actually wired up - EDAC or the equivalent collecting correctable-error counts per DIMM, exported to your monitoring, with alerting on rate rather than absolute count, since a burst of corrections on one rank is the strongest hammering signal you will get. Confirm firmware and OS handle uncorrectable errors by isolating rather than silently continuing. Retire DIMM SKUs that show elevated correctable-error rates. The structural fix is the same as every other entry here: do not share a memory controller between untrusted tenants.","references":["https://www.vusec.net/projects/eccploit/","https://download.vusec.net/papers/eccploit_sp19.pdf"],"status":"curated"},{"id":"NCVD-2018-004-microsoft-bitlocker-windows-edri","cve":null,"aliases":["ADV180028","BitLocker hardware encryption trusted by default","eDrive offload"],"title":"Microsoft BitLocker / Windows eDrive hardware-encryption offload on any TCG Opal or IEEE-1667 self-encrypting drive: BitLocker's default was to hand encryption to the drive whenever the drive…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Microsoft BitLocker / Windows eDrive hardware-encryption offload on any TCG Opal or IEEE-1667 self-encrypting drive","year":"2018","cvss_score":null,"severity":"unscored","kev":false,"impact":"BitLocker's default was to hand encryption to the drive whenever the drive claimed to support it, and to then skip software encryption entirely. Every SED firmware weakness therefore became a full BitLocker bypass, silently, on machines whose operators believed they were encrypted. For a GPU operator this is the general lesson in its sharpest form: a self-attested hardware security claim was accepted with no verification, so the whole encryption posture of the fleet inherited the weakest drive firmware in it. BREAKS TENANT HANDOFF wherever Windows bare-metal nodes are re-let, and it fails silently - the management console reports the volume as encrypted and compliant the entire time.","attack_vector":"Anyone who obtains the physical drive from a Windows node that used hardware offload - next tenant, RMA path, decommission channel. The attacker exploits whatever SED firmware flaw the drive has; BitLocker simply removed the software layer that would have stopped them.","remediation":"Policy change, no firmware needed, but a full re-encrypt: set Group Policy 'Configure use of hardware-based encryption for fixed/operating system data drives' to Disabled, then fully decrypt and re-encrypt each volume - toggling the policy alone does NOT re-encrypt already-provisioned disks, which is the step operators most often miss and which leaves the fleet reporting compliant while still using drive crypto. Verify per host with 'manage-bde -status' and confirm the encryption method is a software AES-XTS value, not 'Hardware Encryption'. Budget a full re-encrypt window per node; on a large Windows bare-metal estate this is a rolling multi-week drain-and-re-image campaign, and there is no way to sample it - a node you did not re-encrypt is a node still relying on the drive.","references":["https://msrc.microsoft.com/update-guide/vulnerability/ADV180028","https://kb.cert.org/vuls/id/395981","https://www.ru.nl/en/research/research-news/radboud-university-researchers-discover-security-flaw-in-ssd-hard-drives"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2019-001-aspeed-ast2400-ast2500-ast2600-p","cve":null,"aliases":["P2A bridge default-on","iLPC2AHB","X-DMA","pantsdown attack surface"],"title":"ASPEED AST2400 / AST2500 / AST2600 (PCIe VGA P2A bridge, iLPC2AHB, X-DMA, SoC debug UART): The design-level problem behind pantsdown, and it outlives the patch. ASPEED silicon deliberately exposes…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED AST2400 / AST2500 / AST2600 (PCIe VGA P2A bridge, iLPC2AHB, X-DMA, SoC debug UART)","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"The design-level problem behind pantsdown, and it outlives the patch. ASPEED silicon deliberately exposes several host-side windows into the BMC's own AHB address space: a PCIe VGA peer-to-peer (P2A) aperture, the LPC-to-AHB bridge reachable from the host's SuperIO, X-DMA, and a UART debug console. On a lot of shipped firmware these are left open because ODM tooling, vendor flashing utilities and iKVM features rely on them. A tenant that gets kernel or root on a bare-metal GPU node - or anything that reaches the host PCIe config space, which includes a rogue driver in a passthrough VM on some configurations - can write the BMC's RAM and flash directly. That is a host-to-BMC escalation with no BMC credentials and no management-network access, and it yields an implant on a processor that keeps running when the node is powered off and survives OS reimage entirely.","attack_vector":"Local to the host: code with kernel privilege on the server the BMC is attached to, or PCIe config-space access from a passthrough device. No network path to the BMC needed. On a bare-metal GPU rental fleet, this is any tenant who gets root on the box they rented.","remediation":"Not a single patch - a per-platform hardening audit. The BMC firmware must explicitly lock the P2A bridge, disable the iLPC2AHB path via the SuperIO/eSPI configuration, and clear the SoC debug UART enable at boot. OpenBMC upstream added kernel-side gating but ODM builds frequently re-enable pieces for their own flashing flows, so you cannot assume your image is safe because it is 'post-2019'. Verifying requires reading the BMC's own register state per platform, then a BMC firmware flash to fix - out-of-band, per node, ODM-rebase-lagged, and bricking risk if power is lost mid-write. Config-only partial mitigation: on a bare-metal fleet, wipe and re-flash BMC firmware between tenants rather than trusting the running image, and treat any node that has hosted an untrusted tenant as having a potentially dirty BMC.","references":["https://www.flamingspork.com/blog/2019/01/23/cve-2019-6260:-gaining-control-of-bmc-from-the-host-processor/","https://security.netapp.com/advisory/ntap-20190314-0001/","https://www.theregister.com/2019/01/24/bmc_pantsdown_bug/","https://support.lenovo.com/us/en/solutions/ps500241-aspeed-ast-series-bmc-vulnerability"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2019-001-pcie-address-translation-service","cve":null,"aliases":["Thunderclap","PCIe ATS translated-request trust","IOMMU bypass via Address Translation Services"],"title":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable endpoint the host trusts, including GPUs, NICs and DPUs: The IOMMU/SMMU is the entire…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"The IOMMU/SMMU is the entire basis for saying a passed-through GPU or NIC cannot read host or other-tenant memory. ATS undermines it by design: a device marked ATS-capable is trusted to have already translated an address, and the root complex forwards those requests without re-checking them. A malicious or compromised endpoint simply sets the translated bit and issues DMA at any physical address it likes. Thunderclap showed the broader problem too - even without ATS, real OS IOMMU policies map far more than the buffer in question, leaving windows onto adjacent kernel memory. In a GPU-passthrough fleet the consequence is a tenant with device control reading host memory and other tenants' data, which defeats the isolation model the whole product rests on.","attack_vector":"An attacker who controls a DMA-capable PCIe device: a tenant with GPU or NIC passthrough who can flash or exploit device firmware, an attacker with physical access to a slot or an external PCIe/Thunderbolt port, or a supply-chain-modified card. Not reachable from software alone on a well-configured host - the entry point is device control.","remediation":"There is no single patch; this is configuration you must actively verify, and most fleets have never checked. Concretely: disable ATS unless a workload genuinely needs it (`pci=noats` on Linux, or the equivalent BIOS switch), and never leave ATS enabled for a device assigned to a tenant. Confirm PCIe ACS is enabled on every upstream port and switch so peer-to-peer traffic between passed-through devices is forced up through the IOMMU rather than routed directly - many server BIOSes ship ACS off, and several 'GPU peer-to-peer performance' tuning guides tell you to turn it off, which silently deletes the isolation boundary between two tenants' GPUs on the same switch. Verify per-device IOMMU groups are not lumping unrelated functions together. Then decide deliberately whether the peer-to-peer bandwidth you gain by disabling ACS is worth the tenant-isolation guarantee you lose, and write that decision down.","references":["http://thunderclap.io/","http://thunderclap.io/thunderclap-paper-ndss2019.pdf","https://www.kernel.org/doc/html/latest/admin-guide/kernel-parameters.html"],"status":"curated"},{"id":"NCVD-2019-004-hpe-sas-ssds-20-models-incl-vo04","cve":null,"aliases":["HPE 32,768-hour SSD bug","HPE bulletin a00092491en_us","HPD8"],"title":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL) with firmware prior to HPD8 - 16-bit power-on-hours counter overflow: PHYSICAL /…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL) with…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL / FLEET-WIDE AVAILABILITY EVENT. At exactly 32,768 power-on hours - about 3 years, 270 days - the drive fails permanently. HPE states that neither the drive nor the data on it can be recovered. The killer property is SYNCHRONY: drives bought and racked together cross the threshold together, so RAID redundancy provides no protection because you lose several members at once. One administrator reported eight drives failing within MINUTES of each other. For a GPU operator this is the whole-cluster failure mode nothing in your HA design accounts for: your redundancy assumes independent failures, and a firmware counter overflow makes them perfectly correlated. Not an attack - a latent time bomb sitting in the fleet with a known detonation date, which is exactly why it belongs in an operator vulnerability database.","attack_vector":"No attacker. The trigger is elapsed powered-on time, and every affected drive with a similar install date reaches it simultaneously. The exposure is determined entirely by your procurement and racking history - a bulk purchase deployed in one window is the worst case.","remediation":"Flash to firmware HPD8 or later BEFORE the threshold; after the drive fails there is no recovery and you restore from backup. HPE shipped HPD8 for the first eight models from 22 November 2019 and the remaining twelve in mid-December 2019. The urgent operational step is inventory, not patching: pull power-on hours for every SAS SSD in the fleet (HPE Smart Storage Administrator, or smartctl attribute 9) and sort by hours remaining, because your window is defined by the oldest drives. Across a large estate this is a rolling drain-and-flash campaign of weeks, and it must be sequenced by remaining hours rather than by rack - and critically you must STAGGER it, because flashing a whole batch on one day recreates the correlated-failure problem for the next latent bug. Standing control: alert on power-on-hours thresholds fleet-wide, and deliberately mix drive batches and vendors across redundancy groups so a single firmware defect cannot take out every member of a RAID set at once.","references":["https://www.bleepingcomputer.com/news/hardware/hp-warns-that-some-ssd-drives-will-fail-at-32-768-hours-of-use/","https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na-a00092491en_us"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2019-005-hpe-sas-ssds-20-models-incl-vo04","cve":null,"aliases":["HPE 32,768-hour SSD bug","HPE bulletin a00092491en_us","HPD8"],"title":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL) with firmware prior to HPD8 - 16-bit power-on-hours counter overflow: PHYSICAL /…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE SAS SSDs (20 models incl. VO0480JFDGT, VO0960JFDGU, VO1920JFDGV, VO3840JFDHA, MO0400JFFCF-MO3200JFFCL) with…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL / FLEET-WIDE AVAILABILITY EVENT. At exactly 32,768 power-on hours - about 3 years, 270 days - the drive fails permanently. HPE states that neither the drive nor the data on it can be recovered. The killer property is SYNCHRONY: drives bought and racked together cross the threshold together, so RAID redundancy provides no protection because you lose several members at once. One administrator reported eight drives failing within MINUTES of each other. For a GPU operator this is the whole-cluster failure mode nothing in your HA design accounts for: your redundancy assumes independent failures, and a firmware counter overflow makes them perfectly correlated. Not an attack - a latent time bomb sitting in the fleet with a known detonation date, which is exactly why it belongs in an operator vulnerability database.","attack_vector":"No attacker. The trigger is elapsed powered-on time, and every affected drive with a similar install date reaches it simultaneously. The exposure is determined entirely by your procurement and racking history - a bulk purchase deployed in one window is the worst case.","remediation":"Flash to firmware HPD8 or later BEFORE the threshold; after the drive fails there is no recovery and you restore from backup. HPE shipped HPD8 for the first eight models from 22 November 2019 and the remaining twelve in mid-December 2019. The urgent operational step is inventory, not patching: pull power-on hours for every SAS SSD in the fleet (HPE Smart Storage Administrator, or smartctl attribute 9) and sort by hours remaining, because your window is defined by the oldest drives. Across a large estate this is a rolling drain-and-flash campaign of weeks, and it must be sequenced by remaining hours rather than by rack - and critically you must STAGGER it, because flashing a whole batch on one day recreates the correlated-failure problem for the next latent bug. Standing control: alert on power-on-hours thresholds fleet-wide, and deliberately mix drive batches and vendors across redundancy groups so a single firmware defect cannot take out every member of a RAID set at once.","references":["https://www.bleepingcomputer.com/news/hardware/hp-warns-that-some-ssd-drives-will-fail-at-32-768-hours-of-use/","https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na-a00092491en_us"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2019-005-pcie-address-translation-service","cve":null,"aliases":["Thunderclap","PCIe ATS translated-request trust","IOMMU bypass via Address Translation Services"],"title":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable endpoint the host trusts, including GPUs, NICs and DPUs: The IOMMU/SMMU is the entire…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"PCIe Address Translation Services on hosts using an IOMMU/SMMU for device isolation - affects any DMA-capable…","year":"2019","cvss_score":null,"severity":"unscored","kev":false,"impact":"The IOMMU/SMMU is the entire basis for saying a passed-through GPU or NIC cannot read host or other-tenant memory. ATS undermines it by design: a device marked ATS-capable is trusted to have already translated an address, and the root complex forwards those requests without re-checking them. A malicious or compromised endpoint simply sets the translated bit and issues DMA at any physical address it likes. Thunderclap showed the broader problem too - even without ATS, real OS IOMMU policies map far more than the buffer in question, leaving windows onto adjacent kernel memory. In a GPU-passthrough fleet the consequence is a tenant with device control reading host memory and other tenants' data, which defeats the isolation model the whole product rests on.","attack_vector":"An attacker who controls a DMA-capable PCIe device: a tenant with GPU or NIC passthrough who can flash or exploit device firmware, an attacker with physical access to a slot or an external PCIe/Thunderbolt port, or a supply-chain-modified card. Not reachable from software alone on a well-configured host - the entry point is device control.","remediation":"There is no single patch; this is configuration you must actively verify, and most fleets have never checked. Concretely: disable ATS unless a workload genuinely needs it (`pci=noats` on Linux, or the equivalent BIOS switch), and never leave ATS enabled for a device assigned to a tenant. Confirm PCIe ACS is enabled on every upstream port and switch so peer-to-peer traffic between passed-through devices is forced up through the IOMMU rather than routed directly - many server BIOSes ship ACS off, and several 'GPU peer-to-peer performance' tuning guides tell you to turn it off, which silently deletes the isolation boundary between two tenants' GPUs on the same switch. Verify per-device IOMMU groups are not lumping unrelated functions together. Then decide deliberately whether the peer-to-peer bandwidth you gain by disabling ACS is worth the tenant-isolation guarantee you lose, and write that decision down.","references":["http://thunderclap.io/","http://thunderclap.io/thunderclap-paper-ndss2019.pdf","https://www.kernel.org/doc/html/latest/admin-guide/kernel-parameters.html"],"status":"curated"},{"id":"NCVD-2020-001-amd-zen-1-zen-zen-2-l1d-cache-wa","cve":null,"aliases":["Take A Way","Collide+Probe","Load+Reload"],"title":"AMD Zen 1 / Zen+ / Zen 2 - L1D cache way predictor: MULTI-TENANT ISOLATION: AMD's L1D way predictor hashes virtual addresses to predict which cache way holds a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zen 1 / Zen+ / Zen 2 - L1D cache way predictor","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: AMD's L1D way predictor hashes virtual addresses to predict which cache way holds a line. Collisions in that hash are observable, giving an attacker a channel to leak metadata about a victim's memory access pattern - the researchers used it to break KASLR, to recover an AES key from a table-based implementation, and to build a covert channel between processes. It works from JavaScript in a browser and across VMs, which is unusually broad reach for a microarchitectural channel.","attack_vector":"Local, unprivileged, co-resident with the victim on the same physical core. Affects Zen 1, Zen+ and Zen 2 (2017-2019 EPYC generations).","remediation":"**No CVE was assigned and AMD issued no microcode fix**, taking the position that existing side-channel guidance and secret-independent software already cover it. Treat this as unpatchable on affected silicon. Operator-side controls: do not co-schedule tenants on the same physical core, disable SMT on mixed-tenancy nodes, and prefer newer EPYC generations for workloads where cross-tenant leakage is in your threat model. Nothing here requires a reboot or firmware - it is a scheduling and fleet-composition decision.","references":["https://mlq.me/download/takeaway.pdf","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2020-001-nvidia-multi-instance-gpu-mig-pa","cve":null,"aliases":["MIG side-channel caveat"],"title":"NVIDIA Multi-Instance GPU (MIG) partitioning: MULTI-TENANT ISOLATION: MIG gives each instance its own SM slice, L2 slice, memory slice and memory-bandwidth…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Multi-Instance GPU (MIG) partitioning","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: MIG gives each instance its own SM slice, L2 slice, memory slice and memory-bandwidth allocation, and NVIDIA documents it as providing fault and performance isolation between instances. What it does NOT claim is protection against microarchitectural side channels, and it does not encrypt instance memory. Operators routinely sell MIG partitions as if they were independent GPUs. They are strong performance and fault isolation, not a cryptographic boundary - the shared GPU chip, shared power/thermal domain and shared memory controller remain observable.","attack_vector":"A tenant holding one MIG instance, observing shared resources used by a tenant on another instance of the same physical GPU. No exploit is needed for the telemetry channel; the frequency/power domain is shared by construction.","remediation":"Not a patchable defect - it is the documented scope of the feature. If your product promises isolation stronger than performance isolation, back MIG with NVIDIA Confidential Computing (Hopper and later, and note that CC and MIG have version-dependent compatibility constraints), or with whole-GPU allocation. Verify what you tell customers matches what MIG actually guarantees; also confirm MIG instances are actually destroyed and recreated between tenants rather than reused, because instance teardown is what triggers memory scrubbing.","references":["https://docs.nvidia.com/datacenter/tesla/mig-user-guide/","https://docs.nvidia.com/confidential-computing-deployment-guide/"],"status":"curated"},{"id":"NCVD-2020-004-hpe-sas-ssds-ek0800jvypn-eo1600j","cve":null,"aliases":["HPE 40,000-hour SSD bug","HPE bulletin a00097382en_us","HPD7"],"title":"HPE SAS SSDs EK0800JVYPN, EO1600JVYPP, MK0800JVYPQ, MO1600JVYPR (800GB/1.6TB 12G SAS) with firmware prior to HPD7: PHYSICAL / FLEET-WIDE AVAILABILITY EVENT, and the second one in four months…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"HPE SAS SSDs EK0800JVYPN, EO1600JVYPP, MK0800JVYPQ, MO1600JVYPR (800GB/1.6TB 12G SAS) with firmware prior to HPD7","year":"2020","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL / FLEET-WIDE AVAILABILITY EVENT, and the second one in four months - which is the real finding. At exactly 40,000 power-on hours, roughly 4.6 years, these drives fail permanently and HPE states that neither the data nor the drive can be recovered. Same correlated-failure property as the 32,768-hour bug: co-installed drives die together, so RAID fault tolerance is exceeded and the array goes with them. Affects HPE ProLiant, Synergy, Apollo 4200 and StoreEasy platforms - general-purpose server and storage hardware of exactly the kind that gets repurposed into GPU and AI build-outs. That two independent counter-overflow bombs surfaced in one vendor's SAS SSD line within months tells you this is a recurring firmware-engineering failure mode, not a freak event.","attack_vector":"No attacker. Elapsed powered-on hours, hitting every drive from the same deployment batch at the same moment. HPE projected the first failures would begin around October 2020.","remediation":"Flash to HPD7 or later before the threshold; HPE released it on 20 March 2020 with VMware ESXi, Windows and Linux packages. Post-failure there is no recovery - restore from backup. As with the 32,768-hour bug the first action is a fleet-wide power-on-hours audit (HPE Smart Storage Administrator or smartctl attribute 9) to find which drives are closest to the line, then a staggered rolling drain-and-flash sequenced by hours remaining rather than by rack. Make the standing controls permanent rather than treating this as a one-off: monitor power-on hours as a first-class fleet metric with alerting well ahead of any known threshold, subscribe to drive-vendor firmware bulletins as an operational feed, and deliberately mix procurement batches and vendors across redundancy groups so no single firmware defect can reach every member of an array simultaneously.","references":["https://www.bleepingcomputer.com/news/hardware/hpe-warns-of-new-bug-that-kills-ssd-drives-after-40-000-hours/","https://support.hpe.com/hpesc/public/docDisplay?docId=emr_na-a00097382en_us"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2021-001-amd-platform-secure-boot-psb-oem","cve":null,"aliases":["PSB left unfused","AMD Platform Secure Boot not enabled"],"title":"AMD Platform Secure Boot (PSB) OEM key fusing on EPYC server boards: PSB is the fuse-backed root of trust that makes the PSP verify the OEM's BIOS signature before the x86 cores…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Platform Secure Boot (PSB) OEM key fusing on EPYC server boards","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"PSB is the fuse-backed root of trust that makes the PSP verify the OEM's BIOS signature before the x86 cores ever start. It is not enabled by default - the OEM has to burn their key hash into one-time-programmable fuses at manufacture, and a large fraction of shipped EPYC boards, including whitebox and ODM boards common in neocloud builds, arrive unfused. On an unfused board every firmware persistence flaw in this database becomes trivially exploitable: modified BIOS or PSP firmware boots without complaint, and a bare-metal tenant can leave something behind that owns every customer who follows. There is no CVE because it is a configuration state, not a bug, which is exactly why nobody checks it.","attack_vector":"Anyone who can write the SPI flash - a bare-metal tenant with ring 0, a supply-chain touch point between the factory and your rack, or an attacker chaining one of the SPI protection bypasses in this set. On an unfused board no signature check stands in the way.","remediation":"Verify PSB status at hardware intake, before the node ever enters the fleet - it is readable through the PSP mailbox and via open tooling (psb_status in the fwupd/amd-psb tooling family). Fusing is IRREVERSIBLE and is the OEM's action, not yours: you cannot fuse it yourself after the fact in most designs, and a wrong fuse bricks the board. So this is a procurement requirement - make 'PSB fused with the OEM key' a line item in your server RFP and reject boards that ship unfused. For hardware already deployed unfused, compensate with flash write protection, verified-boot measurement of the SPI image between tenancies, and not selling bare metal on those SKUs.","references":["https://www.amd.com/system/files/documents/amd-security-white-paper.pdf","https://github.com/fwupd/fwupd","https://www.amd.com/en/corporate/product-security"],"status":"curated"},{"id":"NCVD-2021-002-ddr4-dram-with-in-dram-trr-a-cou","cve":null,"aliases":["Half-Double"],"title":"DDR4 DRAM with in-DRAM TRR; a coupling effect that reaches rows at distance two rather than immediate neighbours: Google showed Rowhammer coupling is not confined to adjacent rows - hammering a…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DDR4 DRAM with in-DRAM TRR; a coupling effect that reaches rows at distance two rather than immediate neighbours","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Google showed Rowhammer coupling is not confined to adjacent rows - hammering a row disturbs rows two away, and the mitigation logic that only watches immediate neighbours refreshes exactly the wrong rows. Operationally this means the TRR generation your DIMM vendor sold as a fix has a structural blind spot rather than a tuning problem, and mitigations designed around distance-one coupling need redesign. Same downstream consequences as any Rowhammer primitive: page-table corruption to host escape, or silent corruption of a co-tenant's data with no error signalled.","attack_vector":"Unprivileged local code sharing a memory controller with the victim. No CVE was assigned because this is a property of the DRAM, not a defect in a shippable product.","remediation":"Nothing you can install. The industry answer is Refresh Management (RFM) in the DDR5 spec plus revised in-DRAM tracking, which means new DIMMs and a memory-controller generation that drives RFM - capex on a refresh cycle, not a patch window. In the meantime: raise the refresh rate where BIOS allows it, and treat memory-controller sharing between untrusted tenants as a policy you have chosen to accept rather than a boundary you have.","references":["https://security.googleblog.com/2021/05/introducing-half-double-new-hammering.html","https://github.com/google/hammer-kit"],"status":"curated"},{"id":"NCVD-2021-002-discrete-tpm-lpc-spi-bus-unencry","cve":null,"aliases":["TPM bus sniffing","LPC/SPI interposer"],"title":"Discrete TPM (LPC / SPI bus, unencrypted sessions): A discrete TPM talks to the CPU over LPC or SPI in the clear unless the software explicitly uses…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Discrete TPM (LPC / SPI bus, unencrypted sessions)","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"A discrete TPM talks to the CPU over LPC or SPI in the clear unless the software explicitly uses parameter-encrypted sessions, which most does not. An attacker who clips a logic analyser onto the bus - or onto the exposed pins of a socketed TPM header - reads the sealed key as it is released. The demonstrated case is recovering a BitLocker/LUKS volume key in minutes with a cheap probe, on a machine whose owner believed the disk was hardware-protected. For an operator, this is the flaw that turns a physically accessible node into a full data disclosure, and it leaves no trace in any log.","attack_vector":"Physical access to the motherboard for the duration of one boot. In practice: a colo cage neighbour, remote-hands staff, a decommissioning or RMA handler, or hardware intercepted in shipping. No credentials, no software exploit, no persistence needed.","remediation":"No patch exists - it is a property of the bus, not a bug. Mitigations are architectural: enable TPM parameter encryption / encrypted sessions in the software that unseals (recent Linux and Windows stacks support it, older ones do not), require a PIN or second factor so the TPM value alone is not sufficient to unlock, prefer fTPM where the bus is internal to the package, and use tamper-evident chassis with a documented seal check on every physical touch. Treat any node that left your custody as untrusted until re-provisioned.","references":["https://dolosgroup.io/blog/2021/7/9/from-stolen-laptop-to-inside-the-company-network","https://pulsesecurity.co.nz/articles/TPM-sniffing"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2021-005-intel-cpus-with-sgx-attacked-ove","cve":null,"aliases":["VoltPillager","hardware undervolting attack on SGX"],"title":"Intel CPUs with SGX, attacked over the SVID serial bus between the voltage regulator and the CPU package: Re-runs the Plundervolt fault injection using a cheap external microcontroller wired to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel CPUs with SGX, attacked over the SVID serial bus between the voltage regulator and the CPU package","year":"2021","cvss_score":null,"severity":"unscored","kev":false,"impact":"Re-runs the Plundervolt fault injection using a cheap external microcontroller wired to the motherboard's voltage-regulator bus, so it works on hosts that have already applied the Plundervolt microcode fix - that fix only disabled the software MSR, it did not stop voltage manipulation reaching the package. Recovers keys from enclaves and corrupts enclave computation. Intel's position is that physical attacks are outside the SGX threat model, so there is no patch and there will not be one. The operator consequence is concrete: if an attacker can put hands on the chassis, SGX and by extension the confidential-computing guarantee on that machine are void.","attack_vector":"Physical access to the motherboard, roughly 30 dollars of hardware, and a few minutes to attach to the SVID bus. Relevant threat models: colocation facilities where you do not control the cage, hardware in transit, decommissioned or RMA'd nodes, hosting partners, and anyone with datacenter floor access. Not reachable remotely.","remediation":"UNPATCHABLE by design - Intel classifies it as out of scope, so there is no microcode, BIOS or kernel fix to deploy. Mitigation is entirely physical and procedural: chassis intrusion detection wired into the BMC and actually alarmed, tamper-evident seals, controlled cage access with audited entry logs, and refusing to run confidential workloads on hardware whose physical custody you cannot vouch for. If you sell confidential compute, this is a contractual and facility-security question, not an engineering one - and it should shape which sites you are willing to place attested workloads in.","references":["https://www.usenix.org/conference/usenixsecurity21/presentation/chen-zitai","https://github.com/zt-chen/voltpillager"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2023-001-ddr4-chips-from-all-three-major","cve":null,"aliases":["RowPress"],"title":"DDR4 chips from all three major DRAM manufacturers; worsens as process nodes shrink: A different read-disturbance mechanism from Rowhammer: instead of repeatedly opening and closing a row, hold…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DDR4 chips from all three major DRAM manufacturers; worsens as process nodes shrink","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"A different read-disturbance mechanism from Rowhammer: instead of repeatedly opening and closing a row, hold it open. That cuts the activation count needed for a flip by one to two orders of magnitude, and in the extreme a single activation of an adjacent row suffices. The operator consequence is that activation-counting mitigations - which is what essentially every deployed Rowhammer defence is - can be under threshold and still let flips through. Anything you were told about 'activations per refresh window' as a safety argument needs re-deriving.","attack_vector":"Unprivileged local code on a shared host, demonstrated on a real DDR4 system that already had Rowhammer protection enabled. Sharing a memory controller with the victim is the only requirement.","remediation":"Not patchable. Mitigations have to be redesigned to bound how long a row stays open, not just how often it is activated; the authors show existing Rowhammer defences can be adapted at low additional cost, but that lands in future memory controllers and DRAM, not in your installed base. For now the honest answer to a customer asking 'is our memory isolated from the tenant next door' is no, and the only lever you control is not putting them there.","references":["https://arxiv.org/abs/2306.17061","https://github.com/CMU-SAFARI/RowPress"],"status":"curated"},{"id":"NCVD-2023-001-gigabyte-uefi-firmware-oem-updat","cve":null,"aliases":["Gigabyte App Center backdoor"],"title":"Gigabyte UEFI firmware (OEM update-dropper in firmware): Gigabyte firmware shipped a UEFI module that writes a Windows executable to disk at every boot and has it…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Gigabyte UEFI firmware (OEM update-dropper in firmware)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"Gigabyte firmware shipped a UEFI module that writes a Windows executable to disk at every boot and has it fetch and run further code from Gigabyte-controlled URLs - one of them plain HTTP, with weak certificate validation on the others. It is not malware, it is the vendor's own updater, but it behaves exactly like a firmware implant: an unremovable, boot-persistent downloader with no user consent and no way to disable it from the OS. Anyone who can MITM that fetch, or who compromises the vendor's distribution point, gets code execution on every affected machine at every boot.","attack_vector":"A network attacker in path of the update fetch, or a supply-chain compromise of the vendor endpoint. No access to the machine needed.","remediation":"Gigabyte published firmware updates that fix the transport and validation; applying them is a per-board BIOS flash plus reboot. Where the platform allows it, disable the 'APP Center Download & Install' option in BIOS setup - a config-only mitigation you can push faster than a firmware campaign. The broader operator lesson: audit what your board vendor's firmware talks to on the network before a node ever carries tenant workload, and block outbound egress from the provisioning network by default.","references":["https://eclypsium.com/research/supply-chain-risk-from-gigabyte-app-center-backdoor/","https://kb.cert.org/vuls/id/287178"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2023-001-intel-processors-with-linear-add","cve":null,"aliases":["SLAM","Spectre based on Linear Address Masking"],"title":"Intel processors with Linear Address Masking (LAM): SLAM: Linear Address Masking, a feature intended to let software store metadata in unused address bits, makes…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors with Linear Address Masking (LAM)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"SLAM: Linear Address Masking, a feature intended to let software store metadata in unused address bits, makes previously impractical Spectre gadgets exploitable by widening the set of usable pointer-chasing gadgets - so a performance and usability feature became a security regression. Notable as a case where the mitigation was to ship the hardware feature disabled by default in Linux.","attack_vector":"Local unprivileged code on a host with LAM enabled.","remediation":"Linux disables LAM by default in response; keep it disabled unless you have a specific requirement and have assessed the tradeoff. Kernel-level configuration - a kernel update and reboot, no microcode or BIOS. Verify LAM state on your nodes rather than assuming the default held through a kernel upgrade.","references":["https://www.vusec.net/projects/slam/"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2023-001-msi-intel-boot-guard-oem-key-lea","cve":null,"aliases":["no CVE"],"title":"MSI / Intel Boot Guard OEM key leak: The Money Message ransomware dump exposed MSI's firmware image-signing private keys for 57 products and Intel…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"MSI / Intel Boot Guard OEM key leak","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"The Money Message ransomware dump exposed MSI's firmware image-signing private keys for 57 products and Intel Boot Guard KM/BPM private keys for 116 products, reportedly touching Intel, Lenovo and Supermicro platforms. An attacker can sign a firmware image that the hardware root of trust accepts — Boot Guard is effectively void on affected silicon and the implant survives any OS reinstall","attack_vector":"Supply chain / local flash","remediation":"There is no patch. Boot Guard keys are fused into the CPU at manufacture, so revocation is impossible on shipped hardware. The only response is to treat Boot Guard as non-authoritative on affected platforms and add an independent firmware-measurement/attestation layer","references":["https://www.helpnetsecurity.com/2023/05/08/msi-private-keys-leaked/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2023-003-integrated-gpu-graphics-data-com","cve":null,"aliases":["GPU.zip","graphics data compression side channel"],"title":"Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA): MULTI-TENANT ISOLATION: GPUs apply data-dependent lossless compression to framebuffer traffic even when…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"Integrated GPU graphics data compression (Intel, AMD, Apple, Arm, Qualcomm, NVIDIA)","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: GPUs apply data-dependent lossless compression to framebuffer traffic even when software never asked for it, and the resulting DRAM traffic pattern is measurable from a co-resident context. The published attack recovered pixel data cross-origin from a browser. For a datacenter operator the realistic worry is remote-desktop, cloud-gaming and VDI fleets where the rendered surface is a customer's screen. No CVE was assigned and the affected vendors declined to ship fixes.","attack_vector":"A co-resident attacker able to render and time - in the original work, a web page in another browser tab. On a shared render host, another tenant's session.","remediation":"UNPATCHABLE. Vendors treated the compression as working-as-designed and the browsers mitigated the specific web attack by restricting cross-origin iframe rendering. There is no GPU driver or firmware update. If you sell shared remote-desktop or cloud-gaming capacity, the only real control is not co-residing untrusted sessions on the same GPU.","references":["https://www.hertzbleed.com/gpu.zip/","https://arxiv.org/abs/2310.01187"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2023-004-nvidia-confidential-computing-h1","cve":null,"aliases":["NVIDIA CC DevTools mode"],"title":"NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode: MULTI-TENANT ISOLATION: NVIDIA GPU confidential computing has three modes - CC-Off, CC-On and CC-DevTools.…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Confidential Computing (H100/H200/B100/B200/GB200) - CC-DevTools operating mode","year":"2023","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: NVIDIA GPU confidential computing has three modes - CC-Off, CC-On and CC-DevTools. DevTools mode exists so that profilers and debuggers keep working, and it does so by disabling the protections: the GPU still produces attestation material and still looks like a confidential GPU to tooling, but the confidentiality guarantee is not in force. This is a configuration state, not a bug, and it is the single most likely way an operator ships a 'confidential' GPU that is not one.","attack_vector":"No attacker skill required - the risk is that the mode is set wrong, or set right and then changed by an operator debugging a performance problem and never changed back. Anyone with host root can set it.","remediation":"Not patchable; it is a control you must enforce. Verify CC mode per GPU as a continuous check rather than a build-time one, and make the attestation policy reject DevTools mode explicitly rather than accepting any signed report. Changing CC mode requires a GPU reset and therefore a node drain, which is also why operators are tempted to leave a node in DevTools mode once they set it.","references":["https://docs.nvidia.com/confidential-computing-deployment-guide/","https://docs.nvidia.com/cc-deployment-guide-tee.pdf"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2024-001-amd-global-history-register-side","cve":null,"aliases":["Branch History Leak","AMD-SB-7026"],"title":"AMD - Global History Register side channel: MULTI-TENANT ISOLATION: A side channel through the branch predictor's Global History Register, disclosed by…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD - Global History Register side channel","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: A side channel through the branch predictor's Global History Register, disclosed by Harbin Institute of Technology. Branch history is shared microarchitectural state, so an attacker observes the shape of a co-resident victim's control flow - which for an inference workload can reveal what model is running and what path a request took through it.","attack_vector":"Local, co-resident with the victim.","remediation":"**No CVE and no fix** - AMD's response is software best practices. As with the other predictor-state channels, the enforceable control is not co-scheduling mutually untrusted tenants on the same physical core. No patch, no reboot; this is a scheduling and fleet-composition decision.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7026.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2024-001-intel-sgx-cache-side-channel-on","cve":null,"aliases":["TeeJam"],"title":"Intel SGX (cache side channel on sub-cacheline access): MULTI-TENANT ISOLATION: TeeJam: shows that SGX's cache-based side-channel resistance is weaker than assumed…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX (cache side channel on sub-cacheline access)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: TeeJam: shows that SGX's cache-based side-channel resistance is weaker than assumed at sub-cacheline granularity, enabling practical key recovery against enclave implementations previously believed to be constant-time. Operationally this matters because 'we run it in an enclave' is often the entire argument for putting a key on a shared host.","attack_vector":"Local code on the same machine as the victim enclave, with the scheduling control a privileged host has.","remediation":"No single patch - mitigation lives in the enclave software (constant-time implementations hardened at sub-cacheline granularity) and in keeping the SGX SDK/PSW current. Operator action is to require enclave vendors to state which side-channel hardening they apply, and to keep microcode and PSW at current TCB so attestation reflects reality.","references":["https://www.intel.com/content/www/us/en/developer/topic-technology/software-security-guidance/overview.html"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"NCVD-2024-001-platform-attestation-as-an-opera","cve":null,"aliases":[],"title":"Platform attestation as an operational control (fTPM vs discrete TPM trust): Design-level: on most GPU servers the TPM that backs measured boot is a firmware TPM inside the CPU package…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Platform attestation as an operational control (fTPM vs discrete TPM trust)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Design-level: on most GPU servers the TPM that backs measured boot is a firmware TPM inside the CPU package or chipset, whose own integrity depends on the same firmware the attestation is supposed to be measuring. If the platform firmware or the management engine is compromised, the fTPM reports whatever that firmware tells it to, and a verifier cannot distinguish a clean node from a compromised one. Every entry in this database that yields SMM, BMC or CSME code execution collapses the attestation guarantee alongside it. Operators selling 'verified clean bare metal' or confidential GPU compute are usually asserting something their hardware cannot independently prove.","attack_vector":"Any attacker who reaches the firmware layer beneath the TPM - SMM code execution, BMC takeover, or a management-engine flaw. The attestation does not fail loudly; it keeps passing.","remediation":"No patch. Architectural: root attestation in a device that is independent of the firmware it measures (discrete TPM on its own bus, or a separate platform root-of-trust device such as an OCP Cerberus-style controller), pin expected measurements rather than accepting any well-formed quote, verify the freshness and provenance of quotes rather than just their signature, and pair attestation with an independent firmware-integrity scan. Document to customers what attestation does and does not prove rather than over-claiming it.","references":["https://www.opencompute.org/projects/security","https://trustedcomputinggroup.org/resource/tpm-library-specification/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2024-002-amd-zen-2-zen-3-and-zen-4-platfo","cve":null,"aliases":["ZenHammer"],"title":"AMD Zen 2, Zen 3 and Zen 4 platforms with DDR4 (7/10 Zen 2 and 6/10 Zen 3 devices flipped) and DDR5 (1/10 devices): Removes the 'we run EPYC, Rowhammer research is all Intel' excuse. ETH Zurich…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zen 2, Zen 3 and Zen 4 platforms with DDR4 (7/10 Zen 2 and 6/10 Zen 3 devices flipped) and DDR5 (1/10 devices)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Removes the 'we run EPYC, Rowhammer research is all Intel' excuse. ETH Zurich reverse-engineered AMD's DRAM address mapping and got bit flips on Zen 2/3/4, then demonstrated the standard exploit chain - page-table manipulation, RSA key corruption and sudo privilege escalation. It is also the first public DDR5 bit flip on a commodity system. For a GPU cloud this matters because EPYC is the host CPU under most HGX and 8-way GPU nodes: the box holding your control plane, your tenant credentials and every tenant's data on its way to the GPUs is hammerable from a guest.","attack_vector":"Unprivileged local code on an AMD Zen 2/3/4 host sharing DRAM with the victim. A container tenant or VM on the CPU side of a GPU node is sufficient; no GPU access needed.","remediation":"No vendor patch. The DDR4 case is the same non-answer as TRRespass and Blacksmith: refresh-rate increase where BIOS exposes it, or do not co-tenant behind one memory controller. The DDR5 result is a single device out of ten, so DDR5 EPYC platforms are better but not clear. Before you argue you are safe, run the published ZenHammer fuzzer against a sample of your own DIMM SKUs - vendor and date code determine your exposure far more than the CPU does, and this is the cheapest ground truth available.","references":["https://comsec.ethz.ch/research/dram/zenhammer/","https://github.com/comsec-group/zenhammer"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2024-002-intel-processors-indirect-branch","cve":null,"aliases":["Indirector"],"title":"Intel processors (Indirect Branch Predictor structure): MULTI-TENANT ISOLATION: Indirector: reverse-engineering the Indirect Branch Predictor and Branch Target…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel processors (Indirect Branch Predictor structure)","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Indirector: reverse-engineering the Indirect Branch Predictor and Branch Target Buffer on recent Intel parts yielded high-precision branch target injection that works despite existing eIBRS/IBPB deployment, and showed that IBPB does not clear as much predictor state as assumed. The operator takeaway is that the barrier primitive hypervisors use for tenant separation is weaker than the documentation implies.","attack_vector":"Local code on an affected processor; the work targets cross-process and cross-privilege leakage.","remediation":"No single CVE or microcode fix maps cleanly to this research. Mitigation is the existing toolkit applied more aggressively: keep microcode current, enable IBPB on context switch where your workload can absorb the cost, and treat co-tenancy of untrusted workloads on the same physical core as unsupported. Both cost throughput.","references":["https://indirector.cpusec.org/"],"status":"curated","fleet":{"pain_class":"microcode + reboot"}},{"id":"NCVD-2024-003-nvidia-remote-attestation-servic","cve":null,"aliases":["NRAS dependency","GPU attestation availability"],"title":"NVIDIA Remote Attestation Service (NRAS) / NVIDIA attestation SDK: Verifying an NVIDIA GPU attestation report by the default path means calling NVIDIA's hosted Remote…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA Remote Attestation Service (NRAS) / NVIDIA attestation SDK","year":"2024","cvss_score":null,"severity":"unscored","kev":false,"impact":"Verifying an NVIDIA GPU attestation report by the default path means calling NVIDIA's hosted Remote Attestation Service and fetching Reference Integrity Manifests from NVIDIA infrastructure. That places a live third-party dependency inside the trust decision for every confidential workload you start: if NRAS is unreachable, your orchestration either blocks workload admission or fails open, and most implementations fail open because blocking looks like an outage. There is no CVE here - it is an architectural exposure that operators consistently discover during their first NRAS incident rather than during design.","attack_vector":"Not an attacker in the usual sense. The realistic events are an NRAS outage, a network egress policy that blocks it, or an air-gapped deployment where it was never reachable at all.","remediation":"Deploy the local verifier path rather than the remote one where your threat model allows: NVIDIA supports local attestation verification with cached RIMs and the device identity certificate chain, which removes the runtime dependency. Cost: you take on RIM caching and freshness management. Whichever you choose, test the failure mode deliberately - block NRAS in staging and confirm your admission controller denies rather than admits.","references":["https://docs.attestation.nvidia.com/","https://github.com/NVIDIA/nvtrust"],"status":"curated"},{"id":"NCVD-2025-001-amd-zen-3-zen-4-new-exploitation","cve":null,"aliases":["AMD-SB-7031"],"title":"AMD Zen 3 / Zen 4 - new exploitation method for SRSO (CVE-2023-20569): MULTI-TENANT ISOLATION: Google's security team demonstrated a new way to exploit the existing Inception/SRSO…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Zen 3 / Zen 4 - new exploitation method for SRSO (CVE-2023-20569)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Google's security team demonstrated a new way to exploit the existing Inception/SRSO defect on Zen 3 and Zen 4. No new CVE was assigned and AMD asserts the original mitigations still hold. Recorded here because it is the kind of update that never reaches a patch queue: nothing to install, but your risk assessment of an already-known issue should move if the exploitation bar just dropped.","attack_vector":"Local, cross-privilege speculative execution on Zen 3 and Zen 4.","remediation":"**No new action** if you already applied the original SRSO mitigations (microcode plus AGESA, MilanPI 1.0.0.C / GenoaPI 1.0.0.9 or later, plus the kernel's Safe RET). Verify that you did - read /sys/devices/system/cpu/vulnerabilities/spec_rstack_overflow across the fleet rather than assuming. Note the interaction with AMD-SB-7061 above: Safe RET, the mitigation you are relying on here, has its own open weakness.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7031.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2025-001-aspeed-ast2600-ast2700-hardware","cve":null,"aliases":["ASPEED secure boot disabled by default","OpenBMC RoT not enabled"],"title":"ASPEED AST2600 / AST2700 hardware root of trust in OpenBMC builds: AST2600 has a fuse-backed secure boot that verifies the BMC firmware image against an RSA key, and AST2700…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"ASPEED AST2600 / AST2700 hardware root of trust in OpenBMC builds","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"AST2600 has a fuse-backed secure boot that verifies the BMC firmware image against an RSA key, and AST2700 adds the OCP Caliptra-based scheme. Both ship off. The default OpenBMC build configuration leaves secure boot disabled because that is convenient for board bring-up, and plenty of production images never turn it on. Where it is off, anything that can write the BMC's SPI flash - a compromised BMC userspace process, a host-side AHB bridge write, or a supply-chain-tampered update package - installs code that persists across power cycles, OS reimage, and node redeployment to the next tenant, with nothing in the boot chain to reject it. On a GPU cluster this is the difference between a compromised node you can rebuild and a compromised node you have to physically retire.","attack_vector":"Any write path to BMC SPI flash: code execution on the BMC, a host-side AHB bridge, or a malicious/unsigned firmware update accepted by the update daemon. Not remotely reachable by itself - it is the amplifier that turns a one-shot BMC compromise into a permanent one.","remediation":"Enabling it is a one-way door: it means burning OTP fuses with your signing key on every node, which cannot be undone and cannot be done remotely. Practically this is an order-time decision with the ODM, not something an operator retrofits on a deployed fleet. For fleets already racked, audit whether secure boot is fused on each SKU, demand the answer in writing from the ODM, and where it is off, compensate with SPI flash content attestation (hash the image out-of-band and compare against a known-good) plus strict control over who can push BMC firmware. Note that a hardware root of trust also does not help if the signed image itself has a signature-verification bug - see the Supermicro image-parser entries.","references":["https://eclypsium.com/wp-content/uploads/OpenBMC-Security-in-Practice.pdf","https://developer.nvidia.com/blog/analyzing-baseboard-management-controllers-to-secure-data-center-infrastructure/","https://www.binarly.io/blog/old-but-gold-the-underestimated-potency-of-decades-old-attacks-on-bmc-security"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2025-001-intel-sgx-ddr4-memory-bus-physic","cve":null,"aliases":["WireTap"],"title":"Intel SGX / DDR4 memory bus (physical interposer): MULTI-TENANT ISOLATION: WireTap: a low-cost passive DDR4 interposer reads the memory bus of an SGX machine…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX / DDR4 memory bus (physical interposer)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: WireTap: a low-cost passive DDR4 interposer reads the memory bus of an SGX machine and recovers the platform's attestation key, letting the attacker forge quotes that Intel's own attestation service accepts. The consequence for an operator is that SGX remote attestation no longer proves anything about a machine an adversary has had physical access to - which includes colocation, transit, and any hardware that has left your custody. Notably the published work reports the attack against production DCAP attestation, not a lab-only configuration.","attack_vector":"Physical access to the machine long enough to install an interposer between the CPU and a DIMM. Not remote, but well within reach of anyone in the supply chain, a colo neighbour with cage access, or an insider in a datacenter you do not own.","remediation":"Unpatchable in the field on affected parts - deterministic memory encryption without integrity or freshness is a design property of SGX on these generations, not a bug with a patch. Operator response is procedural: treat physical custody as part of the SGX trust boundary, refuse to accept attestation from hardware outside your custody chain, and plan migration to platforms with stronger memory integrity. Watch for Intel TCB recovery advisories, but do not assume one will close this.","references":["https://wiretap.fail/"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2025-002-amd-sev-es-sev-snp-transient-exe","cve":null,"aliases":["PowerHooK","AMD-SB-3032"],"title":"AMD SEV-ES / SEV-SNP - transient-execution-amplified power side channel: MULTI-TENANT ISOLATION: Graz researchers amplified the power side channel with transient execution to extract…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-ES / SEV-SNP - transient-execution-amplified power side channel","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Graz researchers amplified the power side channel with transient execution to extract AES key bytes from inside SEV-ES and SEV-SNP guests. The significance for an operator is the trend line: power telemetry keeps turning out to be a channel that memory encryption does not cover, and each iteration extracts more with less. If you sell confidential computing, the host's power meter is part of your attack surface.","attack_vector":"Requires a malicious hypervisor with access to power telemetry (RAPL) on a host running confidential guests.","remediation":"**No CVE and no microcode fix** - AMD's answer is to restrict or disable hypervisor RAPL access. That is a configuration change you can make today at no performance cost: ensure the host's energy interfaces are root-only and are not exposed to any process a tenant can influence, and do not pass power telemetry into guests. No reboot, no firmware. Combine with performance-determinism mode if you are also mitigating Collide+Power, accepting the throughput cost that carries.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3032.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"hot-patch"}},{"id":"NCVD-2025-002-intel-sgx-and-amd-sev-snp-dram-i","cve":null,"aliases":["Battering RAM"],"title":"Intel SGX and AMD SEV-SNP / DRAM interposer (memory aliasing): MULTI-TENANT ISOLATION: Battering RAM: a cheap DRAM interposer that aliases physical addresses so the same…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Intel SGX and AMD SEV-SNP / DRAM interposer (memory aliasing)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Battering RAM: a cheap DRAM interposer that aliases physical addresses so the same ciphertext is served for two different locations, defeating the memory-encryption protections of both Intel SGX and AMD SEV-SNP and giving the attacker plaintext access to enclave and confidential-VM memory. Same operator conclusion as WireTap and it generalises across both vendors' confidential-compute stacks, so it is not a reason to switch silicon.","attack_vector":"Physical access to install an interposer on the DIMM path. Applies to any deployment where hardware is not continuously in your physical custody.","remediation":"No firmware or microcode fix on affected parts - the attack targets a design property of deterministic memory encryption. Treat confidential compute as protecting against a remote or software-privileged adversary, not a physical one, and write that into what you tell customers. Physical security, tamper-evident hardware and custody controls are the actual mitigation.","references":["https://batteringram.eu/"],"status":"curated"},{"id":"NCVD-2025-003-amd-sev-snp-ciphertext-side-chan","cve":null,"aliases":["Relocate+Vote","Chosen Plaintext Oracle against SEV-SNP","AMD-SB-3021"],"title":"AMD SEV-SNP - ciphertext side channels amplified by hypervisor page movement: MULTI-TENANT ISOLATION: Two 2025 follow-ups to CipherLeaks - Toronto's Relocate+Vote and ETH Zurich's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - ciphertext side channels amplified by hypervisor page movement","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Two 2025 follow-ups to CipherLeaks - Toronto's Relocate+Vote and ETH Zurich's chosen-plaintext oracle - both combine ciphertext visibility with the hypervisor's ability to move or swap guest pages. Relocating a page changes which address the deterministic encryption is keyed to, giving the attacker the equivalent of a chosen-plaintext oracle against a confidential VM. AMD publishes this as informational with **no CVE**, which is the point worth noting: your vulnerability scanner will never mention it.","attack_vector":"Malicious hypervisor able to observe guest ciphertext and relocate guest pages.","remediation":"**No patch.** The available control is guest policy: SEV-SNP ABI 1.58 and later let a guest forbid hypervisor page move and swap, which removes the amplification. As the operator, support and document that policy bit so tenants can set it; as a tenant-facing claim, be honest that ciphertext visibility on Zen 3 and Zen 4 is architectural. The durable fix is Ciphertext Hiding on Zen 5 (Turin) - a hardware refresh, not a maintenance window.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3021.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2025-003-nvidia-gddr6-gpu-memory-a6000-cl","cve":null,"aliases":["GPUHammer"],"title":"NVIDIA GDDR6 GPU memory (A6000-class and similar discrete GPUs): MULTI-TENANT ISOLATION: the first demonstrated Rowhammer bit flips in GPU memory. Researchers flipped bits in…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GDDR6 GPU memory (A6000-class and similar discrete GPUs)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: the first demonstrated Rowhammer bit flips in GPU memory. Researchers flipped bits in GDDR6 on an NVIDIA A6000 from a co-tenant GPU workload and used a single flip to degrade a victim's neural network accuracy from around 80% to near zero. The attack does not read the victim's data - it corrupts it - which makes it an integrity attack on a co-tenant's model or training run rather than a confidentiality one, and correspondingly hard to detect from inside the victim job.","attack_vector":"A co-tenant workload on the same physical GPU with the ability to allocate and hammer memory. No privileges, no driver bug. GPUs without ECC or with ECC disabled are the exposed population.","remediation":"NVIDIA's published response is to enable System-Level ECC, which is available and on by default on datacenter parts (H100, A100, and the Hopper/Blackwell line) but is off or absent on workstation-class parts. Verify with nvidia-smi -q -d ECC across the fleet and turn it on with nvidia-smi -e 1, which requires a GPU reset - so a node drain. Cost is real: ECC on these parts costs roughly 6-10% inference throughput and around 6% of usable VRAM. HBM3/HBM3e parts with on-die ECC are considered less exposed. There is no firmware patch that removes the underlying DRAM weakness.","references":["https://gpuhammer.com/","https://arxiv.org/abs/2507.08166"],"status":"curated","fleet":{"pain_class":"node-drain"}},{"id":"NCVD-2025-004-amd-sev-snp-rmp-entries-cached-i","cve":null,"aliases":["AMD-SB-3036"],"title":"AMD SEV-SNP - RMP entries cached in L1D/L2 leaking physical address bits: MULTI-TENANT ISOLATION: Reverse-map table entries cached in L1D and L2 leak up to six physical address bits…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - RMP entries cached in L1D/L2 leaking physical address bits","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Reverse-map table entries cached in L1D and L2 leak up to six physical address bits to an **unprivileged** process. AMD and the researchers agree there is no immediate impact, and six bits is not a compromise on its own - but physical address bits are exactly the primitive that makes Rowhammer, cache-eviction-set construction and DMA targeting practical, so it is a building block for other people's attacks rather than an attack itself.","attack_vector":"Local, unprivileged - notably lower than the rest of the RMP family, which mostly needs hypervisor privilege.","remediation":"**No fix planned.** Nothing to install and nothing to reboot for. Treat it as a standing reminder that side-channel primitives accumulate: it lowers the cost of the next attack against your SNP hosts without ever appearing in a patch queue. Track AMD-SB-3036 in case AMD's assessment changes.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3036.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2025-005-amd-sev-snp-dimm-interposer-vari","cve":null,"aliases":["AMD-SB-3024"],"title":"AMD SEV-SNP - DIMM interposer variant of BadRAM (KU Leuven): MULTI-TENANT ISOLATION: A memory-bus interposer variant of the BadRAM aliasing attack against SEV-SNP. AMD's…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV-SNP - DIMM interposer variant of BadRAM (KU Leuven)","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: A memory-bus interposer variant of the BadRAM aliasing attack against SEV-SNP. AMD's response is **WONTFIX** - the attack is declared outside the SEV-SNP threat model because it requires physical interposition on the memory bus.","attack_vector":"Physical access with a memory-bus interposer.","remediation":"**No patch, and none coming** - AMD has scoped it out of the threat model. The operator control is physical: tamper-evident chassis handling, chain of custody, and not making confidential-computing claims that a tenant could reasonably read as covering an adversary with physical access to the DIMM slot. If a customer's threat model includes your own datacenter staff, this is a conversation to have explicitly rather than one to leave to the marketing page.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3024.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2025-006-amd-confidential-computing-ddr5","cve":null,"aliases":["AMD-SB-3040"],"title":"AMD confidential computing - DDR5 memory bus interposition against TEEs: MULTI-TENANT ISOLATION: Compromising trusted execution environments by interposing on the DDR5 memory bus.…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD confidential computing - DDR5 memory bus interposition against TEEs","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Compromising trusted execution environments by interposing on the DDR5 memory bus. Same family as the BadRAM interposer work and the same conclusion for an operator: SEV-SNP's guarantees are software- and firmware-scoped, and an adversary with hands on the hardware sits outside them.","attack_vector":"Physical access with DDR5 bus interposition hardware.","remediation":"**No patch** - physical attacks fall outside the declared SEV-SNP threat model. Mitigation is datacenter physical security, hardware chain of custody, and honest scoping of what confidential computing does and does not promise your tenants.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3040.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2025-007-amd-secure-processor-boot-rom-ph","cve":null,"aliases":["AMD-SB-7044"],"title":"AMD Secure Processor boot ROM - physical attacks bypassing secure boot: MULTI-TENANT ISOLATION: Physical attacks that bypass secure boot in the ASP boot ROM. Boot ROM is…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD Secure Processor boot ROM - physical attacks bypassing secure boot","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: Physical attacks that bypass secure boot in the ASP boot ROM. Boot ROM is mask-programmed silicon, so the flawed logic itself cannot be patched - only worked around by the firmware layered above it. An attacker who defeats ASP secure boot owns the platform's root of trust from power-on.","attack_vector":"Physical access to the platform.","remediation":"**Boot ROM is unpatchable by construction** - any mitigation is compensating logic in AGESA/PI firmware above it, delivered as an OEM BIOS package. The real controls are physical: chain of custody for hardware, tamper evidence, and treating any node returned from third-party hands as untrusted until its firmware is measured and reprovisioned. This is the argument for firmware attestation between bare-metal tenants rather than trusting a reimage.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7044.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2025-008-nvidia-gpus-with-gddr6-and-no-on","cve":null,"aliases":["GPUHammer"],"title":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000. Not reproduced on A100 (HBM), H100 or RTX 5090, which have on-die ECC: The first working Rowhammer against GPU memory, driven…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000. Not reproduced on A100 (HBM), H100 or RTX…","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The first working Rowhammer against GPU memory, driven from ordinary user-level CUDA. It flips bits in a neighbouring tenant's GPU memory. The headline result is the one that should worry an AI datacenter: a single bit flip in the exponent of a floating-point weight took an ImageNet model from 80% to 0.1% accuracy. There is no crash, no XID, no ECC event on an ECC-off GPU - the victim's model simply stops working and they have no way to attribute it to you. In a fleet that time-slices or MIG-partitions GPUs between customers, this is a cross-tenant integrity attack against the actual product you sell, and it breaks tenant handoff: residual flips persist in the physical DRAM after the previous tenant leaves.","attack_vector":"A tenant running unprivileged CUDA on a GPU whose memory is shared with, or was previously allocated to, the victim - MIG partitions, time-sliced sharing, MPS, or simply the next tenant on a re-let card. No driver exploit and no privilege escalation required.","remediation":"Enable ECC on every GPU that does not have on-die ECC: `nvidia-smi -e 1`, then reset or reboot the GPU. It is not free - roughly a 10% inference slowdown on an A6000 and about 6.25% of memory capacity gone, which on a rental fleet is directly billable capacity you stop selling. Cards with on-die ECC (H100, GB200-class HBM3e, RTX 5090) are not affected in the demonstrated form. Practical fleet policy: refuse to share a physical GPU between untrusted tenants at all, scrub and reset GPU memory between tenants, and audit that ECC has not been disabled by a tenant with elevated access. Verify ECC state per GPU rather than assuming the fleet default held.","references":["https://gpuhammer.com/","https://arxiv.org/abs/2507.08166"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2025-011-nvidia-gpus-with-gddr6-and-no-on","cve":null,"aliases":["GPUHammer"],"title":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000. Not reproduced on A100 (HBM), H100 or RTX 5090, which have on-die ECC: The first working Rowhammer against GPU memory, driven…","layer":"gpu-stack","layer_name":"NVIDIA / GPU stack","component":"NVIDIA GPUs with GDDR6 and no on-die ECC - demonstrated on RTX A6000. Not reproduced on A100 (HBM), H100 or RTX…","year":"2025","cvss_score":null,"severity":"unscored","kev":false,"impact":"The first working Rowhammer against GPU memory, driven from ordinary user-level CUDA. It flips bits in a neighbouring tenant's GPU memory. The headline result is the one that should worry an AI datacenter: a single bit flip in the exponent of a floating-point weight took an ImageNet model from 80% to 0.1% accuracy. There is no crash, no XID, no ECC event on an ECC-off GPU - the victim's model simply stops working and they have no way to attribute it to you. In a fleet that time-slices or MIG-partitions GPUs between customers, this is a cross-tenant integrity attack against the actual product you sell, and it breaks tenant handoff: residual flips persist in the physical DRAM after the previous tenant leaves.","attack_vector":"A tenant running unprivileged CUDA on a GPU whose memory is shared with, or was previously allocated to, the victim - MIG partitions, time-sliced sharing, MPS, or simply the next tenant on a re-let card. No driver exploit and no privilege escalation required.","remediation":"Enable ECC on every GPU that does not have on-die ECC: `nvidia-smi -e 1`, then reset or reboot the GPU. It is not free - roughly a 10% inference slowdown on an A6000 and about 6.25% of memory capacity gone, which on a rental fleet is directly billable capacity you stop selling. Cards with on-die ECC (H100, GB200-class HBM3e, RTX 5090) are not affected in the demonstrated form. Practical fleet policy: refuse to share a physical GPU between untrusted tenants at all, scrub and reset GPU memory between tenants, and audit that ECC has not been disabled by a tenant with elevated access. Verify ECC state per GPU rather than assuming the fleet default held.","references":["https://gpuhammer.com/","https://arxiv.org/abs/2507.08166"],"status":"curated","fleet":{"pain_class":"node-reboot"}},{"id":"NCVD-2026-001-dmtf-libspdm-get-measurement-ext","cve":null,"aliases":["DMTF-2026-0002"],"title":"DMTF libspdm (GET_MEASUREMENT_EXTENSION_LOG offset/length wrap): Wrapping addition of the Offset and Length fields in GET_MEASUREMENT_EXTENSION_LOG lets a requester read…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DMTF libspdm (GET_MEASUREMENT_EXTENSION_LOG offset/length wrap)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Wrapping addition of the Offset and Length fields in GET_MEASUREMENT_EXTENSION_LOG lets a requester read memory outside the measurement log from a libspdm responder. The attacker is the host or fabric peer asking a device for its measurements, and what it gets back is device firmware memory - which is where attestation keys and secrets live. No CVE was assigned, so this will not appear in NVD, OSV or any scanner your fleet runs.","attack_vector":"Any SPDM requester that can reach an affected responder with MEL_CAP and CHUNK_CAP set - in practice the host talking to its own accelerators or NICs, so host-side root on a bare-metal node is enough.","remediation":"libspdm update embedded in device or platform firmware, plus - importantly - integrator hygiene: the bug only bites when the integrator's libspdm_copy_mem() has its assertions compiled out, which is a build-configuration decision your device vendor made and you cannot see. There is no config mitigation and no way to detect affected devices from the outside. Ask vendors directly for their libspdm version and build flags as part of hardware acceptance; that question is the only real control here.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-m4wc-xmvg-369f"],"status":"curated"},{"id":"NCVD-2026-001-linux-safe-ret-srso-mitigation-o","cve":null,"aliases":["Safe RET Interrupt Vulnerability","AMD-SB-7061"],"title":"Linux Safe RET SRSO mitigation on AMD Zen 1-Zen 4 - interrupt-induced weakening: MULTI-TENANT ISOLATION: An attacker executing code on the machine injects an interrupt at a precise moment to…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Linux Safe RET SRSO mitigation on AMD Zen 1-Zen 4 - interrupt-induced weakening","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: An attacker executing code on the machine injects an interrupt at a precise moment to disrupt **Safe RET**, which is the default SRSO/Inception mitigation on Linux. Disrupting it weakens the mitigation and can leak information across privilege boundaries. The reason to care operationally is that your tooling will report SRSO as mitigated while the mitigation is being defeated - demonstrated on Zen 1 and Zen 2, suspected on Zen 3 and Zen 4.","attack_vector":"Local, requires code execution on the system and precise interrupt timing - so reachable from a tenant workload, not from the network.","remediation":"**No fix published as of August 2026** - this is the live, open gap in the SRSO story. AMD's assessment attributes it to the Linux Safe RET implementation rather than to silicon, so the eventual fix is expected to be a kernel change rather than microcode or BIOS. Until then, treat SRSO as partially mitigated on Zen 1 through Zen 4 rather than closed: for workloads where cross-privilege speculative leakage is genuinely in your threat model, dedicated nodes rather than shared ones is the only control that holds. Track AMD-SB-7061 for the fix.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7061.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-002-amd-sev-firmware-arbitrary-code","cve":null,"aliases":["AMD-SB-3033"],"title":"AMD SEV firmware - arbitrary code execution on the AMD Security Processor (physical): MULTI-TENANT ISOLATION: An academic disclosure achieving arbitrary code execution on the AMD Security…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD SEV firmware - arbitrary code execution on the AMD Security Processor (physical)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: An academic disclosure achieving arbitrary code execution on the AMD Security Processor itself. Code execution in the ASP means control of SEV key management and attestation for every confidential guest on the node. AMD scopes it out as requiring physical access, but shipped defence-in-depth firmware anyway - which is a reasonable signal about how seriously to take it.","attack_vector":"Physical access to the platform.","remediation":"AMD shipped **defence-in-depth** PI updates - MilanPI 1.0.0.J and GenoaPI 1.0.0.H (both December 2025) - so there is something to deploy despite the WONTFIX-adjacent framing. Delivered as an OEM SBIOS package with the usual lag and a power cycle. The mitigation state is **tenant-verifiable via Platform Info Bit 5** in the attestation report, which is worth advertising to confidential-computing customers: they can check you applied it rather than taking your word.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-3033.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-002-dmtf-libspdm-cryptlib-mbedtls-cs","cve":null,"aliases":["DMTF-2026-0001"],"title":"DMTF libspdm (cryptlib_mbedtls CSR generation, stack overflow): An over-long Common Name in a GET_CSR request writes past a stack array inside the responder, corrupting…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"DMTF libspdm (cryptlib_mbedtls CSR generation, stack overflow)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"An over-long Common Name in a GET_CSR request writes past a stack array inside the responder, corrupting stack data and opening the door to code execution in the device's root-of-trust firmware. The responder here is the thing that is supposed to prove the device is trustworthy, so a compromise at this point does not just break one device - it makes that device's attestation statements attacker-authored. No CVE assigned, so scanners will not flag it.","attack_vector":"Any SPDM requester able to send GET_CSR to a responder that supports CSR_CAP and builds on the mbedTLS crypt backend. On a server that is the host, so local root on a bare-metal node reaches it.","remediation":"libspdm update, delivered as device or platform firmware - per-node flash with vendor rebase lag, no package path, no config toggle. Where a device exposes CSR generation you do not use, ask the vendor whether CSR_CAP can be disabled in their build; turning off an unused capability is the only lever an operator has short of the firmware update.","references":["https://github.com/DMTF/libspdm/security/advisories/GHSA-j54w-759w-xj3m"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-003-amd-rep-string-execution-unit-sc","cve":null,"aliases":["AMD-SB-7069"],"title":"AMD - REP-string execution unit scheduler contention side channel: MULTI-TENANT ISOLATION: A newer variant of the SQUIP scheduler-contention channel, this time reached through…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"AMD - REP-string execution unit scheduler contention side channel","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"MULTI-TENANT ISOLATION: A newer variant of the SQUIP scheduler-contention channel, this time reached through REP-string instructions. Same shape as SQUIP - contention on execution-unit scheduler queues observable from a co-resident thread - and the same operator consequence: cross-tenant leakage that no software isolation boundary sees.","attack_vector":"Local, requires co-residency with the victim on the same physical core.","remediation":"**No fix.** AMD points at existing best practices - constant-time, secret-independent code - which you cannot impose on a tenant's workload. The controls that actually work are yours: do not co-schedule distinct trust domains on sibling SMT threads, or disable SMT on mixed-tenancy nodes. Both are scheduling decisions; disabling SMT needs a reboot, whole-core allocation policy needs a kubelet restart and a drain. On a GPU fleet where the CPU is rarely the bottleneck, the throughput cost is smaller than it looks.","references":["https://www.amd.com/en/resources/product-security/bulletin/amd-sb-7069.html","https://www.amd.com/en/resources/product-security.html"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-005-das-u-boot-fit-image-signature-v","cve":null,"aliases":["BRLY-2026-038","BRLY-2026-041"],"title":"Das U-Boot (FIT image signature verification): Binarly disclosed a cluster of flaws in U-Boot's FIT image handling - stack buffer underflow and NULL…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Das U-Boot (FIT image signature verification)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Binarly disclosed a cluster of flaws in U-Boot's FIT image handling - stack buffer underflow and NULL dereference during signature verification, unbounded recursion in FIT validation, and an SPL out-of-bounds write while loading a FIT image. U-Boot is the first-stage bootloader on essentially every ASPEED-based BMC, so these sit below the BMC firmware itself: an attacker who lands here owns the root of trust for the management controller and nothing running on the host can see it.","attack_vector":"Requires the ability to present a crafted FIT image to the bootloader - in BMC terms, an attacker who already achieved a firmware write, or a malicious/compromised firmware update image. It is the persistence layer of a BMC compromise rather than the initial entry.","remediation":"No CVE IDs assigned as of disclosure. Fix arrives as a U-Boot update embedded in a full BMC firmware image from your board vendor, which means an out-of-band per-node BMC flash and an ODM rebase lag measured in months. There is no config mitigation - the compensating control is signed-update enforcement plus strict isolation of the management network.","references":["https://www.binarly.io/advisories"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-006-uefi-secure-boot-microsoft-2011","cve":null,"aliases":["Secure Boot 2026 certificate expiry"],"title":"UEFI Secure Boot (Microsoft 2011 CA/KEK expiry): Not an exploitable flaw but a fleet-wide trust-anchor deadline. The 2011-era Microsoft UEFI CA and KEK…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"UEFI Secure Boot (Microsoft 2011 CA/KEK expiry)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Not an exploitable flaw but a fleet-wide trust-anchor deadline. The 2011-era Microsoft UEFI CA and KEK certificates baked into most server firmware reach end of validity in 2026. Nodes that never receive updated certificates stop receiving valid Secure Boot revocation updates and, depending on firmware behaviour, may refuse to boot newly signed bootloaders - meaning Secure Boot either silently stops being enforceable or turns into an outage. For a GPU operator this hits the exact control the bare-metal tenant-handoff story depends on.","attack_vector":"No attacker required. The risk is that stale certificates leave you unable to revoke a future vulnerable bootloader, so every bypass in this cluster becomes permanent on affected nodes.","remediation":"Inventory Secure Boot certificate validity per node now, before the deadline. Updated certificates arrive by OEM firmware/BIOS update or OS vendor channel and require a reboot; older boards may never get them, in which case plan hardware refresh or accept that those nodes have no working revocation path. Treat certificate expiry as a scheduled fleet program, not an incident.","references":["https://techcommunity.microsoft.com/blog/windows-itpro-blog/act-now-secure-boot-certificates-expire-in-june-2026/4426856"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-013-community-open-source-sonic-soni","cve":null,"aliases":["community SONiC","sonic-net","open-source network OS blind spot"],"title":"Community / open-source SONiC (sonic-net): Community SONiC — the open-source NOS that a growing share of cost-optimised GPU-cluster fabrics run on — has…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Community / open-source SONiC (sonic-net)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Community SONiC — the open-source NOS that a growing share of cost-optimised GPU-cluster fabrics run on — has essentially no published CVE history. A CPE-based NVD query returns zero records, and the project's own process routes reports privately to its security committee. This is not evidence that SONiC is secure; it is evidence that you have no vulnerability feed for the operating system running your leaf/spine. Every downstream commercial distribution that has been examined (Dell Enterprise SONiC) has produced multiple 9.x-severity findings including authentication bypass, command injection, hard-coded and default credentials — which is what you would expect the shared upstream to look like too. Operators running community SONiC are patching blind.","attack_vector":"Not a specific vulnerability. The exposure is process-level: an operator cannot subscribe to a feed that tells them when their switch OS needs patching, so known-vulnerable images stay in production indefinitely.","remediation":"No patch to apply. Practical controls: pin to a distribution that issues advisories (Dell, Edgecore or a vendor-supported build) rather than self-built community images; track the sonic-buildimage git history for security-relevant commits since there is no advisory feed; scan the SONiC container images for known-vulnerable component versions, because most of the real risk is the Debian base and the bundled daemons (lldpd, FRR, redis, the SDK) rather than SONiC-specific code; and keep switch management interfaces on an isolated OOB network on the assumption that you will not learn about the next flaw in time.","references":["https://github.com/sonic-net/SONiC/security","https://www.dell.com/support/kbdoc/en-us/000245655/dsa-2024-449-security-update-for-dell-enterprise-sonic-distribution-vulnerabilities"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-014-weka-data-platform-and-vast-data","cve":null,"aliases":["WEKA","VAST Data","AI-storage advisory blind spot"],"title":"WEKA Data Platform and VAST Data (published-advisory coverage): Neither WEKA nor VAST Data — two of the most commonly deployed storage platforms under large AI training…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"WEKA Data Platform and VAST Data (published-advisory coverage)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Neither WEKA nor VAST Data — two of the most commonly deployed storage platforms under large AI training clusters — has a meaningful public CVE record. NVD keyword searches return nothing attributable to either product. As with community SONiC, this is an absence of disclosure rather than an absence of vulnerabilities: both are complex distributed systems with kernel-level clients, RDMA data paths and multi-tenant namespace separation, and comparable platforms (Lustre, Spectrum Scale, Ceph, BeeGFS) all have significant published histories including cross-tenant access-control failures. An operator running WEKA or VAST has no external feed telling them when to patch the layer that holds every tenant's training data.","attack_vector":"Not a specific vulnerability. The exposure is that these platforms' security posture is visible only to the vendor, so an operator's patch decisions depend entirely on vendor-issued release notes and on asking directly.","remediation":"No patch. Make it contractual and operational: require the vendor to notify you of security-relevant fixes in release notes and to state a disclosure policy; ask for their most recent third-party penetration-test summary at renewal; keep the storage cluster's management plane on an isolated network; and verify multi-tenant namespace separation yourself with an actual cross-tenant read test rather than trusting the product claim. Track the kernel-client packages these platforms install, since those inherit Linux CVEs on your normal patch cycle even when the platform itself publishes nothing.","references":["https://nvd.nist.gov/vuln/search/results?query=weka","https://nvd.nist.gov/vuln/search/results?query=vast+data"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-018-rack-pdu-and-ups-management-esta","cve":null,"aliases":["shared PDU/UPS credentials","commissioning credential reuse"],"title":"Rack PDU and UPS management estates as a class (all vendors): PHYSICAL, and the most common real-world finding in this whole layer. PDU and UPS management interfaces are…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Rack PDU and UPS management estates as a class (all vendors)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL, and the most common real-world finding in this whole layer. PDU and UPS management interfaces are almost universally commissioned with a single shared credential set, applied by whoever racked the gear, never rotated, and documented in a spreadsheet or a runbook. There is no per-device identity, no rotation, and no logging that would show misuse. Anyone who obtains that one credential - a departing contractor, a leaked runbook, one of the many device-side credential-disclosure CVEs listed above - can switch outlets across the entire hall. No CVE will ever be assigned to this, and it is more likely to be used against you than any of the memory-corruption bugs in this file.","attack_vector":"Anyone who reaches the PDU/UPS management network with the shared credential. On many builds that network is the same one the BMCs sit on, which is the same one a tenant with host root can sometimes see.","remediation":"Not a patch. Per-device credentials issued from a secret manager, PDU management interfaces on a VLAN unreachable from any tenant-facing network or from the BMC network, outlet switching disabled on PDUs that do not need it, and authentication logs from power devices shipped somewhere you actually read. Doing this across an existing hall is a few engineer-weeks and touches every rack, which is why it does not get done.","references":["https://www.cisa.gov/news-events/cybersecurity-advisories/aa24-060a"],"status":"curated"},{"id":"NCVD-2026-019-leased-colocation-facility-infra","cve":null,"aliases":["facility VLAN ownership gap","colo landlord equipment"],"title":"Leased colocation facility infrastructure (power, cooling, access control) as a class: PHYSICAL. Most GPU operators lease space rather than own buildings, which means the UPS, switchgear, PDU…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Leased colocation facility infrastructure (power, cooling, access control) as a class","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"PHYSICAL. Most GPU operators lease space rather than own buildings, which means the UPS, switchgear, PDU upstream feeds, CRAH units and access control that determine whether their GPUs stay powered and cooled are the landlord's equipment, on the landlord's network, patched on the landlord's schedule - which is frequently never, because facility gear is treated as a fixed asset rather than as software. Every CVE in this file is therefore, for many operators, a vulnerability they carry the consequences of and have no authority to fix. An availability event here takes out a hall regardless of how well the compute plane is run.","attack_vector":"Whoever can reach the landlord's facility network - which typically includes the landlord's own vendors, remote-monitoring contractors, and any building-management remote access path the operator has never seen.","remediation":"Contractual, not technical. Require in the colocation agreement: a current inventory of facility control equipment with firmware versions, evidence of a patch cadence, segmentation of the facility network from any operator-reachable network, notification of remote-access paths granted to third parties, and the right to audit. Then verify rather than trust. Where a landlord will not commit, price the risk into the site decision - this belongs in site selection, not in the security backlog.","references":["https://www.cisa.gov/topics/critical-infrastructure-security-and-resilience"],"status":"curated"},{"id":"NCVD-2026-023-nvme-admin-command-set-firmware","cve":null,"aliases":["NVMe Firmware Image Download","NVMe Firmware Commit","firmware downgrade attack","unsigned drive firmware"],"title":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the host OS on bare-metal nodes: CLASS ENTRY, and the single highest-leverage…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY, and the single highest-leverage tenant-handoff risk in this whole category. The NVMe specification puts firmware download and firmware commit in the standard admin command set, meaning a host with root can push a firmware image to the drive using nothing more exotic than nvme-cli. Whether that is safe depends entirely on the drive enforcing signature verification and anti-rollback - and that enforcement is uneven across vendors and generations, is not something the host can independently verify, and has been shown breakable even where present (the Kioxia CM6/PM6/PM7 JTAG work bypasses the RSA signature check outright). Where verification is weak or absent, a departing tenant flashes a modified image and owns the drive controller permanently: it outlives your wipe, your reimage and your reprovision, sees everything the next tenant writes, and lies to every host-side tool about its own version, its lock state and whether a sanitize succeeded. Even where signature checking holds, downgrade to an older SIGNED image with known vulnerabilities is a live path unless the drive enforces anti-rollback - and the RPMB replay flaw (CVE-2020-13799) undermines exactly that anti-rollback state across eMMC, UFS and all NVMe versions.","attack_vector":"A tenant with root on the bare-metal host, issuing standard NVMe admin commands to a locally attached drive. No physical access, no exotic tooling, no exploit needed where the drive does not enforce signing - the command path is a documented, supported feature. Also relevant in SR-IOV and DPU/computational-storage designs, where whether the admin queue and Security Send/Receive are properly filtered from a tenant-controlled function is a per-platform question most operators have never actually tested.","remediation":"Policy and platform configuration; there is no patch because the command path is by design. (1) Block it at the platform: filter Firmware Image Download and Firmware Commit - and vendor-specific and Security Send/Receive pass-through - so tenant-controlled hosts and virtual functions cannot reach them. Do not expose raw NVMe admin queues to tenants unless the product genuinely requires it. (2) TEST it rather than assuming: on a representative node, try to flash a drive from a tenant-equivalent shell and confirm you are refused. Most operators have never run this check and will be surprised by the result on at least one SKU. (3) Make signed firmware and enforced anti-rollback a written procurement requirement, and get the vendor to state it per SKU. (4) Inventory expected firmware version per drive serial in a store the host cannot write to, and alert on any change or any version that decreases. (5) For sensitive tenancies, retire local media at end of tenancy instead of recycling it. Assume that verifying firmware integrity across a 10,000-drive fleet is not achievable with host-side tooling - the controller is the thing answering your questions - so the control has to be preventing the write and controlling the media's lifecycle, not detecting the implant afterwards.","references":["https://www.nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0c-2022.10.04-Ratified.pdf","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh","https://www.kb.cert.org/vuls/id/231329","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-024-bacnet-bacnet-ip-as-a-protocol-f","cve":null,"aliases":[],"title":"BACnet / BACnet IP as a protocol (facility control plane): BACnet has no authentication, no integrity protection and no encryption at the network layer. Any device that…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"BACnet / BACnet IP as a protocol (facility control plane)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"BACnet has no authentication, no integrity protection and no encryption at the network layer. Any device that can put a BACnet frame on the wire can issue a WriteProperty to any object on any controller that will accept it - fan speed, damper position, chilled-water setpoint, occupancy schedule, alarm enable. There is no credential to steal because none exists, and there is no log entry that distinguishes a legitimate command from a forged one. This is the single highest-leverage weakness in the whole facility stack for an AI datacenter: an attacker who reaches the BACnet segment does not need to exploit anything, they simply operate the building. Against a hall of 40-140 kW GPU racks, writing setpoints or zeroing fan commands crosses accelerator thermal-shutdown thresholds in minutes, taking down every in-flight training job and stressing hardware through repeated thermal cycles. It also breaks tenant handoff in a shared building: BACnet gives no way to scope one tenant's control authority away from another's equipment, so a compromised neighbour on the same building segment can command your cooling.","attack_vector":"Any host on the BACnet/IP segment, unauthenticated, using off-the-shelf tooling (YABE, Wireshark's BACnet dissector, the open-source BACnet stack utilities). Reachability is everything: the segment typically includes mechanical rooms, IDF closets, the fire and lighting integrators' gear, the landlord's building network, and every controls contractor's laptop that has ever been plugged in. BACnet/IP uses UDP 47808 and relies on broadcast, so it also crosses VLANs wherever a BBMD (BACnet Broadcast Management Device) has been configured to bridge them - operators routinely do not know where their BBMDs are. BACnet MS/TP behind a BACnet router is reachable from IP through that router.","remediation":"Unpatchable by design; the protocol will never authenticate. Three real options, in order of what most operators can actually do. First, segmentation: BACnet on a dedicated VLAN with no route to tenant, corporate or internet networks, an inventory of every BBMD, and switch-level port security or 802.1X on ports serving mechanical spaces. Second, physical security: locked mechanical rooms and control panels, because RS-485 field bus access is a wirecutter away. Third, the actual protocol fix - BACnet Secure Connect (BACnet/SC), which adds TLS and certificate-based device identity; it is supported by newer controller generations and is a controller-replacement project, so treat it as a capital line item for any new build and a multi-year migration for an existing one. For a leased colo the honest answer is that you cannot fix this yourself: it is the landlord's control network. Put it in the contract - require BACnet segment isolation, require disclosure of BBMD placement, require that no tenant network can route to it, and require evidence rather than assurance.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-26-078-08","https://nvd.nist.gov/vuln/detail/CVE-2026-32666","https://www.cisa.gov/resources-tools/resources/secure-design-alert-security-design-improvements-scada-devices"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-024-nvme-of-fabric-authentication-as","cve":null,"aliases":[],"title":"NVMe-oF fabric authentication as deployed - host NQN allowlisting on Linux nvmet, SPDK and most storage appliances: There is no CVE for this because it is the specification working as designed…","layer":"control-plane","layer_name":"Control plane, storage & DevOps","component":"NVMe-oF fabric authentication as deployed - host NQN allowlisting on Linux nvmet, SPDK and most storage appliances","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"There is no CVE for this because it is the specification working as designed, and it is the single largest tenant-isolation gap in NVMe-oF as it is actually deployed. Outside of in-band DH-HMAC-CHAP (NVMe TP-8006, Linux 6.0+) and NVMe/TCP TLS (Linux 6.7+), the only thing binding a namespace to a tenant is the host NQN the initiator asserts in its Connect command. The NQN is a self-declared string. Any host that can reach the target's transport port can claim another tenant's NQN and be granted that tenant's namespaces - full read and write of another customer's dataset or checkpoint volume, with no exploit, no memory corruption and nothing anomalous in the target logs. Most neocloud NVMe-oF deployments run exactly this way because auth and TLS are off by default and cost throughput.","attack_vector":"Any host with IP or RDMA reachability to the target's transport port (TCP 4420 or the RDMA CM port) - a tenant bare-metal node, a compromised BMC bridged onto the storage VLAN, or anything that lands on the storage network. The attacker needs to learn or guess the victim's host NQN, which is usually derived from a predictable pattern or readable from orchestration metadata.","remediation":"Not patchable - this is a configuration and architecture decision. Enable DH-HMAC-CHAP in-band authentication on every subsystem with per-host keys, provisioned by the same system that provisions the namespace, and enable NVMe/TCP TLS where the kernel and target support it. Both cost CPU and some latency, which is why they are skipped; measure it rather than assuming. Independently: put the storage fabric on its own VLAN/VRF with per-initiator ACLs so an unknown host cannot open a connection at all, and treat host NQNs as secrets rather than as inventory labels. Note that turning on DH-HMAC-CHAP exposes you to the nvmet-auth parsing bugs elsewhere in this database, so patch the target kernel in the same change.","references":["https://nvmexpress.org/specifications/","https://docs.kernel.org/admin-guide/nvme-multipath.html"],"status":"curated"},{"id":"NCVD-2026-025-landlord-owned-facility-control","cve":null,"aliases":[],"title":"Landlord-owned facility control network in a leased colo or wholesale hall (governance gap): Almost every neocloud and bare-metal GPU provider runs in space it does not own, which means the entire…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Landlord-owned facility control network in a leased colo or wholesale hall (governance gap)","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Almost every neocloud and bare-metal GPU provider runs in space it does not own, which means the entire cooling and BMS layer - the thing that can kill a whole hall of accelerators in minutes - is operated by a third party the tenant has no technical visibility into and no ability to patch. The tenant carries the full financial consequence of a thermal event (lost training runs, missed SLAs, damaged accelerators, customer churn) while holding none of the controls. In practice the tenant does not know which BMS vendor is deployed, what firmware the controllers run, whether the BACnet segment is isolated from the building's office network, whether the mechanical contractor has a permanent remote-access tunnel in, or who else has a badge to the mechanical rooms. In multi-tenant halls it is worse: cooling is shared infrastructure, so a compromise driven through a neighbouring tenant's contractor lands on your racks. This also directly breaks tenant handoff - when a hall is re-let, nothing in the facility control layer is reset, re-flashed or credential-rotated between occupants.","attack_vector":"Not a single vector - a structural one. The realistic entry points are the mechanical contractor's remote-support path, the landlord's corporate network being flat with the building VLAN, an unmonitored BBMD bridging BACnet across zones, a shared BMS supervisor serving all tenants, and physical access to mechanical rooms and control panels by staff and contractors who are outside the tenant's security program entirely.","remediation":"You cannot patch this - it is the landlord's equipment. The honest remediation is contractual and it needs to be in the lease or the master services agreement, not a security questionnaire answered once. Ask for and get in writing: the BMS/BAS vendor and product versions serving your halls; a network diagram showing the facility VLAN and every route off it; confirmation that no tenant or corporate network can route to the BACnet/Modbus segments; the list of remote-access paths into building controls and who holds them; a commitment to notify you within a defined window when CISA publishes an ICS advisory affecting deployed equipment, with a patch SLA; the right to have a third party validate the segmentation; and physical access logs for mechanical rooms. Also negotiate what you actually need operationally - independent temperature telemetry you own (your own sensors on your own network, not the landlord's BMS feed), so you can detect a thermal excursion without trusting a system you cannot audit, and a documented, tested time-to-thermal-shutdown figure for your rack density so you know how many minutes of margin you are buying.","references":["https://www.cisa.gov/news-events/ics-advisories","https://www.cisa.gov/resources-tools/resources/layering-network-security-through-segmentation"],"status":"curated"},{"id":"NCVD-2026-025-raid-hba-controller-firmware-upd","cve":null,"aliases":[],"title":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip Adaptec SmartRAID/SmartHBA, and their OEM rebadges (Dell PERC, HPE SR/Smart Array…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"On most of these SKUs the controller firmware can be written from the running host OS by a root process using the vendor CLI (StorCLI, arcconf, perccli, ssacli) over the normal driver ioctl path. There is no host-verifiable attestation of what firmware the controller is actually running - the version string you read back is reported by the same firmware you are trying to verify. The controller is a PCIe device with DMA to host memory that sits below the operating system and below the boot chain, and its firmware is not covered by UEFI Secure Boot or by any measured-boot chain that a normal reimage re-establishes. For a bare-metal GPU rental, that means a tenant with root - which every bare-metal tenant has - can write persistent code beneath the next tenant's operating system, and a wipe-and-reimage handoff does not remove it. This is the concrete mechanism behind 'bare-metal tenant handoff is not a reimage', on a component operators rarely inventory at all.","attack_vector":"Local root on the bare-metal host: the legitimate tenant during their rental, or anyone who achieved root through any other path. No physical access, no BMC access, and no reboot required to stage the flash on most controllers.","remediation":"Not patchable and largely unmitigated on current SKUs. What operators can actually do: (1) make controller firmware version and checksum part of the handoff checklist and reflash from a vendor-signed image between tenants rather than trusting the reported version - budget the reboot into the OEM update utility and the node drain, and expect the OEM package to trail Broadcom/Microchip by months; (2) prefer platforms where the controller participates in a platform root of trust that the BMC can attest, and make that a procurement requirement rather than a hope; (3) blacklist or restrict the management ioctl path from tenant workloads where the workload does not need it; (4) accept and price the residual risk for SKUs where the firmware cannot be independently verified, and keep those nodes out of the pool you offer for security-sensitive tenants.","references":["https://www.broadcom.com/support/resources/product-security-center","https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-026-ses-scsi-enclosure-services-encl","cve":null,"aliases":[],"title":"SES (SCSI Enclosure Services) enclosure management on shared SAS JBODs and expanders: SES is how a host controls a drive enclosure - slot power, locate LEDs, fan and power-supply state - and it…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"SES (SCSI Enclosure Services) enclosure management on shared SAS JBODs and expanders","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"SES is how a host controls a drive enclosure - slot power, locate LEDs, fan and power-supply state - and it has no authentication or authorization of any kind. In a shared JBOD or a multi-initiator SAS topology, every attached host can issue SES commands to the enclosure, including commands affecting slots that belong to another host's drives. An operator running a dense shared-enclosure build therefore has a path where one tenant's node can power-cycle or spin down drives serving a different tenant, or induce enough enclosure-level disruption to fail arrays, with nothing more exotic than a standard SCSI command through /dev/sg. The kernel-side ses driver has also had its own out-of-bounds bugs (CVE-2023-53675, CVE-2023-53431) that make malformed enclosure descriptors a host-crash primitive, which matters when the enclosure firmware is the thing supplying those descriptors.","attack_vector":"Local root on any host attached to the shared SAS fabric, issuing SES commands through the generic SCSI interface. Requires no privileged position beyond being one of the initiators the enclosure already trusts - which is the design.","remediation":"Not patchable at the protocol level; SES has no notion of an authenticated initiator. Mitigations are topological: use SAS zoning on the expander so each host only sees its own drive groups and the enclosure services it needs, avoid multi-tenant shared JBODs entirely for bare-metal rentals, and keep enclosure firmware current for the expander-side parsing bugs. Patch the host kernel for the ses driver out-of-bounds issues so a misbehaving or malicious enclosure cannot crash the initiator. Treat SAS zoning configuration as tenant-isolation configuration and audit it the way you would audit FC zoning.","references":["https://nvd.nist.gov/vuln/detail/CVE-2023-53675","https://nvd.nist.gov/vuln/detail/CVE-2023-53431","https://www.t10.org/"],"status":"curated"},{"id":"NCVD-2026-027-modbus-tcp-as-an-unauthenticated","cve":null,"aliases":[],"title":"Modbus TCP as an unauthenticated control channel on facility gear: Modbus TCP has no authentication, no authorization and no integrity checking. A write-single-register from…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Modbus TCP as an unauthenticated control channel on facility gear","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Modbus TCP has no authentication, no authorization and no integrity checking. A write-single-register from any host on the segment is indistinguishable from a legitimate one. In a datacenter Modbus is the lingua franca for exactly the equipment you least want touched: CDU and liquid-cooling controllers, chiller and CRAH controllers, generator and transfer-switch controllers, power meters, and the gateways that expose all of the above to the BMS and DCIM. Where a register is writable, an attacker can command an ATS to transfer, a chiller to stop, or a CDU pump to change speed. Where it is read-only, they can still feed false telemetry to the systems that make automated decisions from it. The physical consequence in a GPU hall is immediate: coolant flow or air handling commanded away from 40-140 kW racks means thermal shutdown of the fleet in minutes and hardware stress from the cycle. Modbus DoS is equally cheap - malformed frames crash many embedded Modbus stacks outright, as the Socomec DIRIS Digiware cluster shows.","attack_vector":"Any host that can reach TCP 502 on the device. No credentials. Many facility devices additionally expose Modbus RTU tunnelled over TCP on non-standard ports, which operators forget to inventory. The gear is on the facility VLAN, generally managed by the landlord or the mechanical contractor rather than by the datacenter operator, and frequently reachable from the DCIM collector - so a compromised monitoring server is a direct control path. Internet-exposed Modbus on port 502 remains a standing Shodan finding for building and industrial gear.","remediation":"Unpatchable by design. Controls in order of effectiveness: put Modbus devices on a dedicated segment with a default-deny policy that permits TCP 502 only from the specific poller address; where the device supports it, disable Modbus writes entirely and run read-only, which many facility integrations do not actually need; where a Modbus-to-BACnet or Modbus-to-SNMP gateway exists, terminate Modbus at the gateway and never route it further; and monitor for write function codes (5, 6, 15, 16, 22, 23) on a segment that should only ever see reads - that is a cheap, high-signal detection most operators do not have. For leased space, the equipment and the segment belong to the landlord, so this becomes a contractual item: demand to know which facility devices speak Modbus, whether writes are enabled, and what filters exist in front of port 502.","references":["https://www.cisa.gov/news-events/cybersecurity-advisories/aa22-103a","https://nvd.nist.gov/vuln/detail/CVE-2024-48882","https://www.cisa.gov/resources-tools/resources/secure-design-alert-security-design-improvements-scada-devices"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-028-nvme-admin-command-set-firmware","cve":null,"aliases":["NVMe Firmware Image Download","NVMe Firmware Commit","firmware downgrade attack","unsigned drive firmware"],"title":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the host OS on bare-metal nodes: CLASS ENTRY, and the single highest-leverage…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"NVMe admin command set - Firmware Image Download (opcode 11h) and Firmware Commit (opcode 10h) reachable from the…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"CLASS ENTRY, and the single highest-leverage tenant-handoff risk in this whole category. The NVMe specification puts firmware download and firmware commit in the standard admin command set, meaning a host with root can push a firmware image to the drive using nothing more exotic than nvme-cli. Whether that is safe depends entirely on the drive enforcing signature verification and anti-rollback - and that enforcement is uneven across vendors and generations, is not something the host can independently verify, and has been shown breakable even where present (the Kioxia CM6/PM6/PM7 JTAG work bypasses the RSA signature check outright). Where verification is weak or absent, a departing tenant flashes a modified image and owns the drive controller permanently: it outlives your wipe, your reimage and your reprovision, sees everything the next tenant writes, and lies to every host-side tool about its own version, its lock state and whether a sanitize succeeded. Even where signature checking holds, downgrade to an older SIGNED image with known vulnerabilities is a live path unless the drive enforces anti-rollback - and the RPMB replay flaw (CVE-2020-13799) undermines exactly that anti-rollback state across eMMC, UFS and all NVMe versions.","attack_vector":"A tenant with root on the bare-metal host, issuing standard NVMe admin commands to a locally attached drive. No physical access, no exotic tooling, no exploit needed where the drive does not enforce signing - the command path is a documented, supported feature. Also relevant in SR-IOV and DPU/computational-storage designs, where whether the admin queue and Security Send/Receive are properly filtered from a tenant-controlled function is a per-platform question most operators have never actually tested.","remediation":"Policy and platform configuration; there is no patch because the command path is by design. (1) Block it at the platform: filter Firmware Image Download and Firmware Commit - and vendor-specific and Security Send/Receive pass-through - so tenant-controlled hosts and virtual functions cannot reach them. Do not expose raw NVMe admin queues to tenants unless the product genuinely requires it. (2) TEST it rather than assuming: on a representative node, try to flash a drive from a tenant-equivalent shell and confirm you are refused. Most operators have never run this check and will be surprised by the result on at least one SKU. (3) Make signed firmware and enforced anti-rollback a written procurement requirement, and get the vendor to state it per SKU. (4) Inventory expected firmware version per drive serial in a store the host cannot write to, and alert on any change or any version that decreases. (5) For sensitive tenancies, retire local media at end of tenancy instead of recycling it. Assume that verifying firmware integrity across a 10,000-drive fleet is not achievable with host-side tooling - the controller is the thing answering your questions - so the control has to be preventing the write and controlling the media's lifecycle, not detecting the implant afterwards.","references":["https://www.nvmexpress.org/wp-content/uploads/NVM-Express-Base-Specification-2.0c-2022.10.04-Ratified.pdf","https://github.com/google/security-research/security/advisories/GHSA-3hh8-94j4-62rh","https://www.kb.cert.org/vuls/id/231329","https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-88r1.pdf"],"status":"curated","fleet":{"pain_class":"unpatchable / mitigate-only"}},{"id":"NCVD-2026-030-raid-hba-controller-firmware-upd","cve":null,"aliases":[],"title":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip Adaptec SmartRAID/SmartHBA, and their OEM rebadges (Dell PERC, HPE SR/Smart Array…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"RAID/HBA controller firmware update path as a class - Broadcom MegaRAID and LSI 9400/9500/9600 HBAs, Microchip…","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"On most of these SKUs the controller firmware can be written from the running host OS by a root process using the vendor CLI (StorCLI, arcconf, perccli, ssacli) over the normal driver ioctl path. There is no host-verifiable attestation of what firmware the controller is actually running - the version string you read back is reported by the same firmware you are trying to verify. The controller is a PCIe device with DMA to host memory that sits below the operating system and below the boot chain, and its firmware is not covered by UEFI Secure Boot or by any measured-boot chain that a normal reimage re-establishes. For a bare-metal GPU rental, that means a tenant with root - which every bare-metal tenant has - can write persistent code beneath the next tenant's operating system, and a wipe-and-reimage handoff does not remove it. This is the concrete mechanism behind 'bare-metal tenant handoff is not a reimage', on a component operators rarely inventory at all.","attack_vector":"Local root on the bare-metal host: the legitimate tenant during their rental, or anyone who achieved root through any other path. No physical access, no BMC access, and no reboot required to stage the flash on most controllers.","remediation":"Not patchable and largely unmitigated on current SKUs. What operators can actually do: (1) make controller firmware version and checksum part of the handoff checklist and reflash from a vendor-signed image between tenants rather than trusting the reported version - budget the reboot into the OEM update utility and the node drain, and expect the OEM package to trail Broadcom/Microchip by months; (2) prefer platforms where the controller participates in a platform root of trust that the BMC can attest, and make that a procurement requirement rather than a hope; (3) blacklist or restrict the management ioctl path from tenant workloads where the workload does not need it; (4) accept and price the residual risk for SKUs where the firmware cannot be independently verified, and keep those nodes out of the pool you offer for security-sensitive tenants.","references":["https://www.broadcom.com/support/resources/product-security-center","https://www.microchip.com/en-us/solutions/embedded-security/how-to-report-potential-product-security-vulnerabilities"],"status":"curated","fleet":{"pain_class":"firmware-flash"}},{"id":"NCVD-2026-033-wiegand-reader-to-controller-wir","cve":null,"aliases":[],"title":"Wiegand reader-to-controller wiring and legacy 125 kHz proximity / MIFARE Classic credentials: Two structural weaknesses in nearly every badge system installed before roughly the last decade, and…","layer":"firmware-bmc-fabric","layer_name":"Firmware, BMC & network fabric","component":"Wiegand reader-to-controller wiring and legacy 125 kHz proximity / MIFARE Classic credentials","year":"2026","cvss_score":null,"severity":"unscored","kev":false,"impact":"Two structural weaknesses in nearly every badge system installed before roughly the last decade, and plenty installed since. First, the Wiegand protocol between the card reader and the door controller is unauthenticated, unencrypted plaintext over a few wires - anyone who can reach the back of the reader, which is on the unsecured side of the door, can splice in a small logger to capture every credential presented, or inject a previously captured credential to open the door. Second, 125 kHz proximity cards and MIFARE Classic credentials are cloneable in seconds with a sub-hundred-dollar reader from a pocket's distance, so an attacker who stands near an employee in a coffee shop can walk into the hall an hour later. Neither of these produces an anomalous event: the badge system records a valid credential at a valid door at a plausible time. What follows from being inside the cage is the whole point - drives with model weights and customer data, server console ports, an unlocked out-of-band switch, and the ability to plant a hardware implant that survives everything your software security program looks at. For a multi-tenant operator this also defeats the cage boundary you sell to customers, and it defeats it in a way that leaves no forensic trace.","attack_vector":"Physical presence. For Wiegand tapping: brief access to the reader housing, which is mounted outside the secured area by definition and usually held on with a security screw. For card cloning: proximity to any credential holder, or access to a credential left in a desk. No network access is required for either, which is exactly why network-centric security programs miss them. Note that many datacenter cages are protected by a single badge reader with no second factor, and that contractor and landlord staff badges frequently open more doors than the tenant realises.","remediation":"Not patchable - these are design properties of the installed hardware. The real fixes are hardware replacements and they cost money: move reader-to-controller communication from Wiegand to OSDP v2 with Secure Channel (encrypted and mutually authenticated), which requires readers and controllers that support it and a rewiring pass; and migrate credentials from 125 kHz prox and MIFARE Classic to a cryptographic credential (DESFire EV2/EV3 with a properly managed site key, or mobile credentials) which requires new cards and new readers. In the meantime: add a second independent factor at the hall and cage doors (PIN pad or biometric, on a separate system from the badge reader), fit tamper switches on reader housings and alarm on them, put a camera covering every cage door with retention long enough to be useful, and run a periodic reconciliation of who actually holds a badge that opens your hall - including landlord and contractor staff. For leased space, cage-door credential technology is a lease negotiation item; ask what credential format is in use and treat '125 kHz prox' as a finding.","references":["https://www.cisa.gov/news-events/ics-advisories/icsa-23-061-03","https://nvd.nist.gov/vuln/detail/CVE-2022-40633","https://www.securityindustry.org/industry-standards/open-supervised-device-protocol/"],"status":"curated"}]}