Intel Gaudi / vLLM hardware plugin: Malformed input to the Gaudi vLLM plugin crashes or wedges the serving process
Impact
Malformed input to the Gaudi vLLM plugin crashes or wedges the serving process. On an inference fleet this is a request-triggered outage of the model server rather than a data-confidentiality problem, but a single tenant can repeatedly take down a shared endpoint that is pinned to expensive accelerators.
Who can reach it
Anyone who can reach the vLLM endpoint with an authenticated request - so any tenant of a shared inference service, or any workload inside the cluster if the endpoint is not network-segmented.
What to do
Upgrade the vLLM hardware plugin for Gaudi to 0.16.0 or later. Pure Python/userspace package update - restart the serving process, no node reboot, no firmware.
References
This entry is curated: imported from vendor advisories with machine assistance, not yet individually verified. Confirm against your vendor's advisory before acting, and report anything wrong.