llama.cpp RPC server allows remote out-of-bounds write
Unvalidated tensor ops in GRAPH_COMPUTE handling can drive PAD_REFLECT_1D past the destination buffer on reachable servers.
A vulnerability report against llama.cpp’s ggml RPC server describes a remote out-of-bounds write reachable when the service accepts untrusted clients.
The server rebuilds and runs compute graphs from network GRAPH_COMPUTE messages. Deserialization of tensor nodes does not validate the operation type or its parameters, so a crafted request can force an illegal PAD_REFLECT_1D path that writes past the destination tensor. Write location, length, and contents are attacker-controlled. The flaw shows up in release builds of the RPC backend used to offload inference to a remote machine.
Anyone exposing ggml-rpc-server beyond a trusted network is in scope: a remote peer can corrupt memory in the server process without authentication beyond reachability of the RPC port. The report, filed by ZZ2266 against ggml-org/llama.cpp, ties the issue to missing checks on inbound graph payloads rather than a local-only misuse of the padding kernel.
Operators should keep RPC endpoints off untrusted networks until the deserialization path rejects illegal ops and bounds, and treat any internet-facing deployment as high risk.