freenode
AI & ML

PyTorch CPU triu/tril write out of bounds on strided batch out

Non-contiguous batch dimensions on the out tensor make the CPU path step wrong and can corrupt the heap; CUDA is fine.

PyTorch's torch.triu and torch.tril can write out of bounds on CPU when the caller supplies an out tensor whose batch dimensions are not contiguous, according to a bug report against the project.

The CPU path flattens a batch of matrices and advances between them using only the innermost batch stride. That works when outer batch strides match a dense layout. It fails when an outer batch dimension has been permuted or otherwise strided, so the kernel walks the wrong addresses. The same inputs on CUDA fill the tensor correctly.

A small reproduction with a permuted zeros buffer as out produced a wrong partial result on CPU and then aborted with heap corruption ("double free or corruption"). Both the upper- and lower-triangular ops are affected.

Anyone relying on in-place or out-of-place triangular fills into a non-contiguous batched CPU tensor is exposed. Contiguous batch layouts and the GPU implementations are not.