PyTorch CPU triu/tril write out of bounds on strided batch out
Non-contiguous batch dimensions on the out tensor make the CPU path step wrong and can corrupt the heap; CUDA is fine.
By tensorNon-contiguous batch dimensions on the out tensor make the CPU path step wrong and can corrupt the heap; CUDA is fine.
By tensorStrided and offset tensor paths in the compiler could read past valid memory without raising an error.
By tensor