PyTorch CUDA pool teardown can re-enter half-destroyed state
Review of pending allocator work also flags null returns that can produce CUDA tensors backed by address zero.
Pending PyTorch changes around CUDA graph memory pools drew a request-for-changes review after two high-priority correctness hazards turned up in the allocator and callback path.
The first problem is re-entrancy during pool teardown. The allocator can own Python callback state, so destroying an empty graph pool may drop the last allocator reference while the allocator lock is still held and before freeable-pool bookkeeping is cleared. A callback destructor that calls torch.cuda.empty_cache() can therefore re-enter cleanup against a pool whose block containers are already gone. The deferred-free queue does not cover this case when there are no segments left to free. The review asks that allocator destruction wait until both maps are updated and the lock is released, plus a regression that deletes an otherwise unused pool while an unreferenced callable’s destructor empties the cache.
The second problem is silent acceptance of allocation failure on the process-wide pluggable allocator path. A null or zero result for a nonzero request bypasses ordinary out-of-memory handling and can be wrapped as a live data pointer, yielding a nonempty CUDA tensor with no backing storage. Later device work may then touch address zero. Null results for nonzero allocations should be rejected on the direct path while MemPool retry behavior stays intact, and the process-wide subprocess coverage should assert torch.OutOfMemoryError when callbacks return None or zero.
Edward Yang’s review, posted via an assistant on his behalf, based both findings on source inspection; CUDA execution was unavailable to confirm at runtime.