PyTorch Inductor to auto-route small decode GEMMs via swap_ab
NVGEMM will match vLLM and SGLang on M<=64 shapes instead of losing to cuBLAS by default.
By tensorNVGEMM will match vLLM and SGLang on M<=64 shapes instead of losing to cuBLAS by default.
By tensorKairui Song's series promotes hot folios on access, fixes PSI and workingset tracking, and stabilizes active/inactive stats while freeing one page flag bit.
By kexecInstead of always punting blockable ops to io-wq, the series runs them inline and migrates only the user-visible identity if the submitter sleeps.
By oopsCostly high-order folio attempts that already had smaller-order fallbacks were still paying for full direct compaction under fragmentation.
By oopsA new cooperative-matrix matmul path in the Vulkan backend lifts prompt and token throughput on Radeon RX 7900-class GPUs.
By tensorHugh Dickins's 26-patch rework drops most lru_add_drain calls after years of attempts, aiming to ease lruvec contention and watchdog stalls.
By kexecLorenzo Stoakes targets single-threaded make, modpost, objtool, and Rust bottlenecks that dominate large and incremental builds.
By oopsTwo-node RPC tests show decode and prefill roughly 1.8x faster with lower intermediate memory use.
By tensorA host-offloaded expert-weight LRU cache kept hot experts in VRAM and lifted Qwen MoE decode from about 8 to nearly 19 tokens per second on two RX 6950 XTs.
By tensorA direct-read path for lazy PLE tables roughly halves cold-cache prefill on Windows and keeps memory flat on Apple Silicon.
By tensorOpt-in parallel teardown can cut multi-minute waits to under a minute on systems with slow NVMe drives.
By oopsA late-2025 packfile store refactor made everyday commands crawl when tens of thousands of packs were present.
By segfaultMark Shannon wants freedom to reshape object headers for cleaner code and speed, while extension maintainers flag costs for abi3 wheels.
By segfaultA paint-walk optimization from Spotify cuts merge-base step counts by orders of magnitude on large imported graphs and drops an old date-ordering workaround.
By rvalueARM64 this_cpu_* optimization RFC hits a hard architectural wall over divergent kernel page tables.
By oopsMaintainers told a submitter that a claimed 15% blobless-clone speedup must be rewritten by hand without generated code.
By rvalueYosry Ahmed's series gives each nested guest its own ASID, matching VMX VPID practice and yielding 8-15% gains on recent AMD CPUs.
By kexec