bnxt_en DMA faults after link-layer headroom change
A latent short-packet padding bug in Broadcom’s NetXtreme-E driver, exposed by a kernel headroom bump, was taking interfaces down within minutes.
A regression in the Broadcom bnxt_en driver is causing IOMMU DMA faults and repeated interface teardown on NetXtreme-E hardware, including Dell systems with the BCM57412. The fault was bisected by Stefan Fleischmann to a kernel change that raised reserved link-layer headroom from 48 to 64 bytes. Interfaces came up normally after boot, then failed after anywhere from a few minutes to under an hour, recoverable only until the next fault.
Eric Dumazet traced the symptom to a DMA read overrun: the NIC’s DMA engine was reading past the end of a mapped buffer into an adjacent unmapped page, often landing on a 4KB boundary. The headroom increase did not introduce a new DMA rule so much as expose a long-standing mistake in how bnxt padded undersized frames before mapping them for transmit. Offload toggles (TSO, VLAN, USO) did not help. The problem reproduced on current net development kernels and without containers running, though Fleischmann first saw it on untagged plus VLAN interfaces used with macvlan and LXC.
The driver had been padding short packets without consistently tying the padded length into the DMA map and transmit descriptor. Dumazet posted a fix that switches the path to skb_put_padto() so length and head length stay in sync with the hardware view. Michael Chan noted that an intermediate approach would still leave the hardware seeing a frame that was too short. Fabian Grünbichler reported that Proxmox had already seen multiple users hit the same failure after picking up related 7.2-stable updates. Fleischmann’s multi-hour soak of the corrected path stayed clean.
The patch is under review on the netdev list. Operators on bnxt_en with recent stable or distribution kernels who see unexplained DMA faults and link flaps should watch for the fix in forthcoming updates.