Files
bwn-a1370 a8418431bb bwn BCM43224 (MacBook Air A1370) driver fixes for FreeBSD 15.1
Patches + build scripts + firmware sources to make the Broadcom BCM43224
(PCI 0x14e4:0x4353, subvendor 0x106b:0x00d1) work on FreeBSD 15.1.

Fixes:
- Fix 1: compile GPL N-PHY code (BWN_GPL_PHY build wiring)
- Fix 2: NULL-guard bwn_dma_ringfree on partial attach
- Fix 3: wire bwn_update_mcast to ic->ic_update_mcast
- Fix 5: load bwn_v4_n firmware at loader stage (loader.conf)
- Fix 6: reclaim pending TX frames at bwn_dma_stop (detach double-free)
- Fix 7: driver-owned node ref + remove dr_usedslot early-return +
        bhndb_pci BHNDB_PCI_QUIRK_NO_MSI for BCM43224 (legacy INTx)
- Fix 8: restrict reclaim to TX rings only (RX reclaim caused heap
        corruption / TCP panic)

README.md  = rebuild guide (apply.sh + build.sh, manual steps)
RESEARCH.md = root-cause investigation log (Fixes 5-8)
2026-08-16 16:08:44 +03:00

134 lines
6.7 KiB
Markdown

# Research & Root-Cause Log
This document records the investigation that produced the fixes in this
repository. It is kept separate from `README.md` (the rebuild guide) so the
how-to stays clean. Each section maps to a fix number used throughout the
project.
Hardware under test: Apple MacBook Air A1370, Broadcom BCM43224
(PCI `0x14e4:0x4353`, subvendor `0x106b:0x00d1`), N-PHY rev-23 core behind a
BCMA/bhnd bus attached via the `bhndb_pci` PCIe-G1 bridge, on an NVIDIA MCP89
host bridge. FreeBSD 15.1-RELEASE-p2 (amd64).
---
## Fix 5 — Load Broadcom v4 firmware at loader time (firmware "not found" panic)
**Problem.** `if_bwn.ko` has a `pci/bwn_pci` devmatch modalias, so the kernel
AUTO-LOADS it when `bwn_pci0` attaches. The firmware-only `bwn_v4_n.ko` has NO
modalias and loads only via `kld_list`. If the firmware module is missing at
attach time (e.g. `rc.conf` `kld_list` got corrupted/shortened),
`bwn_fw_get()` fails and the driver panics while cleaning up
("firmware bwn_v4_ucode16_mimo not found").
**Fix.** Add `bwn_v4_n_load="YES"` to `/boot/loader.conf` so the firmware
module is loaded at the loader stage, before any device attach. This makes
firmware availability independent of `rc.conf`/`kld_list` ordering.
**Verification.** With `bwn_v4_n` loaded, creating `wlan1` on `bwn0` loads
`ucode16_mimo` with no "not found" error.
---
## Fix 6 — Kernel panic on kldunload after a "device timeout" (vmcore.4)
**Symptom.** `kldunload if_bwn` (or any detach while a wlan existed on bwn0)
panicked with a page fault in `node_cleanup` / `ieee80211_ageq_remove`
(fault addr `0x488`), always after a `bwn0: device timeout`.
**Root cause.** TX never completes because MSI interrupts are not delivered,
so the TX slot keeps its `mbuf` + `ieee80211_node` (`mt_ni`) reference —
`bwn_dma_handle_txeof` (the only place that clears `mt_ni`) never runs.
`bwn_dma_stop()` only halts the DMA engine and never frees queued frames, so
the node reference lingers. When net80211 later frees that node (scan cancel /
VAP teardown) the slot points at freed memory; then `bwn_detach` ->
`bwn_dma_free` -> `bwn_dma_ringfree` -> `bwn_dma_free_descbufs` ->
`ieee80211_free_node(stale mt_ni)` walks freed memory and panics.
**Fix.** Reclaim pending frames (and their node references) in
`bwn_dma_stop()`, which runs from both the detach path
(`bwn_stop` -> `bwn_core_exit`) and every HW reset. This drops node references
while the node is still alive (before `ieee80211_ifdetach` frees nodes).
`bwn_dma_free_descbufs()` is idempotent (NULLs `mt_m`/`mt_ni`), so the later
detach free finds empty slots.
---
## Fix 7 — TX node reference ownership + unconditional reclaim + MSI/INTx (vmcore.5)
**Symptom.** After Fix 6 a fresh reboot still panicked on `kldunload`
(vmcore.5): page fault in `ieee80211_ratectl_node_deinit` (fault addr `0x30`,
`vap->iv_rate` deref), `ni->ni_refcnt = 1`, node's `ni_vap` already torn down.
Also, with MSI never delivered, the chip showed "device timeout" and passed no
traffic.
**Root cause.** The driver relied on the mbuf's `rcvif` node reference (owned
by the net80211 stack) for the lifetime of a queued TX frame. The stack can
release that reference at any time (scan cancel, VAP teardown, node reclaim)
while the frame is still in the ring, because TX completion never runs when
interrupts are not delivered. Also `bwn_dma_free_descbufs()` early-returned on
`if (!dr->dr_usedslot)`; `dr_usedslot` is only a free-slot accounting hint and
is not a reliable indicator of pending frames, so the reclaim could be
skipped.
Separately, `bhndb_pci` enables MSI for the BCM43224, but on this host (NVIDIA
MCP89 Apple MacBook Air) MSI is never delivered.
**Fix.**
1. `bwn_dma_tx_start()` / `bwn_pio_tx_start()` now take the driver's OWN
reference via `ieee80211_ref_node(ni)` before stashing it in `mt->mt_ni` /
`tp->tp_ni`. `ieee80211_tx_complete()` in the txeof path releases exactly
one reference, so accounting stays balanced and detach cannot double-free.
2. Removed the `if (!dr->dr_usedslot) return;` early-return from
`bwn_dma_free_descbufs()` so reclaim unconditionally walks every slot.
3. Added `BHNDB_PCI_QUIRK_NO_MSI` for BCM43224 in `bhndb_pci.c` /
`bhndb_pcivar.h` so the bridge falls back to legacy INTx — TX interrupts
are then delivered and the "device timeout" root cause is gone.
---
## Fix 8 — Remove RX-ring reclaim from bwn_dma_stop (vmcore.6 heap corruption)
**Symptom.** After Fix 6 the bwn *detach* panic was gone (clean `bwn0:
detached` cycles), but ~47 min into a session the box panicked with a
DIFFERENT crash: a TCP stack NULL-deref inside `tcp_m_copym` while processing
an inbound packet on wlan0/rtwn0 (the SSH link). The SSH socket's send buffer
was corrupted (`so_snd.sb_mb == NULL` while `sb_ccc == 126404`; `tp` seq
numbers garbage).
**Root cause.** Fix 6's reclaim in `bwn_dma_stop()` walked **all six rings
including the RX ring**. `bwn_dma_stop()` runs on **every** `bwn_hwreset`
(every "device timeout" — and on this chip MSI TX interrupts are not
delivered, so timeouts fire constantly). `bwn_dma_free_descbufs()` frees the
RX mbuf WITHOUT clearing the DMA descriptor's physical address. The RX ring is
persistent (refilled in place by `bwn_rxeof()` -> `bwn_dma_newbuf()`, never
re-allocated) and `bwn_dma_cleanup()` only zeroes the ring *control* register.
So after a reset the descriptor still pointed at the freed page; when RX was
re-enabled the NIC DMAd received frames into freed kernel memory -> heap
corruption that eventually clobbered the SSH socket. The console log showed
~60 lines of `bwn0: bwn_dma_free_descbufs: not TX?` (one per RX slot whose
mbuf had just been freed) immediately before the panic — the smoking gun.
**Fix.** Restrict the reclaim in `bwn_dma_stop()` (factored into a new helper
`bwn_dma_reclaim_tx()`) to the **TX rings only** (wme BK/BE/VI/VO + mcast).
The RX ring is never reclaimed there. `bwn_dma_reclaim_tx()` is also called
from `bwn_detach()` BEFORE `ieee80211_ifdetach()` (defence-in-depth on top of
the ref_node fix in Fix 7).
**Verification.** Disassembly: `bwn_dma_stop` has 6 ringstops (RX + 5 TX) +
exactly 5 `bwn_dma_free_descbufs` calls (TX only; RX ring no longer
reclaimed).
---
## Why the MSI quirk matters (Fix 7, point 3)
On this host the `bhndb_pci` bridge enables MSI for the BCM43224, but the
interrupts are never delivered (no `bwn`/`MSI` line in `vmstat -i`). Because
the only place that clears a TX slot's node reference is the TX-completion
path (`bwn_dma_handle_txeof`), TX frames never complete -> "device timeout"
and the cascade of use-after-free / heap-corruption panics. Forcing legacy
INTx (the `NO_MSI` quirk) makes TX interrupts arrive, so completions run
normally. The reclaim/ref_node fixes (6-8) are defence-in-depth that keep the
driver robust even if completions are delayed.