Repository navigation
fix: disambiguate empty service-filtered scans from adapter wedges - #49
Conversation
The discovery scan is service-filtered (only devices advertising the Reticulum service UUID fire the detection callback). At boot, no peer has started advertising yet, so the filtered scan legitimately returns zero callbacks while the peer's GATT server comes up. The detector treated 3 empty filtered scans as 'adapter corrupted / system reboot required' and fired on_error(critical) - a false positive that tore the BLE interface down in the exact window the peer was about to come online. Observed on two Pi Zero 2 W (BCM43430): both wedged at boot (~3-5 empty filtered scans), then recovered on their own - which a genuinely wedged adapter can never do without a reboot. Confirms the detector was misattributing 'no peer advertising yet' to 'adapter corrupted'. Fix: when a service-filtered scan returns zero callbacks, run a short unfiltered health scan (BleakScanner.discover, no service filter). If it sees any device, the adapter is alive and we are simply waiting for a Reticulum peer - reset the streak, do NOT fire critical. Only a fully blind adapter (unfiltered scan also empty, or the health scan itself raises) counts toward the genuine-wedge streak. This is detection-logic only; it does not change scan cadence, the service filter, or connection behavior.
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
|
test_empty_filtered_but_healthy_adapter_is_not_a_wedge asserted on m.BleakScanner.discover.await_count AFTER the patch.object context had exited, at which point m.BleakScanner has reverted to the real class whose .discover is a plain function (AttributeError: 'function' object has no attribute 'await_count'). This failed CI's integration suite on every python version. Bind the mock to a local (BS) and assert on BS.discover.await_count. Verified all 5 tests pass with CI deps (bleak + pytest-asyncio + rns).
…gression; guard health scan against mid-scan connection Addresses both Greptile P1 findings on PR #49: 1. Overclaim: an empty UNFILTERED scan in a genuinely quiet RF environment (no BLE devices in range at all) is not proof the adapter is broken - both scans legitimately return nothing. Track whether the adapter was ever proven healthy (healthy_ever) and only escalate to the 'system reboot required' critical when a previously-healthy adapter goes blind (a true regression). A never-healthy adapter that sees nothing warns without mandating a reboot. A health scan that FAILS (D-Bus/adapter error, return -1) is distinct from a clean zero and still escalates, since a scan that cannot run is direct fault evidence. 2. Race: a connection that starts during the main-scan window is not caught by the pause check at the top of _perform_scan. Re-check _should_pause_scanning() before running the second (unfiltered) health scan so it cannot collide with an active connection ('Operation already in progress'). Tests: +quiet-RF-benign, +was-healthy-then-blind-wedge, +health-scan-failure, +mid-scan-connection-suppresses-health-scan, +successful-scan-marks-healthy. All changed driver lines covered; 306 integration tests green.
| # and mandating a system reboot would be a false alarm. A truly | ||
| # wedged adapter is one that USED to see devices and no longer | ||
| # can, so gate the critical on healthy_ever. | ||
| if self.consecutive_empty_scans >= 3 and self.healthy_ever: |
There was a problem hiding this comment.
Startup scan fault stays hidden
If an adapter is already blind when the driver starts but both scans finish without errors, healthy_ever stays false. Even after many empty scans, the driver only warns and never raises the critical fault signal. A real startup fault that returns empty results can leave discovery dead without that signal. The driver needs a way to distinguish this case from a quiet room.
Knowledge Base Used: Linux adapter recovery and error handling
Prompt To Fix With AI
This is a comment left during a code review.
Path: src/ble_reticulum/linux_bluetooth_driver.py
Line: 761
Comment:
**Startup scan fault stays hidden**
If an adapter is already blind when the driver starts but both scans finish without errors, `healthy_ever` stays false. Even after many empty scans, the driver only warns and never raises the critical fault signal. A real startup fault that returns empty results can leave discovery dead without that signal. The driver needs a way to distinguish this case from a quiet room.
**Knowledge Base Used:** [Linux adapter recovery and error handling](https://app.greptile.com/torlando-tech/-/custom-context/knowledge-base/torlando-tech/ble-reticulum/-/docs/linux-adapter-recovery.md)
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.…owered state Addresses both Greptile P1 findings on PR #49, which were the symmetric pair of the same flaw: the previous healthy_ever heuristic (device-count based) fails in BOTH directions. Issue 1: an adapter dead from boot (empty scans, no D-Bus error, never 'healthy') kept healthy_ever false, so it only warned and never fired the critical fault signal - discovery could be dead with no signal. Issue 2: healthy_ever latched true after any single healthy observation, so a working adapter in a quiet room (peer walked out of range) would false-fire the 'reboot required' critical after 3 clean-zero scans. Root cause: device count CANNOT distinguish a genuinely quiet RF environment (healthy adapter, no BLE devices in range) from a dead/wedged adapter (no devices because the adapter is down) - both produce 'empty, no error.' Fix: use the adapter's OWN BlueZ 'Powered' state (a new _adapter_is_powered() D-Bus probe) as the independent discriminator, since it does not depend on what the adapter can see: * Powered=False (adapter present but not powered) -> genuine fault -> fire the reboot/power-on critical. Catches both an adapter that was never healthy and one that went down (fixes Issue 1). * Powered=True (working adapter in a quiet room) OR unknown (could not determine) -> no positive fault evidence -> warn only, never mandating a reboot, so a healthy self-recovering adapter is never torn down over a quiet environment (fixes Issue 2). * a health scan that RAISES (can't run) is still direct fault evidence -> critical, independent of Powered. This removes the flawed device-count heuristic entirely. The powered-off critical now also suggests 'bluetoothctl power on' (often a non-reboot fix), and a healthy adapter that self-recovers in a quiet room is left alone - matching the operating requirement that a clean radio room is realistic and the stack must recover without a forced reboot. Tests: rewrote the wedge-detection suite around _adapter_is_powered (not-powered fires, powered/unknown warn, scan-failure fires, mid-scan-connection skips, successful-scan resets) and added TestAdapterIsPowered driving the real D-Bus probe with a mocked bus (powered-true/false, query-error=unknown, no-dbus=unknown, bus always disconnected). All changed driver lines covered; 311 integration tests green.
…t wedge discovery Addresses Greptile P1 finding on PR #49 (iteration 3): _adapter_is_powered() is called from inside _perform_scan after 3 empty scans, but its D-Bus round-trip had no overall timeout. If BlueZ stops replying (e.g. the adapter is wedged at the D-Bus level), the query could hang and discovery would stay stuck instead of retrying. Fix: wrap the whole query in asyncio.wait_for with a bounded wall-clock timeout (POWERED_QUERY_TIMEOUT_S = 5s, a module constant so tests can shorten it). A timeout is treated as an UNKNOWN power state (no positive fault evidence -> warn only), matching the existing no-dbus / query-error handling, and returns control to the scan loop. The bus is connected and disconnected inside the bounded coroutine's own finally, so cancellation on timeout still releases the D-Bus connection. This also closes the dead-from-boot gap the previous review flagged: a dead-from-boot adapter that is Powered=False now fires the critical fault signal (test_not_powered_adapter_declares_wedge_after_3), so discovery being dead is no longer silent - while a powered adapter in a genuinely quiet room (or an indeterminate state) still only warns, never false-rebooting. Test: +test_hung_bluez_times_out_as_unknown - connect() made to sleep 60s, timeout patched to 0.05s; asserts the query returns None promptly (not hung) and does not stall the scan loop. All changed lines covered; 312 integration tests green.
Adds the Keep-a-Changelog entry for the v0.2.3 release covering the scanner wedge health-check, the ifac_size inheritance fix that was dropping all inbound BLE packets, announce-rate stats, duplicate- identity normalization, the package rename, and related fixes (#45, #47, #48, #49 plus the #29-#44 window). This was the missing precondition blocking the v0.2.3 release workflow: the tag points at a commit where pyproject.toml still reads 0.2.2 and no CHANGELOG.md entry exists. Co-authored-by: torlando-tech <torlando-tech@users.noreply.github.com>
What
The discovery scan is service-filtered - only devices advertising the Reticulum service UUID trigger the detection callback. At boot, no peer has started advertising yet, so the filtered scan legitimately returns zero callbacks while the peer's GATT server comes up.
The wedge detector treated 3 empty filtered scans as 'adapter corrupted / system reboot required' and fired
on_error('critical')- a false positive that tore the BLE interface down in the exact window the peer was about to come online.Symptom (observed on two Pi Zero 2 W, BCM43430)
After reboot, both Pis logged:
...within ~15s of boot, then recovered on their own. A genuinely wedged adapter can never recover without a reboot, so the detector was misattributing 'no peer advertising yet' to 'adapter corrupted' - and the resulting
on_error('critical')was detaching the interface right when the peer was coming up.Fix
When a service-filtered scan returns zero callbacks, run a short unfiltered health scan (
BleakScanner.discover(timeout=1.0), no service filter):This is detection-logic only. It does not change scan cadence, the service filter, or connection behavior. A real wedge (adapter blind to all devices) is still detected and still reports 'reboot required'.
Tests
tests/test_adapter_health_check_wedge.py- 5 tests driving the real_perform_scan()with a mockedBleakScanner:_adapter_health_checkreturns correct device countAll pass locally (verified via direct event-loop run); CI will run under pytest-asyncio.
Scope
src/ble_reticulum/linux_bluetooth_driver.py(+74/-11), new test file (+183). No changes toBLEInterface.py, so no impact on Android/iOS - this is Linux driver-only wedge detection.