Skip to content

fix: disambiguate empty service-filtered scans from adapter wedges - #49

Merged
torlando-tech merged 5 commits into
mainfrom
fix/ble-scanner-wedge-health-check
Oct 1, 2026
Merged

torlando-tech merged 5 commits into
mainfrom
fix/ble-scanner-wedge-health-check

Conversation

@torlando-tech

Copy link
Copy Markdown
Owner

What

The discovery scan is service-filtered - only devices advertising the Reticulum service UUID trigger the detection callback. At boot, no peer has started advertising yet, so the filtered scan legitimately returns zero callbacks while the peer's GATT server comes up.

The wedge detector treated 3 empty filtered scans as 'adapter corrupted / system reboot required' and fired on_error('critical') - a false positive that tore the BLE interface down in the exact window the peer was about to come online.

Symptom (observed on two Pi Zero 2 W, BCM43430)

After reboot, both Pis logged:

Scanner corruption detected: 0 callbacks after 1.0s scan (streak: 1)
... (streak: 2, 3)
CRITICAL: Bleak scanner callbacks not firing
System reboot required to restore BLE scanning
[BLE Interface] detached

...within ~15s of boot, then recovered on their own. A genuinely wedged adapter can never recover without a reboot, so the detector was misattributing 'no peer advertising yet' to 'adapter corrupted' - and the resulting on_error('critical') was detaching the interface right when the peer was coming up.

Fix

When a service-filtered scan returns zero callbacks, run a short unfiltered health scan (BleakScanner.discover(timeout=1.0), no service filter):

  • if it sees any device -> adapter is alive, just waiting for a peer -> reset streak, no critical
  • if it sees nothing (or the health scan raises) -> counts toward the genuine-wedge streak -> critical after 3

This is detection-logic only. It does not change scan cadence, the service filter, or connection behavior. A real wedge (adapter blind to all devices) is still detected and still reports 'reboot required'.

Tests

tests/test_adapter_health_check_wedge.py - 5 tests driving the real _perform_scan() with a mocked BleakScanner:

  1. empty-filtered-but-healthy -> NOT a wedge (the RED->GREEN regression: old code fired critical here)
  2. fully-blind adapter -> critical after 3
  3. blind-then-healthy -> streak resets, no critical
  4. health-scan exception -> counts as blind (fail closed)
  5. _adapter_health_check returns correct device count

All pass locally (verified via direct event-loop run); CI will run under pytest-asyncio.

Scope

src/ble_reticulum/linux_bluetooth_driver.py (+74/-11), new test file (+183). No changes to BLEInterface.py, so no impact on Android/iOS - this is Linux driver-only wedge detection.

The discovery scan is service-filtered (only devices advertising the
Reticulum service UUID fire the detection callback). At boot, no peer has
started advertising yet, so the filtered scan legitimately returns zero
callbacks while the peer's GATT server comes up. The detector treated 3
empty filtered scans as 'adapter corrupted / system reboot required' and
fired on_error(critical) - a false positive that tore the BLE interface
down in the exact window the peer was about to come online.

Observed on two Pi Zero 2 W (BCM43430): both wedged at boot (~3-5 empty
filtered scans), then recovered on their own - which a genuinely wedged
adapter can never do without a reboot. Confirms the detector was
misattributing 'no peer advertising yet' to 'adapter corrupted'.

Fix: when a service-filtered scan returns zero callbacks, run a short
unfiltered health scan (BleakScanner.discover, no service filter). If it
sees any device, the adapter is alive and we are simply waiting for a
Reticulum peer - reset the streak, do NOT fire critical. Only a fully
blind adapter (unfiltered scan also empty, or the health scan itself
raises) counts toward the genuine-wedge streak.

This is detection-logic only; it does not change scan cadence, the
service filter, or connection behavior.
@codecov

codecov Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.87500% with 2 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/ble_reticulum/linux_bluetooth_driver.py 96.87% 2 Missing ⚠️

📢 Thoughts on this report? Let us know!

@greptile-apps

greptile-apps Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 4/5

[Medium risk] Refines Bluetooth adapter health detection logic.

The PR is not yet safe to merge because a blind scanner can remain powered and still never report a critical fault.

Fix All in Claude CodeFindings

  1. P1 Startup scan fault stays hidden ▶
Fix with agent prompt
### Issue 1
src/ble_reticulum/linux_bluetooth_driver.py:undefined-761
If an adapter is already blind when the driver starts but both scans finish without errors, `healthy_ever` stays false. Even after many empty scans, the driver only warns and never raises the critical fault signal. A real startup fault that returns empty results can leave discovery dead without that signal. The driver needs a way to distinguish this case from a quiet room.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The Linux Bluetooth driver now checks adapter health when a service-filtered scan finds no Reticulum peers. This separates a peer that has not started advertising from signs of an adapter fault.

  • Nearby BLE devices clear the empty-scan streak; repeated health-scan failures or an unpowered adapter can still trigger a critical error.
  • The PR adds tests for empty scans, connection pauses, and the BlueZ power-state check.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Service-filtered scan sees no peer] --> B[Unfiltered health scan]
  B -->|Sees a device| C[Reset empty-scan streak]
  B -->|Scan fails| D[Count a fault]
  B -->|Sees no device| E[Check adapter power after three scans]
  E -->|Not powered| F[Report critical fault]
  E -->|Powered or unknown| G[Warn without critical fault]
  E -->|Query times out| G
Loading

Reviews (7) · Last reviewed commit: "fix: bound the _adapter_is_powered() D-B..."

Comment thread src/ble_reticulum/linux_bluetooth_driver.py
Comment thread src/ble_reticulum/linux_bluetooth_driver.py
Comment thread tests/test_adapter_health_check_wedge.py Outdated
test_empty_filtered_but_healthy_adapter_is_not_a_wedge asserted on
m.BleakScanner.discover.await_count AFTER the patch.object context had
exited, at which point m.BleakScanner has reverted to the real class whose
.discover is a plain function (AttributeError: 'function' object has no
attribute 'await_count'). This failed CI's integration suite on every
python version.

Bind the mock to a local (BS) and assert on BS.discover.await_count.
Verified all 5 tests pass with CI deps (bleak + pytest-asyncio + rns).
…gression; guard health scan against mid-scan connection

Addresses both Greptile P1 findings on PR #49:

1. Overclaim: an empty UNFILTERED scan in a genuinely quiet RF environment
   (no BLE devices in range at all) is not proof the adapter is broken - both
   scans legitimately return nothing. Track whether the adapter was ever
   proven healthy (healthy_ever) and only escalate to the 'system reboot
   required' critical when a previously-healthy adapter goes blind (a true
   regression). A never-healthy adapter that sees nothing warns without
   mandating a reboot. A health scan that FAILS (D-Bus/adapter error,
   return -1) is distinct from a clean zero and still escalates, since a scan
   that cannot run is direct fault evidence.

2. Race: a connection that starts during the main-scan window is not caught by
   the pause check at the top of _perform_scan. Re-check _should_pause_scanning()
   before running the second (unfiltered) health scan so it cannot collide with
   an active connection ('Operation already in progress').

Tests: +quiet-RF-benign, +was-healthy-then-blind-wedge, +health-scan-failure,
+mid-scan-connection-suppresses-health-scan, +successful-scan-marks-healthy.
All changed driver lines covered; 306 integration tests green.
# and mandating a system reboot would be a false alarm. A truly
# wedged adapter is one that USED to see devices and no longer
# can, so gate the critical on healthy_ever.
if self.consecutive_empty_scans >= 3 and self.healthy_ever:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Startup scan fault stays hidden

If an adapter is already blind when the driver starts but both scans finish without errors, healthy_ever stays false. Even after many empty scans, the driver only warns and never raises the critical fault signal. A real startup fault that returns empty results can leave discovery dead without that signal. The driver needs a way to distinguish this case from a quiet room.

Knowledge Base Used: Linux adapter recovery and error handling

Prompt To Fix With AI
This is a comment left during a code review.
Path: src/ble_reticulum/linux_bluetooth_driver.py
Line: 761

Comment:
**Startup scan fault stays hidden**

If an adapter is already blind when the driver starts but both scans finish without errors, `healthy_ever` stays false. Even after many empty scans, the driver only warns and never raises the critical fault signal. A real startup fault that returns empty results can leave discovery dead without that signal. The driver needs a way to distinguish this case from a quiet room.

**Knowledge Base Used:** [Linux adapter recovery and error handling](https://app.greptile.com/torlando-tech/-/custom-context/knowledge-base/torlando-tech/ble-reticulum/-/docs/linux-adapter-recovery.md)

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Fix in Claude Code

…owered state

Addresses both Greptile P1 findings on PR #49, which were the symmetric pair
of the same flaw: the previous healthy_ever heuristic (device-count based)
fails in BOTH directions.

Issue 1: an adapter dead from boot (empty scans, no D-Bus error, never 'healthy')
kept healthy_ever false, so it only warned and never fired the critical fault
signal - discovery could be dead with no signal.

Issue 2: healthy_ever latched true after any single healthy observation, so a
working adapter in a quiet room (peer walked out of range) would false-fire the
'reboot required' critical after 3 clean-zero scans.

Root cause: device count CANNOT distinguish a genuinely quiet RF environment
(healthy adapter, no BLE devices in range) from a dead/wedged adapter (no devices
because the adapter is down) - both produce 'empty, no error.'

Fix: use the adapter's OWN BlueZ 'Powered' state (a new _adapter_is_powered()
D-Bus probe) as the independent discriminator, since it does not depend on what
the adapter can see:
  * Powered=False (adapter present but not powered) -> genuine fault -> fire the
    reboot/power-on critical. Catches both an adapter that was never healthy and
    one that went down (fixes Issue 1).
  * Powered=True (working adapter in a quiet room) OR unknown (could not
    determine) -> no positive fault evidence -> warn only, never mandating a
    reboot, so a healthy self-recovering adapter is never torn down over a quiet
    environment (fixes Issue 2).
  * a health scan that RAISES (can't run) is still direct fault evidence ->
    critical, independent of Powered.

This removes the flawed device-count heuristic entirely. The powered-off critical
now also suggests 'bluetoothctl power on' (often a non-reboot fix), and a healthy
adapter that self-recovers in a quiet room is left alone - matching the operating
requirement that a clean radio room is realistic and the stack must recover
without a forced reboot.

Tests: rewrote the wedge-detection suite around _adapter_is_powered (not-powered
fires, powered/unknown warn, scan-failure fires, mid-scan-connection skips,
successful-scan resets) and added TestAdapterIsPowered driving the real D-Bus
probe with a mocked bus (powered-true/false, query-error=unknown, no-dbus=unknown,
bus always disconnected). All changed driver lines covered; 311 integration tests
green.
Comment thread src/ble_reticulum/linux_bluetooth_driver.py Outdated
…t wedge discovery

Addresses Greptile P1 finding on PR #49 (iteration 3): _adapter_is_powered() is
called from inside _perform_scan after 3 empty scans, but its D-Bus round-trip
had no overall timeout. If BlueZ stops replying (e.g. the adapter is wedged at
the D-Bus level), the query could hang and discovery would stay stuck instead of
retrying.

Fix: wrap the whole query in asyncio.wait_for with a bounded wall-clock timeout
(POWERED_QUERY_TIMEOUT_S = 5s, a module constant so tests can shorten it). A
timeout is treated as an UNKNOWN power state (no positive fault evidence ->
warn only), matching the existing no-dbus / query-error handling, and returns
control to the scan loop. The bus is connected and disconnected inside the
bounded coroutine's own finally, so cancellation on timeout still releases the
D-Bus connection.

This also closes the dead-from-boot gap the previous review flagged: a
dead-from-boot adapter that is Powered=False now fires the critical fault signal
(test_not_powered_adapter_declares_wedge_after_3), so discovery being dead is no
longer silent - while a powered adapter in a genuinely quiet room (or an
indeterminate state) still only warns, never false-rebooting.

Test: +test_hung_bluez_times_out_as_unknown - connect() made to sleep 60s,
timeout patched to 0.05s; asserts the query returns None promptly (not hung)
and does not stall the scan loop. All changed lines covered; 312 integration
tests green.
@torlando-tech
torlando-tech merged commit ed164a8 into main Oct 1, 2026
14 checks passed
@torlando-tech
torlando-tech deleted the fix/ble-scanner-wedge-health-check branch October 1, 2026 13:42
torlando-tech added a commit that referenced this pull request Oct 3, 2026
Adds the Keep-a-Changelog entry for the v0.2.3 release covering the
scanner wedge health-check, the ifac_size inheritance fix that was
dropping all inbound BLE packets, announce-rate stats, duplicate-
identity normalization, the package rename, and related fixes (#45,
#47, #48, #49 plus the #29-#44 window).

This was the missing precondition blocking the v0.2.3 release workflow:
the tag points at a commit where pyproject.toml still reads 0.2.2 and
no CHANGELOG.md entry exists.

Co-authored-by: torlando-tech <torlando-tech@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant