The previous recheck narrowed the race without closing it. Testing
hotspotHold outside lifecycleLock left this interleaving:
1. the recheck passes, hold not yet set
2. holdForHotspot() sets the flag and calls stop(), which takes the
lock, finds nothing published, and returns having stopped nothing
3. this block takes the lock and publishes the service anyway
Aware ends up running while the hotspot believes it owns the radio,
which is the original failure.
The hold is now tested inside the same synchronized block that
publishes, so the check and the publication are one step against
stop(). Ordering holds because holdForHotspot() sets the flag before
calling stop(): either the publisher observes the flag and abandons,
or stop() observes the published service and tears it down. There is
no interleaving that leaves a service published with the hold set.
stopServices() stays outside the lock since it can block.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Addresses three issues from Codex review on #808.
The Wi-Fi Aware hold could be defeated by a race. startIfPossible()
does a long stretch of async work between checking the hold and
assigning the service, so holdForHotspot() landing in that window
left an in-flight start free to resurrect NAN behind the hotspot's
back, putting every P2P attempt back on BUSY. The hold is now
rechecked before committing the service, and the freshly started
service is torn down if the hotspot claimed the radio meanwhile.
Startup error paths bypassed cleanup. A web-server failure, a null
connection info, or a throw from the outer block stopped the manager
but left the Aware hold set, blocking all mesh starts until the user
happened to retry or close the screen. Worse, an error after the
server had started left it serving the APK on port 9999 -- including
after the device reconnected to an ordinary Wi-Fi network -- because
only stopHotspot() cleared it.
Both follow from the same gap: cleanup lived at the call sites rather
than in one place. All failures now go through failWith(), which
shares teardown() with stopHotspot() and releases the server, the
manager and the hold together.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Starting the APK-sharing hotspot failed intermittently with
"Failed to create hotspot: BUSY", sometimes for minutes, then
succeeded for no apparent reason. Two distinct causes, both
confirmed against a Pixel 9a via dumpsys and HAL logs.
1. Wi-Fi Aware holds the radio. The mesh's NAN interface and
Wi-Fi Direct's P2P interface cannot coexist on common chipsets:
HalDevMgr: bestIfaceCreationProposal is null, requestIface=P2P,
existingIface=[name=wlan0 type=STA, name=aware_nmi0 type=NAN]
WifiP2pNative: Failed to create P2p iface
The P2P state machine then stays in P2pDisabledState and answers
every createGroup with BUSY, while still broadcasting
WIFI_P2P_STATE_ENABLED. Whether sharing worked came down to
whether Aware happened to be attached, which is what made it look
random. WifiAwareController now releases Aware for the duration of
the hotspot and blocks restarts until it finishes.
2. Orphaned groups. A P2P group outlives the process that created
it, so a crash or swipe-away while hosting leaves one behind, and
the framework answers BUSY for as long as it exists. Startup now
removes a stale group first, but only one it can show is ours --
Wi-Fi Direct is shared with Cast, Android Auto and Quick Share.
Ownership is the group name we recorded creating, with the SSID
prefix as a fallback for orphans from older builds.
Also fixed while tracing these:
- Channel leak: initialize() ran on every retry attempt and the
channel was never closed, leaving a binder registration with
WifiP2pService per attempt. Observed climbing to 7 stale clients.
It is now initialised once and closed after removeGroup replies.
- BUSY is the framework's catch-all reply, so retrying was futile
for permanent causes and too impatient for real contention.
Retries now back off 1s/2s/4s/8s and only for genuinely transient
failures; P2P being off fails immediately with a message that says
so rather than 15 seconds ending in "busy".
- Turning Wi-Fi off mid-session left the UI showing an active
hotspot forever; it now aborts cleanly.
Retry and startup decisions are extracted into HotspotStartupPolicy,
which has no Android dependencies and is covered by 13 unit tests.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Receivers hard-cap reassembly at MAX_FRAGMENTS_PER_ID (256), but the
generic send path fragmented packets with no caller cap (0xFFFF), so
broadcast file transfers above ~120 KB were fully transmitted yet
undeliverable. FragmentingPacketSender now caps fragmentation at
MAX_FRAGMENTS_PER_ID and reports failure via a new
TransferProgressEvent.failed flag, which surfaces as
DeliveryStatus.Failed in the UI and as a file_send error in the debug
test hook instead of an indefinite wait.
Adds FragmentingPacketSenderTest and a file_oversize mesh-lab scenario
asserting sender-side rejection.
Debug-only broadcast receiver (app/src/debug) exposes mesh operations over
ADB: scan, connect, Noise handshake, DMs, broadcast, announce, file
send/receive with SHA-256 verification, BLE toggle, state dumps, and raw
packet injection. tools/release_gate/mesh_lab.py orchestrates scenarios
(dm, broadcast, file, file_private, raw) on two live devices and emits
evidence JSON.
Align the background location footer with the battery optimization layout so primary, check-again, and skip stay at the same height when navigating between them.
Co-authored-by: Cursor <cursoragent@cursor.com>
Addresses review feedback: the size-cap error was posted to the main
mesh timeline, so a user sending from a private chat or channel never
saw it. rejectIfOversized now posts to the private conversation
(addPrivateMessageNoUnread) or channel (addChannelMessage) the send
originated from, falling back to the main timeline for public sends.
Adds regression tests for both routings.