If SQLite clearAll fails mid-panic, in-memory state was already cleared
but a process restart could reload encrypted history from disk (#699).
Fall back to deleting the database files before reporting failure.
A relay that fails is retried on an exponential backoff, then abandoned
for the lifetime of the process. Nothing brings it back: the relay layer
registers no connectivity callback, the periodic subscription validator
only repairs subscriptions on sockets that are already open (and returns
immediately when connectedRelayCount is 0, which is exactly the state
after an outage), and connect() runs once from NostrClient.initialize().
The remaining paths that reset reconnectAttempts are a manual retry, a
Tor state change, and a successful open.
Two ways a relay died permanently:
- Any error whose message mentioned DNS returned before scheduling
anything at all. "Unable to resolve host" is what this device reports
when it simply has no network, so a moment in a tunnel killed every
relay at once, with no retry ever.
- Otherwise the schedule stopped at MAX_RECONNECT_ATTEMPTS. With
INITIAL=1s and MULTIPLIER=2 that is nine waits totalling about eight
and a half minutes, after which the relay was dead. MAX_BACKOFF_INTERVAL
was unreachable: attempt 9 asks for 256s and attempt 10 gave up, so the
five-minute ceiling the constant defines never applied to anything.
Let the backoff saturate at MAX_BACKOFF_INTERVAL and keep retrying there.
A name-resolution failure now backs off like any other error. Steady state
costs one connection attempt per relay per five minutes; the previous
behaviour cost the user every internet DM, delivery receipt and geohash
channel until they noticed and restarted the app.
The schedule moves into RelayReconnectPolicy so it is unit-testable
without OkHttp or a Context.
GeohashRepository synchronizes every accessor that touches its shared
maps - eighteen of them - but currentGeohash is a plain var with two
unsynchronized accessors, and it is written from the UI thread when the
user picks a channel while Nostr handler and timer threads read it.
Several of those reads happen inside the @Synchronized methods
(refreshGeohashPeople, updateParticipant, displayNameForNostrPubkey), and
holding the lock on the read side buys nothing when the write side never
takes it: there is no happens-before edge, so a stale value can persist.
Two ways that shows up: startGeoParticipantsTimer loops
`while (repo.getCurrentGeohash() != null)` and keeps refreshing a channel
the user has left, and displayNameForNostrPubkey derives the identity for
the wrong geohash when deciding whether a pubkey is us.
Synchronize the two accessors, matching the rest of the class.
NotificationManager keeps the same state @Volatile for this reason.
The collision case fails on main: two packets sharing a 64-byte prefix and
a timestamp, where the second was silently dropped.
The other two pass before and after on purpose. Replay of an identical
packet must still be caught, and the same packet arriving from two
different peers must still be tracked separately — strengthening the
identity must not quietly weaken either.