If SQLite clearAll fails mid-panic, in-memory state was already cleared
but a process restart could reload encrypted history from disk (#699).
Fall back to deleting the database files before reporting failure.
A relay that fails is retried on an exponential backoff, then abandoned
for the lifetime of the process. Nothing brings it back: the relay layer
registers no connectivity callback, the periodic subscription validator
only repairs subscriptions on sockets that are already open (and returns
immediately when connectedRelayCount is 0, which is exactly the state
after an outage), and connect() runs once from NostrClient.initialize().
The remaining paths that reset reconnectAttempts are a manual retry, a
Tor state change, and a successful open.
Two ways a relay died permanently:
- Any error whose message mentioned DNS returned before scheduling
anything at all. "Unable to resolve host" is what this device reports
when it simply has no network, so a moment in a tunnel killed every
relay at once, with no retry ever.
- Otherwise the schedule stopped at MAX_RECONNECT_ATTEMPTS. With
INITIAL=1s and MULTIPLIER=2 that is nine waits totalling about eight
and a half minutes, after which the relay was dead. MAX_BACKOFF_INTERVAL
was unreachable: attempt 9 asks for 256s and attempt 10 gave up, so the
five-minute ceiling the constant defines never applied to anything.
Let the backoff saturate at MAX_BACKOFF_INTERVAL and keep retrying there.
A name-resolution failure now backs off like any other error. Steady state
costs one connection attempt per relay per five minutes; the previous
behaviour cost the user every internet DM, delivery receipt and geohash
channel until they noticed and restarted the app.
The schedule moves into RelayReconnectPolicy so it is unit-testable
without OkHttp or a Context.
The reconnect case fails on main: a peer served once, then sent mail while
away, receives nothing when it returns.
The other three hold the properties the fix must not break — a reconnect
with no new mail still sends nothing, the trim leaves the cap in place
rather than emptying it, and a recently delivered message is not the one
forgotten.
sendCachedMessages is latched per peer by cachedMessagesSentToPeer, and
nothing ever removed a peer from that set. So a peer received whatever was
held for it on its first connection of the session, and every later
reconnection was refused at the top of the function. Mail queued while it
was away sat in the cache until it aged out.
The latch means "this peer has been handed everything we hold". Caching
new mail for that peer is exactly the event that stops being true, so
release it there. A reconnect with nothing new still sends nothing, which
is what the latch is for.
Also stops the periodic cleanup wiping both tracking sets outright.
Emptying them bounded memory by forgetting: deliveredMessages is what
keeps an already-delivered message from being queued again, and the peer
latch is what keeps a batch from being sent twice. Both are now
insertion-ordered and trimmed to their cap oldest-first.