fetch-georelays.yml rewrites the bundled asset every week; a test
asserting the asset's exact entry count would fail on the first data
update after this merges. The exact-count check moves to a committed
snapshot of online_relays_gps.csv, which never changes: a different
count there is a change in the validator, never a change in the data.
The bundled asset is still tested on every run: it must pass the same
validation a download must pass, including the minimum entry count. A
weekly update the validator refuses, or one that shrinks below that
minimum, fails the test suite. A bad file gets caught in the repo
instead of shipping inside the app.
The snapshot is a copy of online_relays_gps.csv as of Aug 30.
#914 aligned the file both platforms read and how rows are keyed and
ordered; the acceptance rules still differed. iOS rejects a whole
directory on one malformed or conflicting row, validates the header,
caps size and rows, and screens every host. This client skipped bad
rows and accepted almost any host. The bitchat repo validates the file
before it lands; this port bounds what the client accepts as well.
parseCsv and canonicalHost are replaced by a port of
GeoRelayDirectory.validatedEntries and validatedDirectoryAddress. A
rejected download keeps the current list; one that keeps less than
half of the known entries is rejected. Five details follow iOS's
compiled validator rather than a plain reading of the Swift, each with
a test. The cache refetches at startup when its source URL changed.
Cross-checked against the Swift validator on 84 fixture runs and
100,000 fuzzed inputs. Full suite, lint and build pass.
fetch-georelays.yml downloads the relay list from the georelays repo
every week and pushes it straight to main. This branch moves the app's
fetch to online_relays_gps.csv, and a job still reading the old source
would keep the bundled asset on a different list from the one the app
downloads.
Pointed at online_relays_gps.csv, the job keeps the bundled asset on
the app's source. When the file has not changed, the job downloads the
same file and pushes nothing; when it changes, the asset follows. The
asset itself is not changed in this branch.
Review follow-up on the naming: the bundled asset and the download
cache kept their old file names when the fetch moved to
online_relays_gps.csv. A comment now records where both come from; the
names then do not send anyone looking for a georelays fetch that is
gone.
Review follow-up from the automated pass. The selection tests carried a
real relay hostname and real city coordinates; AGENTS.md requires
synthetic, non-identifying fixtures. Both are replaced with clearly
synthetic values that keep the same shapes: the port-variant pair, the
one-coordinate tie group, and the duplicated nearest relay that used to
displace the fifth server.
The selection rules are also now recorded in
docs/client-rewrite-contracts.md, since agreeing clients are the point
of the change: same source file, same host key, same tie order.
Behavior is unchanged; RelayDirectoryTest covers it.
Both platforms use only the five relays nearest a geohash, and a
message crosses platforms only if the two selections share a relay.
Android and iOS read different relay files; measured Aug 11 over 6,013
points against iOS's own selection code, the two selections shared 2.3
of 5 relays on average, and about one in nine points shared none.
Fetch the file iOS reads. Key each row by the host string iOS builds,
which collapses the 115 hosts listed twice. Order distance ties by that
key. With all three, both platforms select the same five, in order, at
all 6,013 points. The bundled asset is left to the weekly job.
If SQLite clearAll fails mid-panic, in-memory state was already cleared
but a process restart could reload encrypted history from disk (#699).
Fall back to deleting the database files before reporting failure.
A relay that fails is retried on an exponential backoff, then abandoned
for the lifetime of the process. Nothing brings it back: the relay layer
registers no connectivity callback, the periodic subscription validator
only repairs subscriptions on sockets that are already open (and returns
immediately when connectedRelayCount is 0, which is exactly the state
after an outage), and connect() runs once from NostrClient.initialize().
The remaining paths that reset reconnectAttempts are a manual retry, a
Tor state change, and a successful open.
Two ways a relay died permanently:
- Any error whose message mentioned DNS returned before scheduling
anything at all. "Unable to resolve host" is what this device reports
when it simply has no network, so a moment in a tunnel killed every
relay at once, with no retry ever.
- Otherwise the schedule stopped at MAX_RECONNECT_ATTEMPTS. With
INITIAL=1s and MULTIPLIER=2 that is nine waits totalling about eight
and a half minutes, after which the relay was dead. MAX_BACKOFF_INTERVAL
was unreachable: attempt 9 asks for 256s and attempt 10 gave up, so the
five-minute ceiling the constant defines never applied to anything.
Let the backoff saturate at MAX_BACKOFF_INTERVAL and keep retrying there.
A name-resolution failure now backs off like any other error. Steady state
costs one connection attempt per relay per five minutes; the previous
behaviour cost the user every internet DM, delivery receipt and geohash
channel until they noticed and restarted the app.
The schedule moves into RelayReconnectPolicy so it is unit-testable
without OkHttp or a Context.