Deluan Quintão f853ca604a
refactor(db): migrate all ids to a uniform canonical 128-bit base62 encoding (#5824)
* refactor(model): extract canonical 128-bit base62 id codec

* feat(model): generate random ids as canonical 128-bit base62 values

* feat(scanner): emit legacy PIDs in canonical base62 encoding

* feat(db): add id canonicalization transform for the uniform-ids migration

* feat(db): migrate all ids to canonical 128-bit base62 encoding

* fix(db): canonicalize ids in junction tables and JSON columns

* chore(jellyfin): update id-family notes for uniform canonical ids

* test(ids): harden codec input contract and migration edge coverage

* refactor(model): use log.Fatal for Encode128 contract guard per project convention

* fix(db): force full rescan after id migration for legacy PID configs

* test(db): guard id-column inventory against schema drift

* refactor(ids): compile-time Encode128 contract and unified column rewrite helper

* refactor(db): apply review feedback to id migration

Filter empty strings in collectColumn's SQL, reuse a prepared statement
for rewriteColumn updates, and clarify the legacy ID functions' comment
now that they emit the canonical encoding.

* feat(auth): split session and public-link JWT secrets, rotating sessions on id migration

* test(subsonic): initialize public token secret in helpers suite

The suite sets auth.TokenAuth directly instead of calling auth.Init, so the
new PublicTokenAuth was nil whenever Ginkgo's spec order ran a helpers spec
before any spec that calls auth.Init, panicking in publicurl.ImageURL.

* refactor(db): inline canonicalID into its only consumer, the uniform-ids migration

* refactor(model): rename Encode128/Decode128 to Encode/Decode

With every id now exactly 128 bits, the width suffix is redundant; the
package-qualified id.Encode/id.Decode carries the same information.

* test(db): make the id-columns guard classify JSON columns too

The guard only inspected columns named id/pid/*_id, so it could not see ids
embedded in JSON. Widen it to *_ids and to every JSON column, and drive the
"covered" set from a new embeddedIDColumns list instead of the inline calls
in the migration.

Every JSON column the schema has now carries a verdict. The four denormalized
caches -- media_file/album.participants, media_file/album.tags,
album.folder_ids and artist.similar_artists -- hold only artist, tag and
folder ids. Those all come from id.NewHash, whose 22-char base62 encoding of
a 128-bit MD5 is already in canonical range, so canonicalID is the identity
on them and the migration correctly leaves them alone. A new codec test pins
that invariant, since the exemptions depend on it.

Verified on a copy of a 727MB/96k-track production database: canonicalizing
those four columns changed zero rows, and artist, tag and folder ids were
themselves unchanged by the migration (only media_file ids moved, 95108 of
96666).
2026-08-02 12:58:53 -04:00

49 lines
1.0 KiB
Go

package id
import (
"crypto/md5"
"crypto/rand"
"fmt"
"math/big"
"strings"
)
func NewRandom() string {
var b [16]byte
_, _ = rand.Read(b[:])
return Encode(b)
}
// Encode renders a 16-byte value as the canonical 22-char zero-padded base62 id.
func Encode(b [16]byte) string {
return fmt.Sprintf("%022s", new(big.Int).SetBytes(b[:]).Text(62))
}
// Decode is the exact inverse of Encode.
func Decode(s string) ([]byte, error) {
if len(s) != 22 {
return nil, fmt.Errorf("invalid id length %d", len(s))
}
v, ok := new(big.Int).SetString(s, 62)
if !ok || v.Sign() < 0 {
return nil, fmt.Errorf("invalid base62 id %q", s)
}
if v.BitLen() > 128 {
return nil, fmt.Errorf("id %q overflows 128 bits", s)
}
return v.FillBytes(make([]byte, 16)), nil
}
func NewHash(data ...string) string {
hash := md5.New()
for _, d := range data {
hash.Write([]byte(d))
hash.Write([]byte(string('\u200b')))
}
return Encode([16]byte(hash.Sum(nil)))
}
func NewTagID(name, value string) string {
return NewHash(strings.ToLower(name), strings.ToLower(value))
}