2 Commits

Author SHA1 Message Date
Deluan Quintão
7b7721f002
feat(matcher): match similar/top songs by multiple artists (#5668)
* feat(matcher): add Song.Artists (agents.Artist) and field-wise song dedup

* refactor(matcher): make song equality an agents.Song.Equals method via hashstructure

Move the sameSong free function from core/matcher into an Equals method on
agents.Song, following the model.MediaFile/Album.Equals convention. Uses strict
hashstructure hashing (nil opts, no IgnoreZeroValue) to preserve the original
whole-value equality contract. Tests moved to core/agents.

* feat(matcher): match by multiple artists with overlap ranking and artist-ID fast-path

* refactor(matcher): rank artist overlap and specificity above the preferred-track flag

Identity signals (specificityLevel, artistOverlap) now outrank the taste
signal (preferredMatch) in betterThan. A starred/4-star track that is a worse
identity match no longer beats a more specific or higher-overlap track.
PreferStarred still breaks ties when specificity and overlap are equal.

* fix(matcher): score artist-MBID specificity against all credited artists, not just the last

sanitizedTrack.artistMBID (string) replaced with artistMBIDs (map[string]struct{}) so
bucketTracks collects all credited owned MBIDs per query instead of last-write-wins.
computeSpecificityLevel tests set membership, letting each of a collaboration's
MBID-bearing artists reach the proper specificity level (4/5) independently.

* feat(plugins): carry multiple artists (with IDs) through SongRef conversions

* feat(plugins): regenerate schemas and PDK wrappers for multi-artist SongRef

* refactor(matcher): tidy bucketTracks accumulator and artist resolution

Replace bucketTracks' two parallel per-track maps (overlapByQuery/mbidsByQuery)
with a named queryAccum struct (F2). Collapse resolveArtists' four hand-mutated
parallel maps into a pendingArtist slice with derived nameToQueries/mbidToQueries
maps (F1). Replace the own() method on resolvedArtists with a package-level
addToSet helper that drops the method/receiver indirection (F3).

* docs(matcher): trim comments that restate the code

* fix(plugins): use Vec::is_empty for slice fields in generated Rust PDK

* fix(matcher): treat a resolved artist ID as an identity match for specificity

* docs(matcher): reflect artist-ID identity in the specificity ladder
2026-06-26 14:46:09 -04:00
Deluan Quintão
6486a27634
refactor(matcher): index-space resolution + batched title lookups (#5635)
* refactor(matcher): resolve matches in song-index space

* test(matcher): pin per-index duration matching for duplicate title+artist songs

* refactor(matcher): drop unreachable specificity sentinel

* refactor(matcher): hoist PreferStarred read out of scoring loop

* docs(matcher): correct config field references in MatchSongs doc

* fix(matcher): log swallowed per-artist DB error in title matching

* fix(matcher): fail title matching when all artist lookups error

* refactor(matcher): simplify loaders and test helpers

* docs(matcher): move algorithm docs to package-level doc.go

* docs(matcher): focus examples on fuzzy matching behavior

* refactor(matcher): store index in dedup map and harden test helper

* fix(matcher): keep exact-phase matches when all title lookups fail

* perf(matcher): batch title-phase artist lookups into one query

matchByTitle issued one GetAll per distinct artist, run serially. On a large
library a batch of similar-songs spans dozens of artists, and profiling against
a 95k-track library showed the matcher was ~90% bound in that serial query loop
(a 100-song batch fired ~89 separate multi-join queries, taking ~6s).

Replace the loop with a single 'order_artist_name IN (...)' query, then group the
returned tracks by artist in memory and score each song against its bucket. This
cuts a 100-song batch from ~6s to ~0.4s (roughly 14x) with less memory.

Grouping keys on order_artist_name (the field the query filters on, matching how
the per-artist queries are keyed), falling back to the sanitized Artist when it
is unset. Because there is now a single query, matchByTitle is all-or-nothing
like the ID/MBID/ISRC loaders: the per-artist best-effort skip is gone, while
resolveMatches still preserves exact-phase matches when the title query fails.

* refactor(matcher): group batched title matches by order_artist_name

After batching the title-phase lookups into one query, the returned tracks must
be grouped back to their artist. Key on MediaFile.OrderArtistName — the exact
field the query filters on — so collaboration/"feat." tracks (whose display
Artist differs from the sort artist) bucket correctly, with a sanitized-Artist
fallback when it is unset.

OrderArtistName is deprecated in favor of Participants, but the bulk GetAll path
does not hydrate participant detail (the rich artist fields come only from the
per-record GetWithParticipants JOIN), so the participant order name is empty here
and the column is the only populated source.

Also adds a TODO in computeSpecificityLevel: its artist-MBID levels read the
deprecated, unpopulated MediaFile.MbzArtistID column, so they never fire today.
2026-06-21 03:45:56 -04:00