SoulSync

Commit Graph

Author	SHA1	Message	Date
Broque Thomas	5bc5fbb662	Add MusicBrainz as a metadata source Register MusicBrainz as a first-class metadata source alongside Deezer, iTunes, Spotify, Discogs, and Hydrabase. Expose the shared client through metadata services, add the settings option, and expand the MusicBrainz search adapter with source-compatible artist, album, track, and detail methods. Carry MusicBrainz IDs through similar-artist discovery, recommended artists, artist map serialization, and personalized playlist selection. Update DB migrations and lookup filters so similar_artist_musicbrainz_id is preserved on older schemas and used for source requirements and library exclusion. Normalize MusicBrainz album adapter output for import context and add regression coverage for registry mapping, typed album conversion, and similar-artist filtering. Verified by user with 120 focused tests passing.	1 week ago
Broque Thomas	e061f12a05	Filter owned artists from discovery recommendations	1 week ago
Broque Thomas	6fe85f2f37	Server playlist sync: append mode (preserve user-added tracks) Discord report (CJFC, 2026-04-26): syncing a Spotify playlist to the server overwrote anything manually added to the server-side playlist. The fix adds a per-sync mode picker next to the Sync button on the playlist details modal — Replace (default, current delete-recreate behavior) or Append only (preserves existing tracks, only adds new ones). Useful when the source platform caps playlist size and the user is manually building beyond it on the server. Implementation: * New `append_to_playlist(name, tracks)` method on Plex / Jellyfin / Navidrome clients. Each uses the server's NATIVE append API: - Plex: `existing_playlist.addItems(new_tracks)` - Jellyfin: `POST /Playlists/<id>/Items?Ids=...&UserId=...` - Navidrome: Subsonic `updatePlaylist?songIdToAdd=...` Falls back to `create_playlist` when the playlist doesn't exist yet (first sync). No delete-recreate, no backup playlist created (preserves playlist creation date + metadata + non-soulsync-managed tracks). * Dedup-by-server-native-id (ratingKey for Plex, GUID for Jellyfin, song-id for Navidrome) — never re-adds a track already on the playlist. Server-native identity, not fuzzy title+artist match, so it can't false-collide. * `sync_service.sync_playlist` accepts `sync_mode='replace'\|'append'` kwarg. Single if/else branch dispatches to `append_to_playlist` or `update_playlist`. Threaded through `core/discovery/sync.run_sync_task` and the `/api/sync/start` HTTP handler. Validation on the API rejects unknown mode strings (defaults to 'replace'). * Frontend: per-playlist `<select id="sync-mode-${id}">` rendered next to the Sync button in both modal renderers (sync-spotify.js for Spotify playlists, sync-services.js for Deezer ARL playlists). `startPlaylistSync` reads the select at click time; missing select (other callers like discover.js) defaults to 'replace' so backward compat preserved without per-call-site updates. * SoulSync standalone has no playlist methods at all and the modal hides the Sync button entirely on it via `_isSoulsyncStandalone` — dispatch never reaches that path, no defensive fallback needed. 15 new tests pin per-server append behavior: - missing playlist → create_playlist delegation - dedup filtering (existing IDs skipped, only new tracks added) - empty new-track set short-circuits without API call - failure paths return False without raising - contract listing (KNOWN_PER_SERVER_METHODS includes 'append_to_playlist'; Plex / Jellyfin / Navidrome all implement) Plus tests/discovery/test_discovery_sync.py fake `sync_playlist` fixture got `sync_mode='replace'` default to match the new signature (was breaking after the kwarg add; now passing). WHATS_NEW entry under new '2.6.0' block (hidden by `_getLatestWhatsNewVersion` until next release bump). Closes CJFC discord request.	2 weeks ago
Broque Thomas	a6bb5f5b43	MS Cin-5: Drop per-server globals — engine owns the clients Per-server web_server.py globals (plex_client / jellyfin_client / navidrome_client / soulsync_library_client) are gone. The engine now owns the per-server client instances; web_server.py constructs them inline into the engine init and routes everything through media_server_engine.client('<name>'). Multi-client consumers refactored to take the engine instead of separate per-server kwargs: - services/sync_service.py: PlaylistSyncService.__init__ now takes media_server_engine. Internal _get_active_media_client resolves the active server's client through self._engine.client(name) instead of the per-server self.X_client attributes. - core/listening_stats_worker.py: ListeningStatsWorker takes media_server_engine. The plex/jellyfin/navidrome dispatch in _poll collapses to engine.client(active_server) (gated to those three servers — SoulSync standalone has no listening data). - core/web_scan_manager.py: WebScanManager takes media_server_engine instead of the hand-keyed media_clients dict that drifted out of sync with the engine. - core/discovery/sync.py: SyncDeps holds media_server_engine instead of plex_client / jellyfin_client. Playlist-image dispatch routes through engine.client(name). Web_server.py: - Per-server globals removed from the chained `= None` init line + their try/except construction blocks. Replaced with a _safe_init_media_client(factory, name) helper that captures per-server init failures + passes the resulting clients straight into the MediaServerEngine init dict. - All construction sites (PlaylistSyncService, WebScanManager, ListeningStatsWorker, SyncDeps, library_check) updated to receive the engine instead of per-server clients. Test fixtures (tests/discovery/test_discovery_sync.py) gain a _FakeMediaServerEngine stub + the SyncDeps build helper passes that instead of separate plex/jellyfin clients.	3 weeks ago
Broque Thomas	77c54ab7a7	Migrate discography + quality scanner to typed Album path Three more album-shape consumers now route through Album.from_<source>_dict() when caller passes a known source: - _build_discography_release_dict (artist discography cards) - _build_artist_detail_release_card (artist detail release cards) - _normalize_track_album (quality scanner result normalization) Legacy duck-typing stays as fallback for unknown source, non-dict input, or converter errors. Pure additive — existing callers without source kwarg unchanged.	3 weeks ago
Antti Kettunen	2bc8e8a27b	Preserve artwork in quality scanner wishlist handoff - carry track-level album art through the quality scanner normalization path - preserve artist artwork when provider results expose it - keep album.image_url and album.images populated so the wishlist UI can render the cover consistently - add a regression test covering provider payloads with image_url on both the track and artist	4 weeks ago
Antti Kettunen	c97a072f54	Refactor quality scanner to respect primary metadata provider - search metadata providers in source-priority order for each generated query instead of caching one client for the whole scan - keep the quality-scanner worker provider-neutral and preserve the no-provider error path - update the quality-scanner tests and remove the obsolete web_server spotify_client injection	4 weeks ago
Broque Thomas	793593de51	Lift _run_tidal_discovery_worker to core/discovery/tidal.py Missed worker from the PR5 discovery-workers series — Tidal sits in the same domain as the deezer / spotify_public / listenbrainz / youtube / beatport workers that were lifted in PR5b–PR5h, follows the same shape, shares the same `_search_spotify_for_tidal_track` helper, and was simply overlooked in the original inventory. Pure 1:1 lift of the 212-line worker. Wrapper keeps the original entry-point name so the existing call sites in web_server.py continue to work without changes. What `run_tidal_discovery_worker` does: 1. Pause enrichment workers (release shared resources). 2. For each Tidal track: - Cancellation gate (state['cancelled']). - Discovery cache lookup; cache hit short-circuits the search. - SimpleNamespace-style track passed straight to `_search_spotify_for_tidal_track` (the shared helper used by every worker in this family). - On Spotify match: build `match_data` preserving track_number / disc_number from raw API data, image extracted from album images or track object fallback, release_date filled from track.release_date when album dict is missing it. - On iTunes match: dict result populated as `match_data` with source set to discovery_source, image extracted from album images. - Save matched result to discovery cache. - On miss: Wing It stub stored as 'wing-it' status (success ticked). 3. After all tracks: phase='discovered', activity feed entry, sync discovery results back to mirrored playlist via `_sync_discovery_results_to_mirrored` with 'tidal' tag. 4. On error: state['phase']='error' + status with error string. 5. Finally: resume enrichment workers. Dependencies injected via `TidalDiscoveryDeps` (13 fields) — tidal_discovery_states, spotify_client, plus 11 callable helpers (pause/resume enrichment, get_active_discovery_source, get_metadata_fallback_client, get_discovery_cache_key, get_database, validate_discovery_cache_artist, search_spotify_for_tidal_track, build_discovery_wing_it_stub, add_activity_item, sync_discovery_results_to_mirrored). Same surface as the deezer worker. Diff vs original after `deps.X` → global X normalization is zero differences — 212 lines orig = 212 lines lifted, byte-identical body (including all whitespace, comments, log strings). Tests: 9 new under tests/discovery/test_discovery_tidal.py covering cache hit short-circuit, Spotify tuple match (track/disc preservation), iTunes dict match path, Wing It fallback, cancellation, completion phase update, activity feed entry, mirrored sync invocation, per-track error handling. Full suite: 1299 passing (was 1290). Ruff clean.	4 weeks ago
Broque Thomas	a38bfcba55	PR5h: lift _run_quality_scanner to core/discovery/quality_scanner.py Final lift in the PR5 discovery-workers series. Pulls the 328-line library quality scanner out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrapper keeps the original entry-point name. What the quality scanner does: 1. Reset scanner state (counters, results), load quality profile + minimum acceptable tier from QUALITY_TIERS. 2. Load tracks from DB based on scope: - 'watchlist' → tracks for watchlisted artists only. - other → all library tracks. 3. For each track: - Stop-request gate (state['status'] != 'running'). - Quality-tier check via _get_quality_tier_from_extension(file_path). - Skip tracks meeting standards (tier_num <= min_acceptable_tier). - For low-quality tracks: matching_engine search query gen, score candidates against Spotify (artist + title similarity, album-type bonus), pick best match >= 0.7 confidence. - On match: add full Spotify track to wishlist via `wishlist_service.add_spotify_track_to_wishlist` with source_type='quality_scanner' and a source_context that captures original file_path, format tier, bitrate, and match confidence. 4. After all tracks: status='finished', progress=100, activity feed entry, emit `quality_scan_completed` event for automation engine. 5. On critical exception: status='error', error message captured. Wishlist service interaction is via the public `add_spotify_track_to_wishlist` API only — no overlap with kettui's planned `core/wishlist/` package extraction (the import lives inside the function, exactly as in the original, and will follow whatever path that package takes). Dependencies injected via `QualityScannerDeps` (8 fields) — quality_scanner_state dict, quality_scanner_lock, QUALITY_TIERS constant, spotify_client, matching_engine, automation_engine, plus 2 callable helpers (get_quality_tier_from_extension, add_activity_item). Diff vs original after `deps.X` → global X normalization is zero differences — 328 lines orig = 328 lines lifted, byte-identical body (including all whitespace, comments, log strings, and the inline `from core.wishlist_service import get_wishlist_service` / `from database.music_database import MusicDatabase` imports at the top of the function). Tests: 11 new under tests/discovery/test_discovery_quality_scanner.py covering state init/reset, no-watchlist-artists short-circuit, unauthenticated Spotify error, high-quality skip, low-quality search trigger, match → wishlist add (with full source_context payload), no-match no-add, mid-loop stop request, completion phase + progress, automation engine event emission, all-library scope load. Full suite: 1152 passing (was 1141). Ruff clean. End of the PR5 series — `web_server.py` lost ~328 lines on this commit alone; total trim across PR5a–PR5h is ~2,400 lines of discovery worker code moved into focused `core/discovery/*.py` modules. The remaining discovery-adjacent worker `_process_watchlist_scan_automatically` was deliberately deferred to avoid overlap with kettui's planned wishlist extraction.	4 weeks ago
Broque Thomas	c9108ef2fe	PR5g: lift _run_listenbrainz_discovery_worker to core/discovery/listenbrainz.py Seventh lift in the PR5 discovery-workers series. Pulls the 286-line ListenBrainz discovery worker out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrapper keeps the original entry-point name. What the ListenBrainz discovery worker does: 1. Pause enrichment workers (release shared resources). 2. For each ListenBrainz track: - Cancellation gate (state['phase'] != 'discovering'). - Discovery cache lookup; cache hit short-circuits the search. - Strategy 1: matching_engine search queries with confidence scoring against Spotify (preferred) or iTunes (fallback). - Strategy 2: swapped artist/title query. - Strategy 3: album-based query (uses album_name when available — unique to LB, since YouTube tracks don't have album metadata). - Strategy 4: extended search with limit=50. - On match → save to discovery cache with image extracted from album images or matched_track.image_url fallback. - On miss → Wing It stub stored as 'wing-it' status. 3. After all tracks: phase='discovered', status='complete', activity feed entry mentioning 'ListenBrainz Discovery Complete'. 4. On error: state['status']='error', phase='fresh'. 5. Finally: resume enrichment workers. Dependencies injected via `ListenbrainzDiscoveryDeps` (16 fields) — listenbrainz_playlist_states, spotify_client, matching_engine, plus 13 callable helpers (pause/resume enrichment, get_active_discovery_source, get_metadata_fallback_client, get_discovery_cache_key, get_database, validate_discovery_cache_artist, extract_artist_name, spotify_rate_limited, discovery_score_candidates, get_metadata_cache, build_discovery_wing_it_stub, add_activity_item). Diff vs original after `deps.X` → global X normalization is zero differences — 286 lines orig = 286 lines lifted, byte-identical body (including all whitespace, comments, log strings). Pre-existing bug preserved (not fixed): if `listenbrainz_playlist_states[ state_key]` raises KeyError on entry, the outer except handler tries to mutate `state` which is unbound → secondary UnboundLocalError. Same bug in the original (and the YouTube discovery worker). Documented here for future cleanup but out of scope for the lift. Tests: 11 new under tests/discovery/test_discovery_listenbrainz.py covering cache hit short-circuit, Strategy 1 confidence match, Wing It fallback, iTunes fallback (Spotify unauthenticated and rate-limited), cancellation (phase change), completion phase update, activity feed entry, per-track error handling, float duration_ms tolerance (regression for the :02d format crash fixed earlier), enrichment workers resume on finally. Full suite: 1141 passing (was 1130). Ruff clean.	4 weeks ago
Broque Thomas	04647eb9f7	PR5f: lift _run_beatport_discovery_worker to core/discovery/beatport.py Sixth lift in the PR5 discovery-workers series. Pulls the 323-line Beatport chart discovery worker out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrapper keeps the original entry-point name. What the Beatport discovery worker does: 1. Pause enrichment workers (release shared resources). 2. For each Beatport track: - Cancellation gate (state['phase'] != 'discovering'). - Clean Beatport text (artist/title) of common annotations via `clean_beatport_text` helper. - Single-string artist normalization for "CID,Taylr Renee"-style entries — split on comma, take the first. - Discovery cache lookup; cache hit short-circuits the search and normalizes cached artists from ['str'] → [{'name': 'str'}] to match the frontend's expected list-of-objects shape. - matching_engine search-query generation (with high min_confidence of 0.9 to avoid bad matches). - Strategy 1: scored candidates from initial Spotify/iTunes searches. - Strategy 4: extended search with limit=50 if no high-confidence match found. - On Spotify match: format artists as [{'name': str}] objects, pull full album object from raw cache when available, fallback to reconstructed album dict otherwise. - On iTunes match: format with image_url-derived album.images entry (300x300 spec), source set to discovery_source. - Save matched result to discovery cache when confidence >= 0.75 (note: lower than search threshold; discovery still benefits from these less-confident matches as user-visible suggestions). - On miss: Wing It stub stored as 'wing-it' status (success ticked). 3. After all tracks: phase='discovered', activity feed entry, sync discovery results back to mirrored playlist via `_sync_discovery_results_to_mirrored` with 'beatport' tag. 4. On error: state['phase']='fresh' + status='error'. 5. Finally: resume enrichment workers. Dependencies injected via `BeatportDiscoveryDeps` (17 fields) — beatport_chart_states, spotify_client, matching_engine, plus 14 callable helpers (pause/resume enrichment, get_active_discovery_source, get_metadata_fallback_client, clean_beatport_text, get_discovery_cache_key, get_database, validate_discovery_cache_artist, spotify_rate_limited, discovery_score_candidates, get_metadata_cache, build_discovery_wing_it_stub, add_activity_item, sync_discovery_results_to_mirrored). Diff vs original after `deps.X` → global X normalization is zero differences — 323 lines orig = 323 lines lifted, byte-identical body (including all whitespace, comments, log strings). Tests: 12 new under tests/discovery/test_discovery_beatport.py covering cache hit short-circuit (with cached-artist normalization), Spotify match formatting (list and string artist inputs), iTunes match (image_url to album.images), Wing It fallback, cancellation (phase change), completion phase update, activity feed entry, mirrored sync invocation, top-level error handler, per-track error handling, comma-separated artist split. Full suite: 1130 passing (was 1118). Ruff clean.	4 weeks ago
Broque Thomas	c5e06691e3	PR5e: lift _run_spotify_public_discovery_worker to core/discovery/spotify_public.py Fifth lift in the PR5 discovery-workers series. Pulls the 278-line public-Spotify-link discovery worker out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrapper keeps the original entry-point name. What the Spotify Public discovery worker does: 1. Pause enrichment workers (release shared resources). 2. For each track: - Cancellation gate (state['cancelled']). - Normalize artists to plain string list (handles dict + str inputs). - Discovery cache lookup; cache hit short-circuits the search and populates display fields from the cached match. - SimpleNamespace duck-type → `_search_spotify_for_tidal_track` (shared search helper, returns tuple for Spotify or dict for iTunes). - On Spotify match: build `match_data` preserving track_number / disc_number from raw API data; image extracted from album images or track object fallback; release_date filled from track.release_date when album dict is missing it. - On iTunes match: dict result → match_data with source set to discovery_source; image extracted from album images. - Save matched result to discovery cache. - On miss: Wing It stub stored as 'wing-it' status. 3. After all tracks: phase='discovered', activity feed entry. 4. On error: state['phase']='error' + status with error string. 5. Finally: resume enrichment workers. This worker is structurally close to the Deezer worker (see PR5d) but intentionally diverges on: - Track-data field names (`spotify_public_track` vs `deezer_track`). - Artist normalization (Spotify Public can pass dicts or strings). - No mirrored-playlist DB writeback (sync is handled separately). Dependencies injected via `SpotifyPublicDiscoveryDeps` (12 fields) — spotify_public_discovery_states, spotify_client, plus 10 callable helpers (pause/resume enrichment, get_active_discovery_source, get_metadata_fallback_client, get_discovery_cache_key, get_database, validate_discovery_cache_artist, search_spotify_for_tidal_track, build_discovery_wing_it_stub, add_activity_item). Diff vs original after `deps.X` → global X normalization is zero differences — 278 lines orig = 278 lines lifted, byte-identical body (including all whitespace, comments, log strings). Tests: 10 new under tests/discovery/test_discovery_spotify_public.py covering cache hit short-circuit, dict-artist normalization, Spotify tuple match (track/disc preservation), iTunes dict match path, Wing It fallback, cancellation, completion phase update, activity feed entry, top-level error handler, per-track error handling. Full suite: 1118 passing (was 1108). Ruff clean.	4 weeks ago
Broque Thomas	2bc665e487	PR5d: lift _run_deezer_discovery_worker to core/discovery/deezer.py Fourth lift in the PR5 discovery-workers series. Pulls the 270-line Deezer discovery worker out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrapper keeps the original entry-point name so the existing call sites continue to work without changes. What the Deezer discovery worker does: 1. Pause enrichment workers (release shared resources). 2. For each Deezer track: - Cancellation gate (state['cancelled']). - Discovery cache lookup; cache hit short-circuits the search and populates display fields from the cached match (artist string, album name). - SimpleNamespace duck-type → `_search_spotify_for_tidal_track` (shared search helper, returns tuple for Spotify or dict for iTunes). - On Spotify match: build `match_data` preserving track_number / disc_number from raw API data, image extracted from album images or track object fallback, release_date filled from track.release_date when album dict is missing it. - On iTunes match: dict result populated as `match_data`, source set to discovery_source, image extracted from album images. - Save matched result to discovery cache. - On miss: Wing It stub stored as 'wing-it' status. 3. After all tracks: phase='discovered', activity feed entry, sync discovery results back to mirrored playlist via `_sync_discovery_results_to_mirrored`. 4. On error: state['phase']='error' + status with error string. 5. Finally: resume enrichment workers. Dependencies injected via `DeezerDiscoveryDeps` (13 fields) — deezer_discovery_states dict, spotify_client, plus 11 callable helpers (pause/resume enrichment, get_active_discovery_source, get_metadata_fallback_client, get_discovery_cache_key, get_database, validate_discovery_cache_artist, search_spotify_for_tidal_track, build_discovery_wing_it_stub, add_activity_item, sync_discovery_results_to_mirrored). Diff vs original after `deps.X` → global X normalization is zero differences — 270 lines orig = 270 lines lifted, byte-identical body (including all whitespace, comments, log strings). Tests: 10 new under tests/discovery/test_discovery_deezer.py covering cache hit short-circuit, Spotify tuple match (track/disc number preservation), iTunes dict match path, Wing It fallback, cancellation, completion phase update, activity feed entry, mirrored sync invocation, top-level error handler, per-track error handling. Full suite: 1108 passing (was 1098). Ruff clean.	4 weeks ago
Broque Thomas	bda0500226	PR5c: lift _run_playlist_discovery_worker to core/discovery/playlist.py Third lift in the PR5 discovery-workers series. Pulls the 323-line mirrored-playlist discovery worker out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrapper keeps the original entry-point name so the existing call site (`_run_playlist_discovery_worker(pls, automation_id=None)` from the automation engine) continues to work without changes. What the playlist discovery worker does: 1. Pause enrichment workers (release shared resources). 2. Pre-compute total track count across all playlists for the automation progress card. 3. For each playlist: - Fast pre-scan separates already-discovered tracks (skipped, unless incomplete metadata or a Wing It stub) from undiscovered ones. - For each undiscovered track: - Cancellation gate via _playlist_discovery_cancelled set. - Discovery cache lookup (with artist validation). - matching_engine search-query generation, then Spotify (preferred) or iTunes (fallback) search + scoring. - Extended search fallback (limit=50) if no high-confidence match. - On match → enrich album from metadata cache (id, images, total_tracks, album_type, release_date, artists, plus track_number and disc_number), build matched_data, write to track.extra_data, save to discovery cache. - On miss → Wing It stub stored as 'wing_it_fallback' provider. 4. After all playlists: emit `discovery_completed` event when at least one new track was discovered, mark automation progress 'finished'. 5. On error → automation progress 'error', traceback printed. 6. Finally: resume enrichment workers. Dependencies injected via `PlaylistDiscoveryDeps` (16 fields) — spotify_client, matching_engine, automation_engine, the cancellation set, plus 12 callable helpers (pause/resume enrichment, get_active_discovery_source, get_metadata_fallback_client/source, update_automation_progress, get_database, get_discovery_cache_key, validate_discovery_cache_artist, discovery_score_candidates, get_metadata_cache, build_discovery_wing_it_stub). Diff vs original after `deps.X` → global X normalization is zero differences — 323 lines orig = 323 lines lifted, byte-identical body (including all whitespace, comments, log strings). Tests: 15 new under tests/discovery/test_discovery_playlist.py covering empty playlists, no-tracks playlist skip, complete-discovery skip, incomplete-discovery re-run, Wing It always re-run, unmatched_by_user respect, cache hit short-circuit, match above threshold (extra_data + cache save), match below threshold falls to Wing It, iTunes fallback, neither-provider error path, cancellation, discovery_completed event emit, no-event on zero-discovered, multi-playlist grand_total aggregation. Full suite: 1098 passing (was 1083). Ruff clean.	4 weeks ago
Broque Thomas	3c1f614b6e	fix: cast duration_ms to int before :02d format in discovery workers yt_dlp sometimes returns float `duration_ms` for YouTube tracks. The discovery workers format the duration with `f"{x // 60000}:{(x % 60000) // 1000:02d}"` — and `:02d` requires an int. When the duration is a float, the format string raises: Unknown format code 'd' for object of type 'float' Caught when running YouTube discovery on a real playlist (bbno$ tracks) — every track failed with status='Error'. Pre-existing bug, surfaced now because of yt_dlp returning float durations on this playlist. Fixed at all 8 sites by casting through `int()` before the `// 60000` and `% 60000` operations: - core/discovery/youtube.py: 2 sites in run_youtube_discovery_worker (cache hit + main result construction). - web_server.py L29238/L29372: 2 sites in _run_listenbrainz_discovery_worker. - web_server.py L40112/L40136/L40161/L40178: 4 sites in the YouTube retry/pre-discovered results assembly path. The `if duration_ms` / `if dur` guard already protects against None and 0, so `int(...)` is only called on truthy numeric values. Tests: 1 new regression test under tests/discovery/test_discovery_youtube.py (`test_float_duration_does_not_crash_format`) — passes a float duration_ms and asserts the worker completes without an error result. Ruff clean.	4 weeks ago
Broque Thomas	27fa96fe97	PR5b: lift _run_youtube_discovery_worker to core/discovery/youtube.py Second lift in the PR5 discovery-workers series. Pulls the 332-line YouTube discovery worker out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrappers keep the original entry-point name so the two callers (`youtube_discovery_executor.submit(_run_youtube_discovery_worker, ...)`) continue to work without changes. What the YouTube discovery worker does: 1. Pause enrichment workers (release shared resources). 2. For each YouTube playlist track: - Cancellation check (phase != 'discovering' aborts). - Discovery cache lookup; cache hit short-circuits the search. - Strategy 1: matching_engine search queries with confidence scoring against Spotify (preferred) or iTunes (fallback). - Strategy 2: swapped artist/title query. - Strategy 3: raw (untokenized) query. - Strategy 4: extended search with limit=50. - On match → save to discovery cache. - On miss → build Wing It stub from raw source data. 3. After loop: phase='discovered', sort results by index, and for mirrored playlists write extra_data back to the DB. 4. Activity feed entry with match summary. 5. On error → state['status']='error', phase='fresh'. 6. Finally: resume enrichment workers. Dependencies injected via `YoutubeDiscoveryDeps` (16 fields) — youtube_playlist_states, spotify_client, matching_engine, plus 13 callable helpers (pause/resume enrichment, get_active_discovery_source, get_metadata_fallback_client, discovery cache key/validate, extract artist name, spotify_rate_limited, discovery_score_candidates, get_metadata_cache, build_discovery_wing_it_stub, get_database, add_activity_item). Diff vs original after `deps.X` → global X normalization is zero differences — 332 lines orig = 332 lines lifted, byte-identical body (including all whitespace). Pre-existing bug preserved (not fixed): if `youtube_playlist_states[url_hash]` raises KeyError on entry, the outer except handler tries to mutate `state` which is unbound → secondary UnboundLocalError. Same bug in the original. Documented here for future cleanup but out of scope for the lift. Tests: 14 new under tests/discovery/test_discovery_youtube.py covering cache hit short-circuit, Strategy 1 confidence match, Wing It fallback, iTunes fallback path (Spotify unauthenticated and rate-limited), cancellation (phase changed), skip_discovery flag, completion phase update, activity feed entry, mirrored playlist DB writeback, non-mirrored no-writeback, enrichment workers pause/resume, error-during-loop resume, results sorted by index after retry. Full suite: 1082 passing (was 1068). Ruff clean.	4 weeks ago
Broque Thomas	bdb7a3139d	PR5a: lift _run_sync_task to core/discovery/sync.py First lift in the new PR5 discovery-workers series. Pulls the 448-line playlist sync background worker out of `web_server.py` into its own focused module under `core/discovery/`. Pure 1:1 lift — wrappers keep the original entry-point name so the four callers (`sync_executor.submit(_run_sync_task, ...)`) continue to work without changes. What the sync worker does: 1. Convert frontend JSON tracks → SpotifyTrack/SpotifyPlaylist objects. 2. Normalize artist/album shapes for downstream wishlist parity. 3. Wire a progress_callback that updates `sync_states` + automation card. 4. Patch sync_service for database-only fallback when no media server is connected. 5. `run_async(sync_service.sync_playlist(...))` and capture the result. 6. Update sync_states to 'finished', push playlist poster image to Plex / Jellyfin / Emby, record sync history (with re-sync vs new-sync branching), emit `playlist_synced` event for automation engine, and persist sync status with a tracks_hash for smart-skip on the next scheduled sync. 7. On exception → mark error in sync_states + automation; finally clear progress callback + drop `_original_tracks_map` from sync_service. Dependencies injected via `SyncDeps` (11 fields) — config_manager, sync_service, plex_client, jellyfin_client, automation_engine, run_async, record_sync_history_start, update_automation_progress, update_and_save_sync_status, sync_states dict, sync_lock. The only structural drift from a pure paste is the top-of-function variable binding: original used `global sync_states, sync_service`, lifted version rebinds them as locals from deps (`sync_states = deps.sync_states` etc.) since the names aren't module-level in the new file. Same behaviour otherwise — diff against the original after `deps.X` → global X normalization is zero differences. Tests: 18 new under tests/discovery/test_discovery_sync.py covering sync history recording (new + resync), setup error path (with and without automation_id), missing sync_service handling, sync_playlist exception handling, successful sync state transition, unmatched-tracks summary, playlist image upload (plex + jellyfin + zero-synced gate), automation engine emit, automation progress finished call, sync history DB persistence (completion + match_details), tracks_hash persistence, and finally-block cleanup (callback clear + map drop). Full suite: 1068 passing (was 1050). Ruff clean. Kicks off the PR5 series — 9 discovery workers totaling ~2,400 lines across `_run_sync_task`, `_run_*_discovery_worker` family, `_run_quality_scanner`, and `_process_watchlist_scan_automatically`. Wishlist-related extractions deliberately skipped to avoid overlap with kettui's planned `core/wishlist/` package.	4 weeks ago

17 Commits (048e4e85d5ceb6bbb9f6a6d2b2efbc259cc331ff)