SoulSync

Commit Graph

Author	SHA1	Message	Date
Broque Thomas	42f3026eef	Reject broken downloads before tagging via universal integrity check Discord report (fresh.dumbledore [VRN]): slskd sometimes ships broken files (truncated transfers, corrupt FLAC, wrong file substituted on filename match). They flowed through post-processing and only surfaced later — Plex/Jellyfin scan failures, dead-air playback, duplicate detector tripping over the wrong length. By that point the file was already tagged, copied, mirrored to the media server, and recorded in provenance. New module `core/imports/file_integrity.py`: - `check_audio_integrity(path, expected_duration_ms=None) -> IntegrityResult` - Three tiered checks, cheapest to most expensive: 1. File size sanity (catches 0-byte stubs and stub transfers) 2. Mutagen parse (catches header damage, wrong-format-with-right-extension) 3. Duration agreement vs. metadata source's expected length, ±3s tolerance (5s for tracks over 10 minutes — long tracks naturally drift more) - Returns IntegrityResult with `ok`, human-readable `reason`, and per-check `checks` dict for debugging - Never raises; pathological inputs return ok=False with explanation Pipeline integration in `core/imports/pipeline.py:post_process_matched_download`: - Hooks between the existing file-stability wait and AcoustID verification - On failure: quarantine via existing `move_to_quarantine` helper, mark task failed with descriptive error, clear matched-context, fire `on_download_completed(success=False)` so the slot is released for retry - Mirrors the existing AcoustID-failure path so retry behavior stays consistent - Wrapped in try/except so an unexpected failure inside the check itself cannot block downloads — logs and continues This is intentionally tier 1: universal across formats, no external deps. A future tier could verify FLAC STREAMINFO MD5 by decoding audio (needs flac binary or libflac wrapper) — skipped for now since tier 1 catches the dominant Discord-reported cases (truncated, 0-byte, wrong file). Tests: - `tests/imports/test_file_integrity.py` — 14 cases covering all three check tiers, edge cases (zero/negative expected duration, long-track wider tolerance, caller tolerance override), and the mutagen-unavailable degradation path - `tests/imports/test_import_pipeline.py` — two existing tests use 5-byte fixture files that the new check would reject; they monkeypatch the integrity check since they're testing plumbing (notification + metadata_runtime forwarding), not integrity behavior WHATS_NEW entry under '2.4.2' dev cycle.	3 weeks ago
Broque Thomas	783c543c3e	Auto-import: live per-track progress + in-progress history row User reported (Mushy / generally) that dropping an album into the staging folder left the auto-import history blank for the entire processing window — sometimes 5+ minutes for a full album. Pre- existing UX gap, not caused by the recent context-builder refactor. Two root causes: 1. ``_record_result`` only fired AFTER ``_process_matches`` returned. For a 14-track album with ~30s/track post-processing, that meant ~7 minutes of zero rows in auto_import_history → nothing for ``/api/auto-import/results`` to return → empty UI. 2. ``_current_status`` only ever transitioned between 'idle' and 'scanning' — never 'processing'. ``get_status()`` had no per- track index/name fields, so the UI had no way to render "Processing track 3/14: Mine" even if it wanted to. Fix: - New ``_record_in_progress`` inserts a status='processing' row up-front (before the per-track loop starts) so the UI sees the import the moment it begins. Returns the row id. - New ``_finalize_result`` updates that same row with the final outcome (completed/failed) when processing finishes. One row per album, not per track — keeps the history list clean. - Both share ``_serialize_match_data`` (extracted from the original ``_record_result``) so the in-progress row carries the same match payload shape the existing review UI already understands. - ``_process_matches`` updates ``_current_track_index``, ``_current_track_total``, and ``_current_track_name`` BEFORE each per-track callback fires, so a polling UI sees consistent "processing N/M: <name>" snapshots. - ``_scan_cycle`` flips ``_current_status`` to 'processing' before the per-album loop, resets it + the per-track fields after. Defensive ``finally`` clears progress even if the inner code path raised. - ``get_status()`` exposes the new fields so the UI's existing /api/auto-import/status polling picks them up. - Frontend (stats-automations.js): renders the new ``current_status='processing'`` state with track index/total/name in the existing progress bar element. New 'processing' status class for styling parity with 'scanning'. 8 regression tests in tests/imports/test_auto_import_live_progress.py: - get_status surfaces the new fields with sane defaults - track_index advances 1, 2, 3 during a 3-track loop - track_total set BEFORE the first callback fires (no '1/0' flicker) - _record_in_progress writes status='processing' with no processed_at - _finalize_result updates the same row to completed + processed_at, no second insert - _finalize_result with failed status leaves processed_at NULL - _finalize_result with row_id=None is a safe no-op - Per-track fields cleared by _scan_cycle's finally block Full pytest 1643 passed; ruff clean.	3 weeks ago
Broque Thomas	29089b35b3	Honor configured Tidal redirect_uri, drop request-host fallback Reported case (Foxxify): Tidal returned error 1002 ("Invalid redirect URI") on every authentication attempt for users accessing SoulSync from a network IP. User had ``http://127.0.0.1:8889/tidal/callback`` registered in his Tidal Developer Portal — matching the SoulSync UI default and docs. Root cause: the /auth/tidal route at web_server.py:5594-5598 had a "fallback: dynamically set based on request host" branch that fired when ``tidal.redirect_uri`` config was empty AND the request didn't come from localhost. That fallback overrode the TidalClient constructor's safe default (``http://127.0.0.1:<port>/tidal/callback``) with a uri built from request.host like ``http://192.168.x.x:8889/tidal/callback``. Tidal compares strings exactly so this never matched the documented portal registration and the user got 1002 before the consent screen even rendered. The trap is the SoulSync settings UI displays the default URI as the placeholder + "Current Redirect URI" display — but the placeholder never gets saved to config unless the user explicitly clicks Save. Most users who follow the docs (register the displayed default with Tidal, then click Authenticate) hit the empty-config path and the broken fallback. Fix: drop the request-host fallback. Empty config falls back to the constructor default that matches the documented portal registration. The existing post-auth swap-step in the instructions page below handles the Docker / remote-access case as designed: 1. SoulSync sends 127.0.0.1:8889 in the authorize URL → matches portal → Tidal accepts. 2. User authorizes → Tidal redirects browser to 127.0.0.1:8889 (which fails locally — nothing on user's machine listens there). 3. Instructions tell user to swap 127.0.0.1 with the host they're accessing SoulSync from. 4. Swapped URL hits the container's exposed callback port → auth completes. 8 regression tests in tests/test_tidal_auth_redirect_uri.py: - Configured redirect_uri sent verbatim (localhost / custom port / explicit network IP) - Empty config falls back to constructor default — NOT request.host (the actual reported scenario, with explicit assertion message warning if the bug returns) - Empty config + localhost access uses the same default (sanity) Full pytest 1635 passed; ruff clean.	3 weeks ago
Broque Thomas	24c2d75c6d	Make extract_external_ids recognize all source-tagging conventions Smoke-testing the just-merged provenance PR against live logs revealed the new ID-match block was silently no-opping: no [ExtID Match] / [Provenance Match] log lines despite the code path being live. Tracing revealed two related gaps in extract_external_ids' source detection: 1. Underscore-prefixed key. Deezer / Discogs / Hydrabase clients tag normalized track dicts with ``_source`` (underscore prefix — convention used in 8+ places across core/). The extractor only looked for ``provider`` and ``source``, so Deezer-sourced tracks silently returned no IDs. 2. No provider field at all. Spotify and iTunes raw API responses carry ``id`` but no provider/source key of any kind. The extractor couldn't disambiguate the native ``id``, so Spotify-primary scans would have hit the same silent miss once the user switched primary sources. Two-part fix: - ``extract_external_ids`` now recognizes ``_source`` as another candidate provider field. - New optional ``source_hint`` parameter lets the caller supply the configured primary source as a fallback when the track dict has no provider field of its own. Track-side provider field still wins when present (defensive against a wrong hint). Watchlist scanner now passes ``get_primary_source()`` as the hint so both naming conventions (Deezer-style _source, Spotify-style no-tag) get handled uniformly. 6 new regression tests cover: - _source recognized for Deezer - _source recognized for Hydrabase (cross-provider mapping) - _source recognized for Discogs (no library column — verifies graceful no-crash) - source_hint disambiguates raw tracks for spotify/itunes/deezer - track-side provider takes precedence over hint - None hint defaults safely Full pytest 1630 passed; ruff clean. After this lands and the server restarts, watchlist scans should produce [ExtID Match] / [Provenance Match] log lines for tracks already on disk regardless of which metadata source the user has configured as primary.	3 weeks ago
Broque Thomas	34ba26f5c8	Persist source IDs at download time + backfill onto tracks on sync Followup to fix/watchlist-external-id-match. The companion PR closed the demand side — the watchlist scanner asks for tracks by external IDs before falling back to fuzzy. But for users on Plex / Jellyfin / Navidrome the supply side was still broken: tracks.spotify_track_id (and the other ID columns) only got populated by the asynchronous enrichment workers, sometimes hours after the file was actually written. During that window the ID match fell through to fuzzy and the bug returned. We were already collecting every ID during post-processing — they live in the `pp` dict in core/metadata/source.py:embed_source_ids and get embedded into file tags. We just dropped the in-memory copy afterwards. This PR persists them and uses them: - Schema migration adds spotify_track_id / itunes_track_id / deezer_track_id / tidal_track_id / qobuz_track_id / musicbrainz_recording_id / audiodb_id / soul_id / isrc columns + indexes to the existing track_downloads table (already keyed by file_path). - core/metadata/source.py:embed_source_ids exposes pp["id_tags"] and the resolved ISRC back to the import context as _embedded_id_tags / _isrc. - core/imports/side_effects.py:record_download_provenance reads those context fields and passes them to db.record_track_download, which now accepts the new ID kwargs and persists them. - New db.get_provenance_by_file_path with exact + basename-suffix fallback (handles container mount-root differences between download-time path and media-server-reported path). - New db.backfill_track_external_ids_from_provenance copies IDs from track_downloads onto a tracks row idempotently — COALESCE on every column preserves any value the enrichment worker already wrote (enrichment is more authoritative for late binding). - database/music_database.py:insert_or_update_media_track (the single insertion point used by every Plex / Jellyfin / Navidrome sync) calls the backfill immediately after each INSERT/UPDATE. - New core/library/track_identity.py:find_provenance_by_external_id used as a second-tier fallback in watchlist_scanner.is_track_missing _from_library — catches the window between download and media-server sync. Caller checks os.path.exists on the provenance file_path before treating it as "already in library" so a deleted file doesn't prevent re-download. Effect: freshly downloaded files become ID-recognizable to the watchlist on the very next scan, no enrichment-wait window. 19 regression tests in tests/test_provenance_id_persistence.py: - Schema migration adds expected columns + indexes - record_track_download persists every ID kwarg - record_track_download backward-compat (old kwargs still work) - get_provenance_by_file_path: exact match, basename fallback for mount-root differences, multi-record latest-wins, defensive None - backfill: copies all IDs, preserves existing via COALESCE, no-op when no provenance exists - find_provenance_by_external_id: per-ID lookup, ISRC cross-bridge, OR semantics, latest-wins on multiple matches Out of scope: backfilling provenance for files downloaded BEFORE this PR (their track_downloads rows don't carry the new IDs). Those continue to wait for enrichment. Acceptable — only affects historical files; new downloads benefit immediately. Full pytest 1625 passed; ruff clean.	3 weeks ago
Broque Thomas	ecb8939c80	Match library tracks by external IDs before fuzzy in watchlist scan Reported case (CAL): a track already on disk got re-downloaded by the watchlist scanner on every scan. Library DB had stale album metadata for the file (track tagged on album "Left Alone") while the metadata source reported it on a different album ("NPC" single). The title+artist+album fuzzy block correctly said the album names didn't match and declared the track missing — but the file's stable external IDs (Spotify ID, ISRC, etc.) unambiguously identified it as the same recording. The earlier compilation-album fix (PR #461) handled qualifier drift ("OST" vs "Music From The Motion Picture"). This case is two genuinely different album names referring to the same song. Fix: provider-neutral external-ID short-circuit before the fuzzy block in `is_track_missing_from_library`. Pulls every recognized ID off the source track (Spotify / iTunes / Deezer / Tidal / Qobuz / MusicBrainz / AudioDB / Hydrabase / ISRC), runs a single SELECT against the indexed external-ID columns on the `tracks` table, and treats any hit as "track exists in library — don't re-download". If no IDs are available (older imports without enrichment, library scans that didn't populate external IDs), falls through to the existing fuzzy logic so the safety net stays intact. New `core/library/track_identity.py` module with two helpers: - `extract_external_ids(track)`: handles dict and object-style track shapes, direct-field aliases (spotify_id / spotify_track_id / SPOTIFY_TRACK_ID), and provider-disambiguated native `id` fields (when track has `provider='deezer'` and `id='X'`, treats X as a Deezer ID). - `find_library_track_by_external_id(db, external_ids, server_source)`: builds an OR of indexed column matches with IS NOT NULL guards, optional server_source filter that also passes legacy NULL rows, single-row LIMIT. ISRC bridges across providers — a library track imported via Deezer can be matched against a Spotify scan when both sides carry the same ISRC. 43 regression tests in `tests/test_library_track_identity.py`: - 9 ID-extraction tests for direct fields (Spotify / iTunes / Deezer / ISRC / MBID / AudioDB / Hydrabase) - 8 ID-extraction tests via the provider field (8 providers + source alias + missing-provider-ignored) - 7 mixed/defensive tests (multiple IDs, object-style, empty strings, None track, numeric coercion) - 8 lookup tests (per-provider + ISRC cross-bridge) - 3 OR-semantics tests - 4 server_source filter tests - 2 ID-column-map sanity tests Full pytest 1606 passed; ruff clean.	3 weeks ago
Broque Thomas	486116c34f	Honor lossy_copy.delete_original after successful conversion Reported case (CAL): with lossy_copy.enabled=True, lossy_copy.delete_original=True, and codec=mp3, every download left both the original FLAC AND the converted MP3 in the target folder. Users opting into a lossy-only library ended up dual-format on every import. Root cause: ``core/imports/file_ops.py:create_lossy_copy`` reads ``lossy_copy.codec`` and ``lossy_copy.bitrate`` from config but never reads ``lossy_copy.delete_original``. The setting is only consulted by the pre-move source-vanished check at ``core/imports/pipeline.py:651`` (so the pipeline knows to look for a lossy variant when the FLAC has already moved on), but no code path actually deletes the source after conversion. Fix: after ffmpeg returns success and the QUALITY tag is written, check ``lossy_copy.delete_original`` and ``os.remove`` the original when enabled. Belt-and-suspenders: - Same-path guard (``os.path.normpath(out_path) != os.path.normpath(final_path)``) prevents accidentally wiping the just-converted file if a future codec choice somehow resolves out_path to the source path. - ``FileNotFoundError`` is treated as success (concurrent worker / dedup cleanup got there first). - Other ``OSError`` (permission denied, locked file) is logged but doesn't propagate — the conversion already succeeded, the user just has to clean up the original manually. Failure paths skip the delete: - ffmpeg returns non-zero → returns None, original stays - lossy_copy.enabled=False → early return before conversion runs - delete_original=False (default) → original stays 7 regression tests cover honored-when-enabled, kept-when-disabled, default-keep, ffmpeg-failure-path, lossy-disabled-path, racing-delete, and locked-file paths. Full pytest 1563 passed; ruff clean. Note: this PR does NOT address the second bug CAL mentioned (track re-downloaded despite already existing on disk). That symptom is caused by stale album metadata on the user's existing files — the library DB has the track tagged on a different album than the metadata source reports — combined with wishlist.allow_duplicate_tracks defaulting to True. Same class of issue partially addressed in PR fix/watchlist-redownload-and-duplicate-detection but compilation- album drift is the only currently-handled case. Tracking separately.	3 weeks ago
Broque Thomas	99dbe265de	Sync Qobuz auth to enrichment worker after login Discord-reported (Foxxify): logging in to Qobuz via the Connect button on Settings showed "Connected: <username> (Active)" but underneath an error said "Qobuz not authenticated...", and the dashboard indicator stayed yellow. Saving settings or reloading the tab didn't help. Root cause: SoulSync runs two QobuzClient instances side by side — one through soulseek_client.qobuz for the /api/qobuz/auth/* endpoints, and a second owned by the enrichment worker thread for thread safety. The login flow only updated the auth-flow instance's in-memory state (plus persisted to config). The dashboard's "configured" check at web_server.py:3371 reads ``qobuz_enrichment_worker.client.user_auth_token`` — the WORKER's instance — which still believed itself unauthenticated. The connection-test step at core/connection_test.py:370 hits the same worker instance for the same reason. Fix: add ``QobuzClient.reload_credentials()`` — a public, network-free method that re-reads the saved session from config and updates the instance's in-memory state + session headers. Call it on the enrichment worker's client immediately after a successful ``/api/qobuz/auth/login``, ``/api/qobuz/auth/token``, or ``/api/qobuz/auth/logout`` so the two instances stay in lockstep without waiting for the next process restart. Unlike the existing ``_restore_session()`` this skips the network probe — the caller has just authenticated, so the token is known good. A small ``_sync_qobuz_credentials_to_worker()`` helper in web_server.py wraps the call so all three endpoints share one path. 10 new regression tests cover the populate / clear / partial-config paths plus the actual two-instance-sync scenario from the bug report. Full pytest 1555 passed (the one pre-existing flake in test_tidal_auth_instructions.py is order-dependent and unrelated).	3 weeks ago
Antti Kettunen	b85a05fb88	Move image URL normalization into metadata helpers - keep existing /api/image-proxy URLs from being wrapped again - reuse the shared metadata package instead of duplicating URL logic in web_server.py - add regression coverage for proxy passthrough and internal URL normalization	3 weeks ago
Antti Kettunen	2b3022f6b0	Fix Spotify source ID fallback - Prefer real Spotify IDs when importing Spotify contexts - Skip numeric fallback IDs so Deezer values do not leak into spotify_* columns - Add regressions for import context and SoulSync library writes - Keep the route test asserting the Spotify album link	3 weeks ago
Antti Kettunen	e2bd0e1871	Split metadata source and Spotify status - Keep the primary metadata provider snapshot generic and move Spotify auth/rate-limit details into a separate status object. - Update the websocket fixture and dashboard/settings consumers to read the two buckets independently.	3 weeks ago
Antti Kettunen	36267618a3	Rename status cache to metadata_source Expose the primary metadata provider status under a generic cache key and update the websocket fixture plus frontend readers to match.	3 weeks ago
elmerohueso	f9f47f978e	fix post-download tagging, and enable tagging for hifi	3 weeks ago
Broque Thomas	7e32618f86	Drop old per-service enrichment routes after registry cutover Followup to the enrichment-bubble registry consolidation. The dashboard polling + click handlers all hit /api/enrichment/<service>/{status,pause,resume} now, so the 30 hand-rolled per-service routes in web_server.py have zero callers and can come out: /api/musicbrainz/{status,pause,resume} /api/audiodb/{status,pause,resume} /api/discogs/{status,pause,resume} /api/deezer/{status,pause,resume} /api/spotify-enrichment/{status,pause,resume} /api/itunes-enrichment/{status,pause,resume} /api/lastfm-enrichment/{status,pause,resume} /api/genius-enrichment/{status,pause,resume} /api/tidal-enrichment/{status,pause,resume} /api/qobuz-enrichment/{status,pause,resume} Worker init blocks stay (they still construct the workers + persist pause state). Section comment headers are preserved with a one-line note pointing readers at the new generic blueprint. Test fixtures in tests/conftest.py and tests/metadata/test_enrichment_events.py also updated to use the new URL paths so they reflect production reality. They were synthetic stubs that never depended on the production routes — purely cosmetic alignment. Net: ~510 lines deleted from web_server.py. Full pytest 1541 passed; ruff clean.	3 weeks ago
Broque Thomas	98c04cf332	Consolidate enrichment bubble routes behind a service registry The dashboard's enrichment-status bubbles (MusicBrainz, AudioDB, Discogs, Deezer, Spotify, iTunes, Last.fm, Genius, Tidal, Qobuz) each had its own copy-pasted /status, /pause, /resume route in web_server.py — 30 routes that differed only in the worker reference and a couple of per-service quirks (Spotify's rate-limit guard, Last.fm/Genius yield-override behavior, Tidal/Qobuz extra status fields). Replace them with a registry-driven blueprint: - core/enrichment/services.py declares an EnrichmentService dataclass with worker_getter, config_paused_key, pre_resume_check, auto_pause_token, and extra_status_defaults — all variation captured as data, no branching on service id. - core/enrichment/api.py exposes a Flask blueprint with three routes (/api/enrichment/<service>/{status,pause,resume}). Per-service quirks are honored via the descriptor: Spotify's rate-limit ban still returns 429 with `rate_limited: true`, Last.fm/Genius still drop the auto-pause token and add the yield override, Tidal/Qobuz still merge `authenticated: false` into the fallback payload. - web_server.py registers all 10 services after their workers initialize, wires the host-side hooks (config_manager.set, _download_auto_paused.discard, _download_yield_override.add), and registers the blueprint. - webui/static/enrichment.js polling + click handlers now hit the generic endpoints. The per-service `update<Service>StatusFromData` functions are unchanged — they still process the same payload. This is the cutover step. Old per-service routes are intentionally left in place as a fallback during the soak period — they currently have zero callers in the codebase and will be deleted in a follow-up patch once production has run on the new pipeline for a few days. 27 new tests in tests/test_enrichment_services.py cover the registry behavior + every quirk path through the generic blueprint (rate-limit guard, auto-pause token cleanup, persisted-pause config keys, extra default fields, worker-not-initialized fallback, exceptions). Full suite 1541 passed; ruff clean.	3 weeks ago
Broque Thomas	0e68109b68	Add filename-pass safety: require duration agreement or artist match Self-review of the previous commit found a real false-positive risk in the new filename-bucket pass: two unrelated songs that happen to share a canonical filename (e.g. ``Yellow.mp3`` by Coldplay vs by some other artist) would be grouped because all metadata gates were dropped. The filename pass now layers a safety net under ``require_metadata_match=False``: - If both rows carry a duration: must agree within 3 seconds. Same source download = identical duration; a 3+ second gap means different recordings. - Else if both rows carry an artist: relaxed 0.6 similarity check — catches dedup orphans that share an artist tag while rejecting strangers-with-same-filename. - Else (no duration AND at least one artist blank): skip — too little signal to safely group. 5 additional regression tests cover the false-positive prevention paths plus the genuine dedup-orphan scenarios that must still be caught after the safety net.	3 weeks ago
Broque Thomas	6e61890551	Stop watchlist re-downloading compilation tracks; catch slskd dedup orphans Two related bugs reported on Discord by Mushy. 1. The watchlist re-downloaded the same OST track up to 7 times. ``is_track_missing_from_library`` compared Spotify's album name and the media-server scan's album name with a raw SequenceMatcher at a strict 0.85 threshold. Compilations and soundtracks routinely fail this — Spotify reports ``"Napoleon Dynamite (Music From The Motion Picture)"`` while the Plex / Navidrome / Jellyfin tag scan saves it as ``"Napoleon Dynamite OST"``. Raw similarity ≈ 0.49, so the scanner declared the track missing on every 30-minute scan and added it back to the wishlist. The wishlist then issued a fresh download. slskd appended ``_<19-digit-ns-timestamp>`` to each new copy because the target file already existed, and the user ended up with seven copies of one song in one folder. Fix: extract two pure helpers — ``_normalize_album_for_match`` strips qualifier parentheticals (Music From X, OST, Deluxe Edition, Remastered, Anniversary, etc.) and trailing dash-clauses; ``_albums_likely_match`` checks equality after normalization, substring containment, and a relaxed 0.6 fuzzy ratio. A volume / part / disc / standalone-trailing-number guard rejects pairs like ``"Greatest Hits Vol. 1"`` vs ``"Greatest Hits Vol. 2"`` so the relaxed threshold doesn't introduce false positives on serialized releases. After this change the Napoleon Dynamite case collapses to ``"napoleon dynamite" == "napoleon dynamite"`` via the equality short-circuit and the redownload loop dies. 2. The duplicate detector found only one of the seven dupe files. The detector buckets tracks by the first 4 chars of their normalized tag title. Files written by slskd directly into a library folder often get inconsistent (or blank) tags from the media-server rescan, so the seven copies were bucketed apart by parsed title and never compared. Fix: refactor the per-bucket comparison into ``_scan_bucket``, then add a second pass — ``_build_filename_buckets`` re-buckets leftover tracks by canonical filename stem (slskd dedup tail stripped via ``_strip_slskd_dedup_suffix``, same regex the import-cleanup PR uses) plus extension. Filename agreement is itself strong evidence the files came from the same source download, so the second pass calls ``_scan_bucket`` with ``require_metadata_match=False`` to skip the title / artist / cross-album gates. The same-physical-file guard still runs so bind-mount duplicates aren't flagged. 72 new regression tests across two files cover the album-match helpers (28 tests including the Napoleon Dynamite scenario, 7 volume disagreements, 8 positive/negative pairs, 5 defensive cases) and the new filename-bucket pass (16 tests across bucket construction, scan integration, and existing title-pass behavior). Full pytest 1509 passed; ruff clean. Reported by Mushy in Discord.	3 weeks ago
Broque Thomas	46d8e15674	Prune slskd dedup orphans after import slskd appends "_<19-digit unix-nanosecond timestamp>" to a downloaded filename when the destination already contains a same-named file (concurrent downloads of the same track, partial-file retries after a connection drop, cancelled-then-redownloaded files, the same track surfacing in multiple synced playlists). The file-finder code already recognized the suffix when matching a download to its source — but after the canonical file moved into the library, the leftover "_<timestamp>" siblings sat orphaned in the downloads folder forever. Reported on Discord by Shdjfgatdif. cleanup_slskd_dedup_siblings() runs at the end of each successful import (3 safe_move_file sites in pipeline.py) and prunes any remaining siblings that strip down to the canonical stem with the same extension. Conservative match (>= 18 trailing digits) keeps legitimate filenames like "Track 5" and "Album 1995" untouched. Per- file unlink failures are swallowed so a single locked file doesn't block the rest. 17 regression tests cover the suffix-strip primitive, orphan removal, no-op cases, mismatched extensions, subdirectories, and partial-failure recovery.	3 weeks ago
BoulderBadgeDad	94b08bbb49	Merge pull request #457 from kettui/refactor/spotify-auth-flow Clarify metadata source gating and Spotify auth flow	3 weeks ago
Broque Thomas	ab85c45785	Restore soulsync logger state between parallel-imports tests Importing web_server fires utils.logging_config.setup_logging at module-init, which clears + re-installs handlers on the shared 'soulsync' logger and pins its level to the user's configured value. That mutation leaks across tests in the same pytest process. This file runs alphabetically before test_library_reorganize_orchestrator, so the leak broke test_watchdog_warns_about_stuck_workers downstream — it relies on caplog capturing soulsync.library_reorganize warnings via root-logger propagation, and the reconfigured logger's new handler chain swallowed those records before they reached caplog (caplog.records came back empty even though pytest's live-log capture clearly showed the warning fired). Adds an autouse fixture that snapshots the soulsync logger's handlers, level, and propagate flag before each test in this file and restores them afterwards. Pollution stays scoped to this file. tests/test_tidal_auth_instructions.py also imports web_server but runs alphabetically AFTER test_library_reorganize_orchestrator so it never tripped this — fix is scoped here, not a project-wide conftest, so we don't change behaviour for unrelated test files.	3 weeks ago
Antti Kettunen	74e3cc460c	Simplify service status and labels - Flatten the Spotify service-status rendering so it shows rate-limit and recovery states explicitly, while otherwise displaying the active metadata provider directly. - Keep the Spotify auth controls and metadata-source picker aligned with the real session state after authenticate and disconnect flows. - Return "Unmapped" for unknown metadata source labels instead of implying iTunes. - Update the metadata registry tests to cover the new label fallback.	3 weeks ago
Antti Kettunen	55603be14c	Clarify Spotify auth flow and sync UI - Send Spotify auth completion back to the opener so the settings page refreshes immediately - Make the local auth flow go straight through to Spotify instead of showing the temporary instruction page - Keep the remote/docker instruction page available for manual callback setups - Sync Spotify status, connect/disconnect buttons, and metadata source selection after auth and disconnect - Keep the disconnect behavior aligned with the active primary metadata source	3 weeks ago
Antti Kettunen	9646f6ca7f	Clarify Spotify auth actions - Hide the auth button when a Spotify session is active - Treat disconnect as a session change, not a provider swap - Share metadata source labels in the registry - Tighten rate-limit copy around Spotify-specific behavior	3 weeks ago
Broque Thomas	f339211654	Parallelize singles-import processing with a 3-worker executor Discord-reported (fresh.dumbledore + maintainer ack): the /api/import/singles/process route iterated staging files through a plain Python for loop. Per-file work is dominated by metadata search round-trips (Spotify/iTunes/Deezer/Discogs), so a multi- track manual import on a typical home network was painfully slow. Adds a dedicated import_singles_executor (3 workers) alongside the existing executor pool, and refactors the route to submit every file at once and aggregate results via as_completed. Worker count balances throughput against any single provider's per-source rate limits — the same shape used by missing_download_executor. Extracts the per-file pipeline into _process_single_import_file which returns a typed (status, payload) outcome: - ("ok", final_title) on success - ("error", message) for missing/malformed input or pipeline failure The worker wraps its own exceptions so a single bad file can't crash the batch; the route adds a belt-and-suspenders try/except around future.result() for any worker-level surprises. Pipeline thread-safety verified: post_process_matched_download already serializes per-file via post_process_locks (one lock per context_key — and each import gets a unique UUID context_key), DB writes serialize through SQLite's WAL + busy_timeout, metadata registry uses RLocks, no bare module-level mutable state. Adds 9 regression tests: - 4 worker-contract tests (missing file, malformed match, pipeline exception wrapping, happy-path return shape) - 2 executor-config tests (worker count, thread name prefix) - 1 integration test that proves the route actually parallelizes by checking wall-clock duration is well under sequential cost - 1 mixed-outcome aggregation test - 1 worker-crash recovery test Doesn't address the related "stops on tab close" complaint — that's a separate request-lifecycle issue that needs job_id + polling, not just parallelism.	3 weeks ago
Broque Thomas	99a38a6201	Route imported singles/EPs through album_path template Discord-reported (winecountrygames + fresh.dumbledore): "Import only makes Albums folder no singles or eps". Users with a ${albumtype}s/$albumartist/... album_path template saw an "Albums" folder fill up correctly but never any "Singles" or "EPs" folder. build_import_album_info detected an album using ``total_tracks > 1`` AND ``album_name != track_title``. Spotify singles fail both — total_tracks is 1 and the album is usually named after the song. The result was that staging/auto-import routed singles through single_path, which doesn't honour $albumtype, so the user's per-type folder layout never applied. Now also treats the metadata source's explicit release-type classification ("single", "ep", "compilation") as evidence that this is an album-shaped release, so it routes through album_path and the user's $albumtype substitution runs. The default fallback value "album" is deliberately excluded from this check so single-track downloads with no real metadata behave exactly as before. Adds 10 regression tests covering the reported scenario, EP and compilation explicit types, and three guards: normal multi-track albums still detected, default 'album' type falls through, and empty/unknown types fall through.	3 weeks ago
Broque Thomas	1e5204a230	Show Tidal callback port (not Spotify's) in auth instructions Discord-reported: clicking the Tidal "Authenticate" button on a Docker setup landed users on a remote-access instructions page that told them their callback URL would look like http://127.0.0.1:8888/tidal/callback?code=... — Spotify's port, hardcoded into the Tidal instructions. Users who followed those instructions literally saved 8888 into their tidal.redirect_uri setting; that mismatched their Tidal Developer App's registered :8889 redirect URI and Tidal returned error 1002 (invalid redirect URI) on every auth attempt. Pull the port from the actual TidalClient.redirect_uri the OAuth URL was just built with (urlparse), with the SOULSYNC_TIDAL_CALLBACK_PORT env var as fallback when the URI can't be parsed. Both the Step 2 example and the Step 3 highlighted URL now reflect whatever Tidal port the user is actually configured to use. Adds 3 regression tests covering the reported scenario, custom callback ports via SOULSYNC_TIDAL_CALLBACK_PORT, and a defensive fallback when redirect_uri is unparseable. Tests hit the real /auth/tidal route through Flask's test client and assert the rendered HTML, so future hardcoded ports get caught immediately.	3 weeks ago
Broque Thomas	ef03901cb4	Bulk watchlist add: fall back through every source ID, not just active The /api/library/watchlist-all-unwatched endpoint required the user's currently active metadata source's ID column on each library artist. A Spotify-primary user with library artists only matched against iTunes or Deezer saw them silently skipped — surfacing on Discord as "Library and Watchlist not syncing correctly". The per- artist Enhanced View sync sometimes "fixed" them because it triggered metadata enrichment that occasionally populated the missing Spotify ID, but couldn't help artists Spotify simply doesn't carry. Extracts the picker as a standalone helper so it can be tested directly: core/watchlist/source_picker.py:pick_artist_id_for_watchlist Picks the active source first when available, then falls back through spotify -> itunes -> deezer -> discogs in registration order. Empty strings count as missing. Numeric IDs are coerced to str so SQLite's TEXT columns store them in the same form library code reads back. Returns (None, None) only when the artist has zero source IDs — the only legitimate skip reason now. Adds 10 regression tests covering active-source priority for each supported primary, fallback ordering through every secondary, the zero-IDs base case, unrecognized active source (e.g. hydrabase still falls through), empty-string handling, and numeric coercion.	3 weeks ago
Broque Thomas	ddef904414	Match featured-artist tracks across discography completion Discord-reported scenario: a single "Super Single" by Artist1 feat. Artist2 is also on Artist1's "Super Album". When the album is fully owned, Artist1's discography correctly shows the single as complete, but Artist2's discography (where the same track also appears as a single) shows it as missing. Two layers needed for the fix: Scanner: the Jellyfin/Emby path was keeping only ArtistItems[0], which is almost always equal to the album artist — so the distinguishing per-track credit was silently suppressed. Now joins every ArtistItems entry with "; " and stores the value when there are multiple credits OR when the single credit differs from the album artist. Plex's originalTitle already carries the full multi- artist tag, so Plex users benefit without needing the scanner change. Scorer: _calculate_track_confidence now splits track_artist on the common multi-artist delimiters real-world tags use (",", ";", "&", "feat.", "ft.", "featuring", "vs.", "x") and scores each piece independently against the search artist, taking the max along with the whole-string similarity as the floor. Never reduces a score — purely additive matching for previously-missed featured-artist credits. Adds 12 regression tests covering the reported scenario, primary- artist back-compat, every delimiter variant (parametrized), no- regression on exact match, and the scanner storing every ArtistItem. Existing Jellyfin-scanned rows persist their old single-artist value until the next library scan rewrites them; Plex rows benefit immediately on next match without needing a rescan.	4 weeks ago
Broque Thomas	f1ec62bad3	Add fallback negative-case test for track-artist matching Re-enabling the previously-dead album-aware fallback could in theory leak false positives if its 0.8 album-title floor were ineffective. Pin the floor with a clearly-mismatched album hint ("Disney Hits" against "Ray of Light") and assert the search returns no match. Distinct artist names with no shared words so the main path actually fails through to the fallback (the prior draft used "Different Artist" / "Real Artist" which both contain "Artist" and scored above the main path's threshold, never reaching the fallback at all).	4 weeks ago
Broque Thomas	345273df22	Match soundtrack tracks against per-track artist, fix dead fallback Two bugs surfacing the same user-reported symptom: a Vaiana OST track ("Where You Are" by Christopher Jackson) wouldn't match against a Plex/Emby library because the album sits under the album artist (Lin-Manuel Miranda). Bug 1: the data was already there but scoring ignored it. The DB schema has a tracks.track_artist column, the scanner populates it from Plex's originalTitle and Jellyfin's ArtistItems[0], and the SQL WHERE clause already searches it — but _rows_to_tracks dropped the column on its way to the Python object, and _calculate_track_confidence only scored against the album-artist JOIN. Candidates whose track- artist matched got returned by the search and then immediately filtered out by the low confidence score. Fix: _rows_to_tracks now propagates row['track_artist'] onto the returned object, and _calculate_track_confidence takes the better of (album-artist similarity, track-artist similarity) so soundtracks match through whichever credit the search query carries. Bug 2: the album-aware fallback path constructed DatabaseTrack with kwargs the dataclass doesn't accept (artist_name, album_title, server_source). Every row TypeError'd, the outer except swallowed it silently, and the fallback never matched anything since the column was added — invisible because nothing logged it. Fix: build DatabaseTrack with valid fields and attach the joined columns afterwards, the same pattern _rows_to_tracks uses. Adds 6 regression tests covering: track-artist match (the OST case), album-artist still matches, scorer takes the better of the two, defensive handling for tracks without track_artist, search-path attribute propagation, and the previously-dead album-aware fallback.	4 weeks ago
Broque Thomas	7698405f58	Surface handler-returned errors in automation last_error The "Clean Search History" automation card kept showing a stale 'DownloadOrchestrator' object has no attribute 'base_url' error even after the underlying handler bug was fixed in `77d20e9`. Root cause is in the engine, not that handler: AutomationEngine only captured uncaught exceptions into last_error. Handlers that report failure by RETURNING {'status': 'error', ...} were treated as successful from the engine's perspective, so subsequent gracefully-failing runs never updated the row to reflect the current state. Both the timer (run_automation) and event (_handle_event_trigger) paths now extract the error string from a result whose status is 'error', falling through 'error' -> 'reason' -> 'message' -> a placeholder so last_error is never None on actual failures regardless of which key the handler chose. Existing behaviour for raised exceptions and successful runs is preserved. Also normalizes _auto_clean_search_history's return key from 'reason' to 'error' so older deployed engines that only check the canonical key still see the failure. Adds 7 regression tests covering every result shape the engine might receive.	4 weeks ago
BoulderBadgeDad	05a4342ac8	Merge pull request #445 from kettui/refactor/remove-quality_scanner-spotify-prio Refactor quality scanner to respect primary metadata provider	4 weeks ago
BoulderBadgeDad	09591ad089	Merge pull request #449 from Nezreka/fix/duplicate-detector-mount-paths Filter same-physical-file duplicates from duplicate detector	4 weeks ago
Broque Thomas	b9c8245c49	Stop config retry tests from writing to the real DB The fixture used the wrong env var name (SOULSYNC_DB_PATH) when trying to redirect ConfigManager at a tmp directory. ConfigManager actually reads DATABASE_PATH (config/settings.py:49), so the test ConfigManager loaded — and then saved — at the user's real database/music_library.db. The retry stub in test_lock_errors_during_retries_log_at_debug_not_error calls the real _save_to_database after its mocked failures, which then clobbered the encrypted app_config row with the test fixture's stub payload {"plex": {"base_url": "http://example.test"}}. Three layers of fix so this can't happen again: - Use the correct env var (DATABASE_PATH). - Pin mgr.database_path / mgr.config_path on the instance after construction, so the test fixture's tmp paths win even if ConfigManager's resolution logic changes. - Assert the resolved database_path is rooted under tmp_path before returning the fixture, so the test refuses to run if it would touch a non-tmp DB.	4 weeks ago
Broque Thomas	382e427117	Filter same-physical-file duplicates from duplicate detector When users bind the same host music directory into both SoulSync (e.g. /app/Transfer) and a media server like Plex (e.g. /media/Music), both scans add a track row pointing at the same physical file via different mount paths. The detector previously flagged those as duplicate groups even though there's only one file on disk. New _is_same_physical_file helper filters pairs where: - The trailing 3 path segments match (filename + album + artist folder), so they're the same release on disk. - The leading mount roots actually differ. - Durations agree within 1s when both rows carry duration data. Adds 10 regression tests covering the reported scenario plus edge cases (Windows separators, case differences, missing durations, sibling-album false-positive guard).	4 weeks ago
Broque Thomas	4238aeb4d9	Add regression tests for config DB retry behaviour (#434 ) Pin the new save-retry contract so future changes can't silently re-introduce the spam reported in #434: - Happy-path saves emit zero ERROR logs. - Transient locks during retries log at DEBUG, not ERROR. - Six attempts run before giving up, with the documented backoff schedule (0.2 + 0.5 + 1.0 + 2.0 + 4.0s). - Genuine exhaustion logs a single ERROR and writes config.json. - sqlite3.OperationalError("database is locked") routes to DEBUG; any other OperationalError still logs ERROR. - _connect_db() actually applies WAL + busy_timeout + synchronous=NORMAL. Also moves `import time` from inside _save_config to the module top so the tests can monkeypatch sleep cleanly.	4 weeks ago
Antti Kettunen	fb190e16ca	Coerce wishlist track counts before category checks - normalize album.total_tracks before comparing it in wishlist classification - avoid mixed-type comparisons when provider payloads serialize track counts as strings - add regression coverage for numeric strings and invalid values	4 weeks ago
Antti Kettunen	2bc8e8a27b	Preserve artwork in quality scanner wishlist handoff - carry track-level album art through the quality scanner normalization path - preserve artist artwork when provider results expose it - keep album.image_url and album.images populated so the wishlist UI can render the cover consistently - add a regression test covering provider payloads with image_url on both the track and artist	4 weeks ago
Antti Kettunen	761fc29523	Mock streaming prep sleep in tests - add an autouse fixture that patches stream-prep sleep to a no-op - keep the polling branches covered while removing the long test delay	4 weeks ago
Antti Kettunen	c97a072f54	Refactor quality scanner to respect primary metadata provider - search metadata providers in source-priority order for each generated query instead of caching one client for the whole scan - keep the quality-scanner worker provider-neutral and preserve the no-provider error path - update the quality-scanner tests and remove the obsolete web_server spotify_client injection	4 weeks ago
Antti Kettunen	fd30d2a0be	Rename wishlist lifecycle helper - Switch the download lifecycle over to the neutral wishlist track helper name - Keep the old Spotify helper as a compatibility alias for older callers - Store track_data as the primary failed-download wishlist payload key and add regression coverage	4 weeks ago
Antti Kettunen	b1a9c1b458	Accept wishlist track_data aliases - Let the wishlist service accept both track_data and spotify_track_data - Preserve the backward-compatible wrapper while avoiding the keyword argument crash - Add a regression test for the alias path	4 weeks ago
Antti Kettunen	0fa692f935	Make wishlist respect configured providers - add neutral wishlist payload helpers while keeping legacy Spotify aliases - route wishlist removal and classification through generic track data - keep API and service compatibility for existing callers	4 weeks ago
BoulderBadgeDad	58a4c1905b	Merge pull request #419 from kettui/refactor/metadata-service-split-and-metadata-client-management-optimizations Split metadata service logic into separate modules, move client management out of web_server	4 weeks ago
Broque Thomas	5c8b8b271a	Lift _prepare_stream_task + playlist_explorer_build_tree to core/ Final lift in the web_server.py extraction effort. Pulls two route handlers + one background worker out of `web_server.py` into new focused packages: - `core/streaming/prepare.py` — 258-line stream-prep worker that downloads a track to the local Stream/ folder for the browser audio player. - `core/playlists/explorer.py` — 305-line route handler for `POST /api/playlist-explorer/build-tree` that streams an NDJSON discography tree from a mirrored playlist. What `prepare_stream_task` does: 1. Reset stream state to 'loading' with the new track info. 2. Clear any prior file from Stream/ (only one stream lives there). 3. Spin up a fresh asyncio event loop and `soulseek_client.download()`. 4. Poll progress every 1.5s. Queue timeout 15s; overall 60s. 5. On succeeded + bytes-match: find the file with retry, move into Stream/, signal slskd completion, mark state 'ready' with file_path. 6. On error/timeout/cancel: state goes to 'error' or 'stopped'. 7. Finally: tear down the event loop cleanly. What `playlist_explorer_build_tree` does: 1. Validate request, load playlist + tracks from DB. 2. Pick active metadata source (Spotify if authed, else fallback). 3. Group tracks by artist using discovered matched_data when the provider matches the active source. 4. Stream NDJSON: meta line → one artist line per group → complete line. 5. Per artist: cache check → resolve discography → tag releases with `in_playlist` flag based on title-similarity match → filter by mode (`albums` = only matches; `discographies` = full disco). 6. Mark playlist as explored on completion. Strict 1:1 byte parity: Both functions exposed their dependencies through proxy patterns established in earlier lifts (PR4–PR8). For prepare_stream_task, `stream_state` is a deps property; for the explorer, Flask `request` / `jsonify` / `Response` are injected via deps so the lifted body keeps its native syntax. Both lifts verified ZERO diff against the original after `deps.X` → global X normalization. 258 lines orig = 258 lines lifted (prepare_stream_task). 305 lines orig = 305 lines lifted (explorer). Bonus cleanup: web_server.py's module-level `import shutil` and `import glob` were now unused (only `_prepare_stream_task` used them at module scope; every other reference is via inline `import shutil` in respective function bodies). Removed both module-level imports — ruff caught the F811 redefinitions and confirmed they're truly redundant. Dependencies for `PrepareStreamDeps` (11 fields): config_manager, soulseek_client, stream_lock, project_root, docker_resolve_path, find_streaming_download_in_all_downloads, find_downloaded_file, extract_filename, cleanup_empty_directories, plus 2 stream_state property delegates. Dependencies for `PlaylistExplorerDeps` (9 fields): Flask request/Response/jsonify, spotify_client, get_database, get_active_discovery_source, get_metadata_fallback_client, get_metadata_fallback_source, get_metadata_cache. Tests: 6 new under tests/streaming/test_prepare.py (state init, Stream/ folder creation + clearing, download-init failure, completed + moved + ready state, partial-bytes incomplete-warning path) plus 9 new under tests/playlists/test_explorer.py (5 validation early-exit paths, streaming response shape with meta/complete lines, mark- explored side effect, discovered-artist grouping using matched_data, provider mismatch falling back to raw artist name). Full suite: 1355 passing (was 1340). Ruff clean. End of the web_server.py extraction effort. Started at ~45,000 lines across PR4–PR8 + this commit; finished around 35,000 lines with the heavy worker + route logic now living in domain-cohesive packages under core/. The remaining bulk in web_server.py is route handlers, service initialization, and the deferred 1530-line `_register_automation_handlers` (startup-only, marginal lift value).	4 weeks ago
Broque Thomas	91978656a5	Lift enhance_artist_quality to core/artists/quality.py Pulls the 284-line artist quality enhancement helper out of `web_server.py` into a new `core/artists/` package. Flask route handler split: route + request parsing stay in web_server.py, the body lifts to a pure function returning `(payload_dict, http_status_code)`. What `enhance_artist_quality` does: 1. Validate request: track_ids must be non-empty, artist must exist. 2. Build a `track_lookup` from `database.get_artist_full_detail` so each selected track resolves with its album context. 3. Per track: - Read current quality tier from the file extension. - Build `matched_track_data` for the wishlist entry, in priority order: - Spotify direct lookup via stored `spotify_track_id` (preferred). Uses raw API data when available; otherwise rebuilds the payload and pulls album images via a follow-up `get_album` call. - Spotify search fallback using matching_engine queries with artist+title similarity scoring (album-type bonus for albums, smaller bonus for EPs). Stops at first >= 0.9 confidence match. - iTunes/fallback source search with the same scoring shape. - Add to wishlist via `wishlist_service.add_spotify_track_to_wishlist` with `source_type='enhance'` and a `source_context` carrying the original file path, format tier, bitrate, original_tier, and artist_name. - Tally `enhanced_count` / `failed_count` / per-track failure reasons. 4. Return `{success, enhanced_count, failed_count, failed_tracks}` 200. Dependencies injected via `ArtistQualityDeps` (7 fields) — spotify_client, matching_engine, get_database, get_wishlist_service, get_current_profile_id, get_quality_tier_from_extension, get_metadata_fallback_client. Diff vs original after `deps.X` → global X normalization is 1 line of cosmetic drift — the success return now uses an explicit `(payload, 200)` tuple to keep all returns shape-consistent for the wrapper. Flask treats `jsonify(x)` and `(jsonify(x), 200)` identically. 284 lines orig = 285 lines lifted, body otherwise byte-identical. Tests: 10 new under tests/artists/test_quality.py covering input validation (empty track_ids, artist not found), Spotify direct lookup via raw_data, Spotify direct lookup with enhanced format requiring album image rebuild, Spotify search fallback, iTunes/fallback source match path, track-not-found and no-file-path failure modes, complete no-match failure, and source_context payload assertions (enhance flag, file path, format tier, bitrate, source_type). Full suite: 1340 passing (was 1330). Ruff clean.	4 weeks ago
Broque Thomas	3a6597561a	Lift _execute_retag to core/library/retag.py Pulls the 258-line retag worker out of `web_server.py` into a new `core/library/` package. Pure 1:1 lift — wrapper keeps the original entry-point name so the retag-trigger endpoint continues to work without changes. What `execute_retag` does: 1. Fetch album + track metadata for the new `album_id` (Spotify or iTunes — the Spotify client transparently falls back). 2. Load existing files in the retag group from the DB. 3. Match each existing track to a new Spotify track: - Priority 1: same disc + track number. - Priority 2: title similarity >= 0.6 (SequenceMatcher). 4. For each matched pair: - Re-write metadata tags via `_enhance_file_metadata`. - Compute the new path via `_build_final_path_for_track` and move the audio file (plus .lrc / .txt sidecars) if the path changes. - Drop an orphaned cover.jpg if it's left in an empty directory. - Clean up empty parent directories left behind. - Download the new cover art into the new album dir. 5. Update the retag group record with new artist / album / image / total_tracks / release_date and the appropriate Spotify-or-iTunes album ID (numeric → iTunes, alphanumeric → Spotify). 6. Mark the retag state 'finished' (or 'error' on exception). Strict 1:1 byte parity: The original mutated `retag_state` as a module global (the function declared `global retag_state` even though it only mutates in place). Here `retag_state` is exposed through the `RetagDeps` proxy as a Python property so the lifted body keeps `name[key] = value` / `name.update(...)` syntax. The property setter rebinds the web_server.py reference if the function ever reassigns it (currently it doesn't, but the setter is wired for parity with the watchlist lift). Diff vs original after `deps.X` → global X normalization is zero differences apart from the dropped `global retag_state` decl and the inline `from database.music_database import get_database` (replaced by deps.get_database()). 258 lines orig = 258 lines lifted, byte-identical body otherwise. Dependencies injected via `RetagDeps` (13 fields) — config_manager, retag_lock, spotify_client, plus 8 callable helpers (get_audio_quality_string, enhance_file_metadata, build_final_path_for_track, safe_move_file, cleanup_empty_directories, download_cover_art, docker_resolve_path, get_database) and 2 property delegates (_get_retag_state / _set_retag_state). Tests: 11 new under tests/library/test_retag.py covering setup error paths (no album data, no album tracks, no existing tracks), track-number priority match, title-similarity fallback, no-match skip, missing file skip, file move when path changes, group record update (spotify vs iTunes ID branching by alphanumeric vs numeric album_id), multi-disc total_discs computation. Full suite: 1330 passing (was 1319). Ruff clean.	4 weeks ago
Broque Thomas	2b2003ba4c	Lift _process_watchlist_scan_automatically to core/watchlist/auto_scan.py Pulls the 390-line watchlist auto-scan orchestrator out of `web_server.py` into a new `core/watchlist/` package. Watchlist (followed-artists scanner that finds new releases) is a separate domain from kettui's wishlist (failed-download retry queue), so this lift does not overlap with the ongoing PR400-style extractions. What `process_watchlist_scan_automatically` does: 1. Smart stuck-detection guard before acquiring the timer lock — prevents deadlock when a previous scan flag is dangling past the 2-hour timeout. 2. Inside the timer lock: re-check + set the active scan flag with the current timestamp. 3. Per-profile expansion (or single-profile when manually triggered): - Watchlist count check + Spotify auth gate. - Backfill missing artist images. 4. Initialize a fresh `watchlist_scan_state` dict (the deps property setter rebinds the web_server.py module-level name so external sentinel checks via id() comparison still detect the swap). 5. Pause enrichment workers, then call `WatchlistScanner.scan_watchlist_artists` with a per-event progress callback that translates scanner events into automation log lines. 6. Post-scan steps (skipped if the scan was cancelled mid-flight): - Populate discovery pool from similar artists (per-profile). - Refresh ListenBrainz playlists. - Update current seasonal playlist (weekly cadence). - Generate Last.fm radio playlists. - Sync Spotify library cache. - Activity feed entry + automation_engine.emit('watchlist_scan_completed'). 7. On exception: mark state['status']='error', re-raise so the automation wrapper records the failure. 8. Finally: resume enrichment workers, clear the scanner's rescan cutoff, reset the auto-scanning flag. Strict 1:1 byte parity: The original mutated `watchlist_auto_scanning`, `watchlist_auto_scanning_timestamp`, and `watchlist_scan_state` as module globals (with a leading `global` decl). Here those names are exposed through the `WatchlistAutoScanDeps` proxy as Python properties so the lifted body keeps the same `name = value` / `name[key] = value` shape. Property setters fan writes back to web_server.py via callback pairs. Diff vs original after `deps.X` → global X normalization is zero differences apart from the dropped `global` declaration line — Python doesn't need it once the names are property accesses on the deps object. 390 lines orig = 390 lines lifted, byte-identical body otherwise. Dependencies injected via `WatchlistAutoScanDeps` (15 fields total) — Flask app, spotify_client, automation_engine, watchlist_timer_lock, plus 5 callable helpers and 6 property delegate callbacks (paired get/set for each of the three globals). Tests: 11 new under tests/watchlist/test_auto_scan.py covering stuck-detection guard, race-check inside lock, zero-watchlist short- circuit, unauthenticated Spotify gate, successful scan with all post- scan steps, automation event emission, activity feed logging, cancellation mid-scan skipping post-steps, profile-scoped trigger, flag reset in finally, rescan cutoff clear in finally. Full suite: 1319 passing (was 1308). Ruff clean.	4 weeks ago
Antti Kettunen	e6c2bee427	Move profile Spotify cache into registry - let core.metadata.registry own per-profile Spotify client caching - register the DB-backed profile credentials provider from web_server.py - invalidate only the affected profile cache entry on save, delete, and auth	4 weeks ago
Antti Kettunen	50e1ae3a3f	Move metadata helpers into package modules - split metadata lookup logic into core/metadata/* - keep core/metadata_service.py as the legacy barrel - update tests and artist-detail code to patch concrete modules	4 weeks ago

1 2 3 4

194 Commits (42f3026eef7ba02aba21d4aef71084011c991c19)