Ingestion Pipeline¶
Durable Operation Tracking¶
Ingestion no longer relies only on in-memory watcher/debounce/worker state for
queue visibility. When a file is discovered, Tuvima Library creates or reuses a
durable media_operations row with operation_type = ingestion.file.
The operation stage is updated as the item moves through discovery, settling, lock checks, queueing, hashing, parsing, scoring, registration, and identity queueing. The batch item endpoint reads these durable rows, so a restart can still show what was waiting, running, interrupted, blocked, or failed.
Durable ingestion status is the source of truth for:
- queued files
- queue position
- active stage
- retry and blocked state
- interrupted work after restart
- per-batch item rows
The legacy ingestion log can still exist as historical/support data, but product
status should use media_operations.
The retail, Wikidata bridge, and quick hydration workers remain durable pollers,
but they also use an in-process signal to wake as soon as upstream work is
created. Timed polling remains the fallback after restarts or missed signals.
Identity retry settings live in config/hydration.json:
identity_retry_max_attempts, identity_retry_base_delay_seconds,
identity_retry_max_delay_seconds, identity_retry_jitter_min_ms, and
identity_retry_jitter_max_ms.
Storage Reset And Reingest¶
The current storage epoch is guid-blob-v1. Internal IDs are stored in SQLite as
16-byte BLOB GUIDs, while API contracts still expose GUIDs as strings. Legacy
TEXT-GUID databases are rejected at startup unless TUVIMA_STORAGE_RESET=1 (or
destructive-reingest) is set, in which case the old database files are renamed
as backups and a clean database is initialized.
For normal development rebuilds, POST /dev/reset-and-seed pauses file watching,
clears catalogued library/database/cache/artwork state while preserving accounts,
profiles, access rules, settings, configured libraries, and provider configuration,
recreates the selected Standard or Stress fixtures, queues them through grouped
normal ingestion scans, and resumes watching. POST /ingestion/rescan remains the
canonical no-wipe operation for configured source media. Reset paths include guards
that refuse destructive cleanup when a library output path overlaps a source folder.
Fresh-ingest validation is the proof path for sequence, artwork, and attribution changes. Do not repair bad historical rows in place for these fixes. Stop the Engine and Dashboard, wipe only the configured validation roots, start both apps again, ingest the representative fixtures, then validate database/API state before visual Dashboard checks.
Validation roots:
C:\temp\tuvima-watch\booksC:\temp\tuvima-watch\audiobooksC:\temp\tuvima-watch\tvC:\temp\tuvima-watch\moviesC:\temp\tuvima-watch\musicC:\temp\tuvima-watch\comicsC:\temp\tuvima-library
This document describes how Tuvima Library discovers, processes, organises, and stages media files - from the moment a file appears in a watched folder to the moment it is promoted into the organised library.
Libraries and sources¶
The Engine is configured with logical Libraries, each declaring:
| Field | Description |
|---|---|
kind |
catalogued or personal |
area |
Read, Watch, Listen, or View |
metadata_policy |
Enriched, local-preferred, local-only, or manual |
sources |
Stable per-folder source identities and safety policy |
primary_destination_source_id |
The explicit managed destination; never inferred from ordering |
accepted_intake_modes |
The direct intake mechanisms accepted by this library |
Configuration lives in config/libraries.json. Schema 6 is required; obsolete top-level watch folders, flat source paths, and read-only flags are rejected.
File watching is source-folder aware. A flush that contains files from one source records that source path on the ingestion batch; a flush that spans more than one source records Multiple source folders. Watcher noise is buffered for Ingestion:FswQuietPeriodSeconds seconds, defaulting to 30 seconds, before the batch is released to the debounce queue.
{
"schema_version": "6.0",
"libraries": [
{
"id": "44444444-4444-4444-8444-444444444444",
"name": "Movies",
"category": "Movies",
"kind": "catalogued",
"area": "watch",
"presentation": "catalogue",
"metadata_policy": "enriched",
"media_types": ["Movies"],
"sources": [{
"id": "44444444-aaaa-4444-8444-444444444444",
"path": "/media/library/movies",
"role": "primary_destination",
"management_mode": "managed_by_tuvima",
"source_type": "local_folder",
"include_subdirectories": true,
"access_mode": "writable",
"participates_in_organization": true,
"intake_role": "direct"
}],
"primary_destination_source_id": "44444444-aaaa-4444-8444-444444444444",
"visibility": "shared",
"accepted_intake_modes": ["incoming_folder", "browser_upload"],
"duplicate_policy": "skip_exact",
"organization_policy": { "mode": "tuvima_standard", "preserve_originals": true }
}
]
}
The category tells the Engine where to organise files on disk. The media_types tell it how to process them - which processor to use, which metadata providers to query, and what confidence prior to apply during identification. A library folder designated for Movies gives any MP4 it finds an 0.80 prior confidence for that media type, skipping much of the heuristic disambiguation that would otherwise be needed.
Ingestion Snapshot¶
The Dashboard's Ingestion page uses GET /ingestion/operations, backed by IIngestionOperationsStatusService, as the application-facing ingestion status model. The service aggregates existing persisted state rather than keeping a separate demo task model:
ILibraryItemRepositoryfor total, registered, provisional, and review lifecycle countsIIngestionBatchRepositoryfor active and recent batchesidentity_jobsandingestion_logfor pipeline stage countsstage_progressrows in the snapshot for the numbered user-facing stage barsingestion_batch_artifactsfor batch-scoped media, metadata, artwork, people, relationship, QID, and review artifacts used by later Activity rollupsreview_queuefor actionable review reason groupsconfig/libraries.jsonfor Read, Watch, Listen, the internal personal-library bridges behind View Personal Spaces, stable sources, destinations, and incoming locationsprovider_healthand provider config files for provider status- runtime
IngestionOptionsandcore.jsonfor organization rule summaries
Live Dashboard updates come from the snapshot endpoint. Existing SignalR Intercom events (BatchProgress and IngestionProgress) trigger a refresh of that same snapshot model instead of maintaining separate UI progress state. Polling is faster while jobs are active and slower while idle.
GET /ingestion/operations includes a stage_progress collection. Each row has:
| Field | Meaning |
|---|---|
stage_number |
User-facing order. Stages 1-8 are progress rows; Stage 9 is review exception state for metrics and Activity rollups. |
stage_key |
Stable key such as scan, retail, wikidata, people, or deep_artwork. |
label |
Display label for the stage. |
completed_files / total_files / percent_complete |
File progress against the active batch total. |
active_count / queued_count |
Current in-flight and queued work where the backend can measure it. |
status_label |
Human-readable state: pending, active, complete, stale, or needs review. |
active_item_label |
Exact item label only when per-file state is known. |
active_group_label / active_group_count |
Group label for batched calls such as Wikidata reconciliation. |
label_accuracy |
ExactItem, GroupedLookup, BatchOnly, Stale, or None. |
artifact_label / artifact_count |
Running artifact count for the stage. |
detail_items |
Optional server-generated detail rows shaped as { label, value, tone?, icon? } for expanded UI content and artifact math. |
last_updated_time / is_stale |
Freshness signal for active worker data. |
The compact Ingestion page renders Stages 1-8 as progress rows:
| # | Stage | Progress Rule | Artifact Count |
|---|---|---|---|
| 1 | Scan | Files accepted, skipped, duplicated, or failed / total files | Accepted files |
| 2 | Read Details | Files parsed, skipped, duplicated, or failed / total files | Parsed files |
| 3 | Retail Match | Files retail matched, review-ready, no-result, skipped, or failed / total files | Provider matches |
| 4 | Wikidata | Files with QID, no QID, not applicable, review-ready, skipped, or failed / total files | QIDs |
| 5 | Ready | Files organized, writeback-ready, visible, or terminal / total files | Added files |
| 6 | People | Files person-enriched, skipped, or not applicable / total files | People |
| 7 | Relationships | Files relationship-enriched, skipped, or not applicable / total files | Links |
| 8 | Artwork | Files deep-artwork completed, skipped, or not applicable / total files | Assets |
Stage 3 is the collapsed retail metadata, quick metadata, and primary cover/poster bar. Stage 8 is later deep artwork enrichment. Stages 6, 7, and 8 may run concurrently after their retail/Wikidata prerequisites exist. For grouped Tuvima.Wikidata or provider work, the backend must show a group label instead of an exact file label unless it has a correlation key for a specific file.
Stage 7 prefers explicit series order values over Wikidata previous/next backlink consistency. Missing or contradictory public Wikidata chain links are diagnostics, not Review Queue work, unless they expose a local conflict that needs a curator decision.
Review remains live, but it is not a progress row in the Dashboard. stage_progress can still expose Stage 9 for API consumers and artifact ledger support; the Dashboard uses the top Need Review metric plus latest-batch review delta, and recent batch rows repeat the review count with other artifact totals.
Music remains a conservative organization lane. The status surface emphasizes tag/fingerprint-first handling and preserving album folders instead of implying aggressive rename/move behavior.
Intake Modes¶
Library source monitoring¶
The Engine monitors every catalogued library source for new files. Existing-library sources are indexed in place and remain read-only. Managed writable sources may organize files when their source and library policies allow it.
The .staging/ directory within the library root is excluded from library source monitoring to prevent re-ingestion loops.
Import Mode¶
Import mode performs a one-time scan of an existing collection. It follows the same processing steps as Watch mode, then either moves or copies the file depending on import_action. Copy mode leaves originals untouched. After an import completes, the folder can optionally be switched to Watch mode for ongoing monitoring.
Processing Steps¶
Every file - regardless of intake mode - goes through the same sequential processing pipeline:
Code-Level Stage Chain¶
IngestionEngine is the lifecycle facade for watcher startup, pause/resume,
scanning, shutdown, and dry-run entry points. Per-file work runs through the
ordered IIngestionStage chain:
settle/detecthash/dedupeprocessscore/identifyorganizewrite-backidentity-job creation
The hash/dedupe stage acquires the content-hash lock, and the coordinator releases it only after the chain terminates. Duplicate resolution, asset registration, organization-gate review creation, safe write-back, and identity job creation therefore retain the same serialized scope.
The organize stage is deliberately a readiness and review decision during
initial ingestion. It does not move the file into the final library. The file
remains in place until the retail-first identity pipeline has enough context for
AutoOrganizeService to promote it. Write-back is likewise deferred for files
still in a monitored source folder because changing their bytes would change the content
hash and trigger re-ingestion.
Nullable integrations are grouped with the stage that consumes them: hash cache, capability planning, provisional-review outcomes, managed artwork/export, and the identity-pipeline wake signal. This keeps optional behavior from leaking into unrelated stages.
1. Settle¶
The Engine waits briefly after detecting a file to confirm it has finished being written to disk. This prevents reading partially-copied files from network shares or slow storage.
2. Lock Check¶
The Engine verifies that no other process has an exclusive lock on the file before attempting to read it.
3. Fingerprint¶
A SHA-256 content hash is computed from the file's bytes. This hash is the file's permanent identity throughout its lifetime in the library. It survives renaming, moving, and metadata edits. If a file is ingested a second time (e.g. after a database rebuild), the hash allows the Engine to recognise it immediately.
4. Scan¶
The appropriate processor for the file's format opens the file and extracts all embedded metadata:
- EpubProcessor - reads OPF package metadata: title, author, publisher, year, series, language, cover image
- AudioProcessor - reads ID3v2 (MP3), iTunes atoms (M4B/M4A), Vorbis comments (FLAC/OGG): title, artist, album, track number, chapter markers, genre, ASIN, embedded artwork
- VideoProcessor - reads container metadata (MP4, MKV): title, resolution, duration, codec, embedded subtitles, chapter list
- ComicProcessor - reads ComicInfo.xml from CBZ/CBR archives: title, series, issue number, writer, artist, publisher
The processor also emits media type candidates when the format is ambiguous. See the Media Type Disambiguation section below.
5. Identify¶
The Priority Cascade Engine scores all available claims for this file - from embedded metadata, filename parsing, and any prior library folder hints - and assigns the file to an existing Collection or creates a new one. This is where the title, author, series, and other canonical values are resolved.
If multiple files from the same source folder have already been processed (e.g. a TV season with 22 episodes), the Engine uses Ingestion Hinting: the first file's resolved metadata is cached as a folder-level prior. Subsequent siblings receive the collection ID, QID, and bridge IDs from that prior as high-confidence claims, dramatically reducing the number of Wikidata lookups needed.
Work deduplication fallback: MediaEntityChainFactory checks whether a Work already exists before creating a new one (matching by title + author + media type via IWorkRepository). When canonical_values has not yet been populated for an in-flight asset, the deduplication check falls back to a raw metadata_claims lookup so that duplicate files arriving close together in time do not bypass the check. Duplicate files create a new Edition under the existing Work rather than creating a duplicate Work.
6. Move to Staging¶
The file is moved from its source location into {LibraryRoot}/.data/staging/, where it waits for hydration and promotion. Cover art is extracted as a claim at this stage and persisted through the managed asset store when the artwork pipeline writes it.
Managed artwork is stored through AssetPathService and entity_assets under .data/assets/...; it is never keyed by provisional QID folders.
Staging-First Flow¶
All ingested files land in .staging/ before reaching the organised library. The library invariant is that every file within the library root (outside .staging/) has been hydrated, has reached a settled identity outcome, and has a settled artwork outcome. That may mean a resolved QID with art present, or a precision-preserving QID-missing result with artwork explicitly confirmed missing.
Library source --(detect + process)--> .staging/ --(hydration + promote)--> Managed destination
|
stays here if:
- low confidence
- unidentifiable
- needs review
- media type ambiguous
Staging Subcategories¶
Files are routed to one of four subcategories based on their overall confidence score after the Identify step:
| Subcategory | Condition | Behaviour |
|---|---|---|
.staging/pending/ |
Confidence >= 0.85, or any user-locked claim | AutoOrganizeService promotes after hydration |
.staging/low-confidence/ |
Confidence 0.40 - 0.85, no user locks | Awaits hydration improvement or manual review |
.staging/unidentifiable/ |
Confidence < 0.40, no user locks | Requires user to provide a title or match |
.staging/other/ |
Resolves to "Other" category | Requires media type classification |
Staging lifecycle logging: The Engine logs staging progress at Information level at four points: (1) when an asset is moved into staging, (2) when a review queue item is created, (3) when a gap is detected in expected review creation (e.g. a confidence score that should have triggered a review but did not), and (4) when an asset is promoted out of staging into the organised library. These log entries allow the staging pipeline to be audited from the activity log.
browse readiness gate: main browse surfaces visibility is no longer a simple "in staging or not" decision. The shared library item projection computes browse visibility, pipeline step, artwork state, and readiness from identity jobs, review state, and canonical artwork flags. An item is visible in the main browse surfaces only after it has a non-placeholder title, a resolved media type, and settled artwork (present, or missing after explicit settlement). Review-only or still-hidden items remain available in Activity, Review, and the Review Queue.
AutoOrganize Gate¶
AutoOrganizeService promotes a staged file to the organised library when:
This threshold (AutoLinkThreshold = 0.85) is defined once in ScoringConfiguration and reused by both the staging router and the promotion gate. It governs filesystem promotion, not main browse surfaces visibility.
Hero Banner¶
Hero banner generation (blur + vignette + grain, via SkiaSharp) runs during promotion by AutoOrganizeService. It is a post-hydration step, not an ingestion step, because it benefits from the enriched metadata and high-resolution cover art that hydration provides.
Manual Reclamation¶
Staged files retain their fingerprint and metadata in the database. A user can manually resolve a staged file from the Dashboard - by dragging it to a Collection or providing a user-locked title - triggering promotion to the organised library structure. The .staging/ directory is excluded from library source monitoring to prevent re-ingestion loops.
On startup, if {LibraryRoot}/.orphans/ exists and .staging/ does not, the Engine renames the directory and updates all database file paths automatically.
File Organisation¶
Data Authority¶
The database is the authoritative data store for all metadata, relationships, and canonical values. User metadata edits are additionally written back into the file's embedded metadata via IMetadataTagger (EPUB OPF, ID3 tags, M4B atoms), ensuring portability - the file carries its own metadata independently of the database.
Wikidata properties are re-fetchable through provider reconciliation. Managed artwork is indexed in the database through entity_assets; optional local sidecars are exports, not runtime fallback reads.
Recovery scenarios: - Standard: Scheduled SQLite backups (by domain: universe, people, library) as primary recovery - Wikidata data loss: Re-fetch via batch Reconciliation API - Full wipe: Re-ingest from library root; file embedded metadata and batch Wikidata reconciliation rebuild the library
Folder Structure Templates¶
The default organisation template is:
Per-media-type overrides:
| Media Type | Template |
|---|---|
| Books | {Category}/{Title} - {QID}/Epub/{Title}{Ext} |
| Audiobooks | {Category}/{Title} - {QID}/Audiobook/{Title}{Ext} |
| TV | {Category}/{Title} - {QID}/S{Season:00}E{Episode:00} - {EpisodeTitle}{Ext} |
| Music | {Category}/{Artist}/{Album} - {QID}/{TrackNumber:00} - {Title}{Ext} |
| Movies | {Category}/{Title} - {QID}/{Title}{Ext} |
| Comics | {Category}/{Title} - {QID}/{Title}{Ext} |
Books and Audiobooks share the same title folder under the Books category, distinguished by their format subfolder. This means an ebook and its audiobook counterpart live at:
{LibraryRoot}/Books/Dune - Q190159/Epub/Dune.epub
{LibraryRoot}/Books/Dune - Q190159/Audiobook/Dune.m4b
{LibraryRoot}/.data/assets/artwork/Work/{workId}/CoverArt/{variantId}.jpg
Cover art is owned by the work-level entity asset, so ebook and audiobook variants can share the same preferred cover without duplicating files beside each media item.
Category Mapping¶
The {Category} path segment is derived from the file's media type:
| Media Types | Category folder |
|---|---|
| Epub, Audiobook | Books |
| TV | TV |
| Movies | Movies |
| Music | Music |
| Comics | Comics |
| Unknown | Other |
Migration Note¶
Existing libraries organised under older folder patterns continue to work. On the next hydration pass or a manual "Re-organise Library" action, files are moved to the current structure automatically.
Media Type Disambiguation¶
Some file formats map to multiple possible media types. Magic bytes identify the container format but not the content type. An MP3 file could be an audiobook chapter or a music track. An MP4 could be a feature film or a TV episode. The disambiguation system resolves this using heuristic signals treated as voted claims.
Signal Sources¶
Media type is resolved using the same Weighted Voter architecture as all other metadata fields. Multiple signals emit competing candidates with associated confidence values:
| Signal source | Confidence range | Examples |
|---|---|---|
| Magic bytes (unambiguous formats) | 0.95-1.0 | EPUB -> Books, CBZ -> Comics, M4B -> Audiobooks |
| Processor heuristics | 0.30-0.80 | File duration, bitrate, chapter markers, genre tag |
| Filename and path patterns | 0.25-0.65 | S01E01 in filename -> TV, audiobooks in path -> Audiobooks |
| User lock | 1.0 | Manual override - always wins |
Confidence Thresholds¶
| Threshold | Behaviour |
|---|---|
>= 0.70 (auto_assign_threshold) |
Accept automatically, proceed normally |
0.40-0.70 (review_threshold) |
Accept provisionally, create AmbiguousMediaType review queue entry |
| < 0.40 | Assign MediaType.Unknown, block auto-organize, create review entry |
AudioProcessor Disambiguation¶
The AudioProcessor runs at priority 95 (above VideoProcessor at 90) and handles audio format detection.
Unambiguous assignments:
- .m4b -> Audiobooks (0.98 confidence)
- .flac, .ogg, .wav -> Music (0.95 confidence)
For ambiguous formats (.mp3, .m4a), the processor emits weighted candidates using additive heuristic signals:
- Duration: Very long files (> 60 min) bias toward Audiobooks; short files (< 5 min) bias toward Music
- Chapter markers: Presence of chapter metadata strongly indicates Audiobooks
- Genre tags: Genre values matching known audiobook indicators (e.g. "Spoken Word", "Audiobook") or music genres (e.g. "Rock", "Jazz") push the score in respective directions
- Album and track metadata: Presence of track numbers and album names strongly indicates Music
- Bitrate: Low bitrate speech-range audio biases toward Audiobooks
- Path keywords: Parent folder names like
audiobooks,musicin the source path - File size: Very large single files bias toward Audiobooks
Each type (Audiobook, Music) starts at a base score of 0.25. Signals are additive and the final scores are normalized to [0.0, 1.0] before comparison against the confidence thresholds.
VideoProcessor Disambiguation¶
The VideoProcessor resolves ambiguity between Movies and TV:
- TV filename patterns:
SxxExxorNxNNpatterns in the filename strongly indicate TV - Duration: Short files bias toward TV episodes; feature-length files bias toward Movies
- Path keywords: Parent folder structures containing season or series names
- Sibling file count: Many similarly-named files in the same folder indicate a TV series
Base score per type (Movie, TV) is 0.35. Signals are additive and normalized to [0.20, 0.90].
Configuration¶
All disambiguation thresholds and heuristic parameters - duration bands, bitrate thresholds, path keywords, genre tag lists, TV filename patterns - are configurable in config/disambiguation.json. No code changes are needed to tune the system's behaviour.
Review Resolution¶
When a file lands in the review queue with an AmbiguousMediaType trigger, the user selects the correct media type from candidate cards in the Needs Review tab. The selected type is saved as a user-locked claim at confidence 1.0, the review item is resolved, and the hydration pipeline re-runs for that entity.
After Stage 3 retail metadata returns 3 or more claims, the pipeline can auto-resolve pending AmbiguousMediaType review items - the provider results provide enough signal to confirm the media type without user input.
Writeback & Auto Re-tag Sweep¶
Once a file is identified and enriched, WriteBackService embeds the canonical metadata back into the file itself so external players, re-ingestion, and library rebuilds see it without consulting the database. The per-media-type field list lives in config/writeback-fields.json - the single source of truth shared by the taggers and the media detail editor.
Per-media-type writeback hash¶
Every media asset carries a writeback_fields_hash column (migration M-084) that combines:
- The SHA-256 of the JSON slice for the asset's media type in
writeback-fields.json. - The version constant of the specific tagger that wrote the file (
VideoMetadataTagger.TaggerVersion,AudioMetadataTagger.TaggerVersion,EpubMetadataTagger.TaggerVersion,ComicMetadataTagger.TaggerVersion).
A file is considered stale when its stored hash differs from the currently-computed hash. Bumping a tagger version or editing the field list for that media type invalidates the hash for every matching file.
Pending diff + Apply flow¶
WritebackConfigState (singleton) watches writeback-fields.json via IConfigurationLoader. When the file changes, the state computes a pending diff (added and removed fields per media type) and surfaces it without running anything. The Auto Re-tag Sweep card on the Maintenance settings tab shows the diff and two buttons:
- Apply - commits the pending diff to
CurrentHashesand signals the worker to start a sweep. - Run Now - re-runs the sweep against the current hashes without applying a new diff (useful if a tagger version was bumped).
No files are touched until the user clicks Apply or Run Now.
RetagSweepWorker¶
A BackgroundService that wakes on either a cron schedule (config/maintenance.json -> schedules.retag_sweep, default 0 3 * * *) or the PendingApplied signal from WritebackConfigState. Each pass:
- Calls
IMediaAssetRepository.GetStaleForRetagAsyncto find identified assets whosewriteback_fields_hashdiffers from the current hash (or is NULL). - Processes in batches, calling
WriteBackService.WriteMetadataAsync(assetId, "config_change")for each asset. On success, the service stamps the new hash on the row. - Classifies failures via
RetagFailureClassifier: - Locked / IoFailed ->
ScheduleRetagRetryAsyncwith the next off-hours window start. The sweep picks these up on the next run. - Corrupt / Unknown (after retries exhausted) -> inserts a
ReviewQueueEntrywith triggerWritebackFailed, routing the file to the Review Queue. - Broadcasts live progress via SignalR (
RetagSweepProgressandRetagSweepCompletedevents) so the Maintenance tab shows a processed / succeeded / transient / terminal counter during a sweep.
Endpoints¶
| Route | Role | Purpose |
|---|---|---|
GET /maintenance/retag-sweep/state |
Admin or Curator | Returns pending diff + current hashes |
POST /maintenance/retag-sweep/apply |
Admin | Commits pending diff, signals worker |
POST /maintenance/retag-sweep/run-now |
Admin | Signals worker without applying a diff |
POST /maintenance/retag-sweep/retry/{assetId} |
Admin | Manual retry for a specific failed asset |
Review Queue integration¶
Review Queue surfaces WritebackFailed review items alongside other review triggers. The message reads "Re-tag failed - file may be locked or corrupt"; resolving the item either manually (fix the file, click retry) or via a successful next-sweep pass clears the review row.
Supported Library Types¶
| Library Type | Formats | Notes |
|---|---|---|
| Books | EPUB, PDF | Combined with Audiobooks under the Books category |
| Audiobooks | M4B, MP3, M4A | Combined with Books under the Books category |
| TV | MP4, MKV, AVI, WebM | Season/episode folder structure |
| Movies | MP4, MKV, AVI, WebM | Single-work, flat folder structure |
| Music | MP3, FLAC, OGG, M4A, WAV | Album = Collection, Track = Work |
| Comics | CBZ, CBR | Sequential art; ComicInfo.xml metadata |
Future library types planned but not yet implemented:
- Other - YouTube videos, lectures, personal recordings, and any media that does not fit the primary types. Files would be stored and user-provided metadata accepted, but automated enrichment would be limited.
- View intelligence - Optional local features such as deeper EXIF/XMP extraction, maps, face/object detection, memories, and event grouping. This future privacy-sensitive scope builds on the isolated local-asset index rather than catalogue ingestion.
Related¶
Series Manifest Hydration¶
After Stage 4 resolves a Wikidata QID and full property claims have been persisted, WikidataBridgeWorker asks WikidataSeriesManifestHydrationService whether the item belongs to a canonical series. The service only uses QID-backed relationship facts such as P179/series_qid, never fuzzy title matching.
For books, audiobooks, comics, and TV, a canonical series QID triggers a Tuvima.Wikidata manifest fetch. Tuvima stores every named item in series_manifest_items, including missing works the user does not own. Later imports from the same series first link against the cached named manifest, so adding another Dune ebook or audiobook usually does not require downloading the whole series again while the cache is fresh.
Manifest hydration is diagnostic unless the external container is sequence compatible for the media type. Wikimedia list articles, publisher production lists, franchises, and universes must not become lane shelves. Provider-backed containers are preferred when they are more specific: Comic Vine volumes for comics, Apple collections for albums, TMDB seasons and film collections for watch media, and Wikidata manifests for books/audiobooks when no retail sequence source exists.