Media & Entertainment
Making a 1.2 PB media archive searchable with generative AI
A Chicago-based media archive and licensing company
Automated metadata tagging and clip discovery across a petabyte of archival video
- Industry
- Stock footage licensing
- Scale
- ~1.2 PB archive
- Engagement
- GenAI metadata pipeline + storage modernization
- Architecture shown
- Google Cloud reference design
The challenge
- Metadata was produced entirely by hand, so most of the archive was effectively unsearchable
- Around 1.2 PB of media sat split across on-premises storage and a separate cloud bucket
- Non-English archival content had no transcription, so it could not be discovered at all
- Storage tiering decisions were manual and rarely revisited
What it had to do
- Generate usable metadata at archive scale, not clip by clip
- Validate tags against the licensing taxonomy already in use
- Surface only the edge cases that need human judgement
- Remove manual storage-tier decisions permanently
What we built
We built a pipeline that processes a video, identifies candidate clips, generates descriptions, validates tags against the existing licensing taxonomy, and surfaces only the cases needing human judgement. Editors review AI output instead of producing metadata from scratch. We also delivered a phased migration plan for the full archive with a security gap analysis and a cost comparison.
Automated tagging pipeline
Clip detection, description generation and taxonomy validation in one pass.
Human review on edge cases only
Editors adjudicate uncertain output rather than writing all of it.
Cross-language discovery
Generated subtitles opened up content that existed only in non-English audio.
Tiering that manages itself
Storage cost tracks real access patterns with no manual audit.
Reference architecture
Archive
- Cloud Storage
- Autoclass tiering
Extract
- Video Intelligence API
- Speech-to-Text
Enrich
- Gemini on Vertex AI
- Cloud Run jobs
Discover
- BigQuery catalogue
- Vector Search
Results
- 10 to 20 times more content processed for tagging per unit of editorial time versus the manual workflow
- Editors shifted from producing metadata to reviewing it, compressing time to a licensable asset
- Non-English archival content made discoverable through generated subtitles for the first time
- A phased migration plan delivered for the full 1.2 PB archive with a security gap analysis and cost comparison
- Storage tiering made automatic, so cost follows real access patterns
For editorial
Review replaces transcription, and the backlog finally moves.
For licensing
More of the archive is findable, so more of it can be sold.
For finance
Storage cost follows usage instead of a stale tiering decision.
- Cloud Storage
- Video Intelligence API
- Speech-to-Text
- Gemini on Vertex AI
- Cloud Run
- BigQuery
- Vector Search
Facing something similar?
Tell us where you are now. A senior engineer replies, usually within one business day.
A senior engineer reads every message and replies within one business day.