Skip to content
CloudMagicTalk to us

Media & Entertainment

Making a 1.2 PB media archive searchable with generative AI

A Chicago-based media archive and licensing company

Automated metadata tagging and clip discovery across a petabyte of archival video

10-20x
more content tagged per editorial hour
Industry
Stock footage licensing
Scale
~1.2 PB archive
Engagement
GenAI metadata pipeline + storage modernization
Architecture shown
Google Cloud reference design

The challenge

  • Metadata was produced entirely by hand, so most of the archive was effectively unsearchable
  • Around 1.2 PB of media sat split across on-premises storage and a separate cloud bucket
  • Non-English archival content had no transcription, so it could not be discovered at all
  • Storage tiering decisions were manual and rarely revisited

What it had to do

  • Generate usable metadata at archive scale, not clip by clip
  • Validate tags against the licensing taxonomy already in use
  • Surface only the edge cases that need human judgement
  • Remove manual storage-tier decisions permanently

What we built

We built a pipeline that processes a video, identifies candidate clips, generates descriptions, validates tags against the existing licensing taxonomy, and surfaces only the cases needing human judgement. Editors review AI output instead of producing metadata from scratch. We also delivered a phased migration plan for the full archive with a security gap analysis and a cost comparison.

Automated tagging pipeline

Clip detection, description generation and taxonomy validation in one pass.

Human review on edge cases only

Editors adjudicate uncertain output rather than writing all of it.

Cross-language discovery

Generated subtitles opened up content that existed only in non-English audio.

Tiering that manages itself

Storage cost tracks real access patterns with no manual audit.

Reference architecture

Archive

  • Cloud Storage
  • Autoclass tiering

Extract

  • Video Intelligence API
  • Speech-to-Text

Enrich

  • Gemini on Vertex AI
  • Cloud Run jobs

Discover

  • BigQuery catalogue
  • Vector Search

Results

  • 10 to 20 times more content processed for tagging per unit of editorial time versus the manual workflow
  • Editors shifted from producing metadata to reviewing it, compressing time to a licensable asset
  • Non-English archival content made discoverable through generated subtitles for the first time
  • A phased migration plan delivered for the full 1.2 PB archive with a security gap analysis and cost comparison
  • Storage tiering made automatic, so cost follows real access patterns

For editorial

Review replaces transcription, and the backlog finally moves.

For licensing

More of the archive is findable, so more of it can be sold.

For finance

Storage cost follows usage instead of a stale tiering decision.

  • Cloud Storage
  • Video Intelligence API
  • Speech-to-Text
  • Gemini on Vertex AI
  • Cloud Run
  • BigQuery
  • Vector Search

Facing something similar?

Tell us where you are now. A senior engineer replies, usually within one business day.