Skip to content
CloudMagicTalk to us

Manufacturing & Industrial

Illustrative example

The predictive maintenance project a warehouse quote killed

A global industrial manufacturer

Shelved as too expensive, then delivered on Databricks at a fifth of the quoted cost

5x
cheaper than the warehouse quote that shelved the project
Industry
Industrial manufacturing
Scale
6 plants, 300+ machines, 1B+ readings a day
Platform
Databricks on Google Cloud
First line live
7 weeks

The challenge

  • Six plants and more than 300 machines produce over a billion sensor readings a day
  • The warehouse-centric design priced always-on compute and nightly full refreshes, and the quote landed near two million dollars a year
  • Vibration-analysis vendors wanted per-machine licensing on top
  • The project was shelved, and maintenance stayed reactive

What it had to do

  • Keep 90 days of telemetry hot and queryable in seconds
  • Score anomaly models on the stream, not in a nightly batch
  • Give engineering, quality, and finance governed access to one copy of the data
  • Land the run cost where the project is obviously worth doing

What we built

The quote died on two assumptions we removed: that telemetry must live in always-on warehouse compute, and that every transformation reruns nightly from scratch. On Databricks, plant historians and sensor gateways stream through Pub/Sub into Structured Streaming, landing as Delta Lake tables on Cloud Storage where storage is cheap and compute only runs when there is work. Delta Live Tables process increments instead of full refreshes, Photon and liquid clustering keep scans tight, and jobs run on spot capacity with serverless for the spiky paths. MLflow manages a per-line anomaly model scored directly on the stream, and Unity Catalog governs who sees what across engineering, quality, and finance. BI kept its tools: BigQuery federation and Looker read the same governed tables.

Streaming lakehouse ingestion

Pub/Sub into Structured Streaming into Delta Lake. Telemetry lands in open tables on object storage, and compute scales to zero between jobs.

Incremental, not full refresh

Delta Live Tables process what changed. The nightly rebuild that priced the warehouse design out of existence never happens.

Models on the stream

MLflow-managed anomaly models score readings as they arrive and retrain nightly on increments, flagging failure signatures roughly 48 hours ahead.

One governed copy

Unity Catalog rules access for engineering, quality, and finance on the same tables. No extracts, no per-machine licenses.

Reference architecture

Plants

  • PLC and SCADA historians
  • Vibration sensors
  • MES

Stream

  • Pub/Sub
  • Structured Streaming

Governed lakehouse

  • Delta Lake on Cloud Storage
  • Delta Live Tables
  • Unity Catalog

Act

  • MLflow anomaly models
  • Maintenance queue
  • BigQuery and Looker

Results

  • The platform runs at roughly a fifth of the shelved warehouse quote
  • 90 days of telemetry stays hot and answers in seconds
  • Failure signatures surface roughly 48 hours before breakdown
  • First production line live in 7 weeks, inside the existing Google Cloud estate

For maintenance teams

Planned interventions with parts on hand, instead of 2am callouts.

For plant leadership

Downtime and line health on one live view across all six plants.

For the data team

Open Delta tables under one catalog, readable from the BI tools they already run.

  • Databricks on Google Cloud
  • Delta Lake
  • Delta Live Tables
  • Unity Catalog
  • Photon
  • Structured Streaming
  • MLflow
  • Pub/Sub
  • Cloud Storage
  • BigQuery
  • Looker

Facing something similar?

Tell us where you are now. A senior engineer replies, usually within one business day.