Manufacturing & Industrial
Illustrative exampleThe predictive maintenance project a warehouse quote killed
A global industrial manufacturer
Shelved as too expensive, then delivered on Databricks at a fifth of the quoted cost
- Industry
- Industrial manufacturing
- Scale
- 6 plants, 300+ machines, 1B+ readings a day
- Platform
- Databricks on Google Cloud
- First line live
- 7 weeks
The challenge
- Six plants and more than 300 machines produce over a billion sensor readings a day
- The warehouse-centric design priced always-on compute and nightly full refreshes, and the quote landed near two million dollars a year
- Vibration-analysis vendors wanted per-machine licensing on top
- The project was shelved, and maintenance stayed reactive
What it had to do
- Keep 90 days of telemetry hot and queryable in seconds
- Score anomaly models on the stream, not in a nightly batch
- Give engineering, quality, and finance governed access to one copy of the data
- Land the run cost where the project is obviously worth doing
What we built
The quote died on two assumptions we removed: that telemetry must live in always-on warehouse compute, and that every transformation reruns nightly from scratch. On Databricks, plant historians and sensor gateways stream through Pub/Sub into Structured Streaming, landing as Delta Lake tables on Cloud Storage where storage is cheap and compute only runs when there is work. Delta Live Tables process increments instead of full refreshes, Photon and liquid clustering keep scans tight, and jobs run on spot capacity with serverless for the spiky paths. MLflow manages a per-line anomaly model scored directly on the stream, and Unity Catalog governs who sees what across engineering, quality, and finance. BI kept its tools: BigQuery federation and Looker read the same governed tables.
Streaming lakehouse ingestion
Pub/Sub into Structured Streaming into Delta Lake. Telemetry lands in open tables on object storage, and compute scales to zero between jobs.
Incremental, not full refresh
Delta Live Tables process what changed. The nightly rebuild that priced the warehouse design out of existence never happens.
Models on the stream
MLflow-managed anomaly models score readings as they arrive and retrain nightly on increments, flagging failure signatures roughly 48 hours ahead.
One governed copy
Unity Catalog rules access for engineering, quality, and finance on the same tables. No extracts, no per-machine licenses.
Reference architecture
Plants
- PLC and SCADA historians
- Vibration sensors
- MES
Stream
- Pub/Sub
- Structured Streaming
Governed lakehouse
- Delta Lake on Cloud Storage
- Delta Live Tables
- Unity Catalog
Act
- MLflow anomaly models
- Maintenance queue
- BigQuery and Looker
Results
- The platform runs at roughly a fifth of the shelved warehouse quote
- 90 days of telemetry stays hot and answers in seconds
- Failure signatures surface roughly 48 hours before breakdown
- First production line live in 7 weeks, inside the existing Google Cloud estate
For maintenance teams
Planned interventions with parts on hand, instead of 2am callouts.
For plant leadership
Downtime and line health on one live view across all six plants.
For the data team
Open Delta tables under one catalog, readable from the BI tools they already run.
- Databricks on Google Cloud
- Delta Lake
- Delta Live Tables
- Unity Catalog
- Photon
- Structured Streaming
- MLflow
- Pub/Sub
- Cloud Storage
- BigQuery
- Looker
Facing something similar?
Tell us where you are now. A senior engineer replies, usually within one business day.
A senior engineer reads every message and replies within one business day.