Retail & E-Commerce
Off five hand-managed VMs and onto containers
A high-volume CRM and order management provider
A high-volume CRM platform migrated out of two data centres, then run as managed operations
- Industry
- E-commerce / CRM and order management
- Scale
- 900+ data tasks across 5 independent VMs
- Engagement
- Full migration + hybrid automation + managed operations
- Architecture shown
- Google Cloud reference design
The challenge
- 900+ scheduled data tasks ran across five independent virtual machines with no orchestration
- Infrastructure sat across two separate hosting providers, one of them a legacy data centre
- Resource allocation for each workload was decided and adjusted by hand
- Every capacity change was a manual operation, so nothing scaled on demand
What it had to do
- Consolidate the workloads onto one orchestrated platform
- Migrate the databases without interrupting a live commerce platform
- Replace hand-tuned resource allocation with scheduling
- Keep the platform running afterwards, not just deliver the move
What we built
We consolidated the task fleet onto an orchestrated container platform so scheduling and resource allocation stopped being a manual decision, migrated the databases off the legacy estate, and codified the whole environment so it is reproducible rather than hand-built. We then stayed on to run it.
Orchestrated task fleet
900+ jobs scheduled on a container platform instead of pinned to five VMs.
Database migration
Production data moved off the legacy estate without interrupting commerce traffic.
Everything as code
The environment is reproducible from source rather than hand-configured.
Managed operations
Monitoring, alerting and logging handed over as a running service, not a document.
Reference architecture
Sources
- Legacy VMs
- FTP & web tiers
Orchestrate
- GKE
- Cloud Scheduler
- Workflows
Data
- Cloud SQL
- Cloud Storage
Operate
- Terraform
- Cloud Monitoring
- Cloud Logging
Results
- 80 to 90% less manual effort spent allocating resources to individual workloads
- 900+ data processing tasks consolidated onto one orchestrated platform
- Databases migrated off the legacy data centre without interrupting the live platform
- Infrastructure defined as code, so environments are reproducible instead of hand-built
- Monitoring, alerting and logging in place, with operations run as an ongoing service
For engineering
Capacity is a scheduling concern, not a person adjusting a VM.
For operations
One platform to watch instead of five machines across two providers.
For the business
Peak order volume no longer needs someone standing by.
- GKE
- Cloud Scheduler
- Workflows
- Cloud SQL
- Cloud Storage
- Terraform
- Cloud Monitoring
Facing something similar?
Tell us where you are now. A senior engineer replies, usually within one business day.
A senior engineer reads every message and replies within one business day.