Case Study

How a European energy utility validated resilience across 500 workloads

A major European energy utility runs roughly 500 workloads across four Azure regions. The environment had grown too large to reason about by hand, leaving optimization reactive and disaster recovery impossible to validate with confidence.

Working with InfrOS, the utility mapped its whole environment as one architecture, compared options with the trade-offs in view, and put every change under continuous conformance. The result: over 30% of in-scope cloud spend identified as reducible, a cross-region rebuild rehearsed at a four-hour Recovery Time Objective (RTO) and a fifteen-minute Recovery Point Objective (RPO), and drift caught in hours instead of at the next audit.

Key Takeaways

Map the architecture, not the resources

Seeing ~500 workloads as one connected architecture, not a list of resources, is what makes both optimization and recovery possible.

Compare whole architectures, not parts

Every change has a ripple effect. Comparing complete designs, weighed on the utility's own priorities, surfaces the real trade-offs.

Prove it before production

The strongest options were validated against the requirements before deployment, so decisions rested on evidence rather than assumption.

Stop drifting, start Evolving™

Because the model stays connected to the live environment, drift is caught in hours, not at the next audit.

No one can hold 500 workloads in their head

The utility supplies electricity to millions of people. Its systems run on Microsoft Azure, where availability, security, and compliance are non-negotiable.

The environment had grown to nearly 500 workloads across four regions. It was too large to reason about by hand. The team could manage individual resources, but not understand how the environment behaved as a system.

Simple questions had become hard to answer:

Traditional tools list resources. They do not show the architecture. Without that view, optimization was reactive, disaster recovery could not be confidently validated, and senior engineers spent weeks analyzing an environment that only grew more complex.

They did not need another inventory tool. They needed architectural control.

  • Which workloads depend on each other?
  • How would a regional failure hit critical services?
  • Where are workloads over- or under-provisioned?
  • Which design choices are driving unnecessary cost?
  • Which changes improve resilience without adding risk?

See, compare, prove, and keep evolving

The team needed to see the environment as one architecture, compare real options, decide with evidence, and keep evolving. They did it in four steps.

1. See the whole environment as one architecture

With read-only access, the team could finally see the environment as a single live model, with ~500 workloads, their dependencies, and data flows, not just a list of resources. This became the basis for every decision that followed.

2. Compare whole-architecture options, with the trade-offs in view

For each workload, the team could weigh the current design against alternatives in other regions or another cloud. Every option was measured against their own priorities - cost, performance, resilience, security, compliance, data residency, and recovery. Because every architecture change has a ripple effect, options were compared as complete architectures, so the team could see the trade-offs and decide with evidence.

3. Prove the decision, then deliver through your own process

The strongest options were validated against the team's requirements before production. Approved decisions became production-ready Terraform, raised as pull requests in the team's own repository, with the requirements travelling alongside the code. The team reviewed, approved, and deployed through its own Git process. The decision stayed in their hands.

4. Stop drifting, start evolving

The model stays connected to the live environment and keeps checking whether the architecture still meets its requirements. When something drifts, the team gets clear options for what to change next, with the impact laid out. They approve; the change ships. Evolution runs on evidence, not guesswork.

Architectural control, and the proof to back it

More than 30% of in-scope cloud spend was identified as reducible - specific, evidenced opportunities tied to individual workloads. Disaster recovery became provable rather than theoretical: a cross-region rebuild rehearsed at a ~4-hour RTO and ~15-minute RPO, with an alternate-cloud path held in reserve.

Because the model stays live, drift, capacity gaps, security issues, and optimization opportunities surface in hours, not at the next audit. Most importantly, the same model used to understand the environment and make the initial decision continues to evaluate it as workloads, technology, costs, and priorities change.

>30%

spend reducible

~4h/~15m

RTO / RPO, rehearsed

~
115
K

resources mapped

Hours

to catch drift