A regulated enterprise cut Azure operations from 20 hours to 23 minutes | Firemind
Case study

A regulated enterprise cut Azure operations from 20 hours to 23 minutes

23 tasks, zero disruptions, every action verified by a human.

When you cannot staff the estate you already have

The client is a regulated Nordic enterprise. For this pilot, Firemind built a new, purpose-built Azure sandbox: a compact environment of around six virtual machines across two resource groups and two networks.

Industry
Regulated industry
Environment
Purpose-built Azure sandbox
Pilot scope
~6 VMs, two resource groups
Delivered by
Firemind IT Operations Engine

Scope: this pilot ran on a single non-production Azure sandbox, built new for the pilot, over an eight week window, in which the engine had access for four and a half weeks. It was not a production deployment and does not cover the client’s wider estate. Every action ran in human in the loop mode. The 20 hours is the per service request baseline agreed with the client for this assessment, not a measured before state.

The work never stops: patching, provisioning, backup, incident response, security remediation, tagging, lifecycle and cost reporting. None of it is optional, and all of it competes for the same engineering hours.

This is definitely welcome. We are not able to manage the software with the people that we have without some AI help.

The client’s IT lead

Challenge

The client’s operations carried the kind of debt that manual effort never quite clears. Nothing about it was dangerous in isolation. Together it added up to risk, cost, and a permanent drag on a team with no spare capacity to absorb it.

  • Routine work was consuming the capacity needed for everything else. Patching, provisioning and incident response arrived continuously. Each request was individually small and collectively enormous, and it fell to the same engineers responsible for improving the platform.
  • The estate had never been codified. There was no infrastructure as code. Tagging standards existed as documentation rather than as enforcement. Several overlapping backup tools were in use at once. The gap between what the documentation described and what was actually running had to be closed by hand, one resource at a time, across the wider estate this pilot was meant to derisk.
  • Nothing could be handed to software on trust. This is a governance-tight organisation that has to evidence every change it makes. Any approach had to come with hard guardrails, human sign off, and a full record of what was done and why.

The ask was specific. Prove that autonomous operations could carry out real operational work safely, in an environment where a mistake would not hurt, before any conversation about production.

Solution

Firemind deployed the IT Operations Engine against a new, non-production Azure sandbox built for the pilot. The scope was deliberately narrow: around six virtual machines across two resource groups and two networks, with human in the loop on every action. Firemind ran the deployment and verified the engine’s output. The engine did the operational work, and showed its reasoning before acting.

The pilot proved three things.

  1. Proof arrived inside four and a half weeks, not eight

    The pilot was scoped for eight weeks, but the engine spent the first two and a half of them locked out. The new Azure subscription landed under a tenant whose policy blocked the connectivity the engine needed, so no test cases could run. Once access cleared, all 23 test cases were completed in the time that was left.

  2. It handled real operational work, not demonstrations

    Across nine categories of cloud operations, the engine did the job an engineer would otherwise have picked up from a queue.

  3. Safety was structural, not procedural

    Three independent layers governed what the engine was permitted to do, and every action carried a full reasoning trace. When something unexpected happened, an unavailable server size, a policy blocking access, a failed patch run, the engine re-planned and continued rather than stopping.

What the second point looked like in practice:

  • Estate-wide patching in nine minutes. Every Windows and Linux machine, honouring an explicit exclusion. When one patch failed, the engine investigated the cause across firewall rules, proxy configuration and patch distribution settings, then fixed it and verified the install.
  • A live incident resolved before a ticket existed. Responding to a CPU alert, the engine identified a runaway process and terminated it, under a policy the client had pre-agreed for that class of action. Load average fell from 115 to 25 inside a minute.
  • A tagging standard applied in twenty minutes. 29 resources across six resource types, preserving existing tags and resolving a case sensitivity conflict on the way.
  • A security gap nobody knew about, closed. Asked to block outbound traffic to a malicious address across the estate, the engine found one machine whose network interface had no security group at all. That gap was closed as part of the same request.
  • Monitoring stood up at scale. Servers provisioned complete with their monitoring stack in a single pass, and a full Zabbix deployment across six servers spanning two separate networks.

The control model is what made the pilot acceptable in an organisation that could not afford a surprise:

  1. Skills

    The client’s own process documents, including their tagging policy, are read by the engine at execution time. Operational standards are encoded without anyone writing code.

  2. Permissible actions

    Every action is risk rated. Reversible, low risk work runs on its own. Anything with data loss potential requires explicit human sign off.

  3. Access boundaries

    Contributor level, time boxed access via Privileged Identity Management forms the hard outer limit. The engine cannot act beyond what its role permits, whatever it is asked to do.

The clearest test of that model was the riskiest task in the pilot. Asked to migrate a database from SQL Server to PostgreSQL, the engine built a fourteen step plan and stopped at every step that would change the schema. Once approved, it resumed and completed the migration with type mapping, constraints and row counts all verified. It took 47 minutes, the longest of any task in the pilot, and it did not proceed a step further than it was permitted to.

Looks like you have very clear steps and you have been doing that multiple times, so nothing on my side. I’m fully trusting you that you are walking that through.

The client’s IT lead, before the database migration

Control stayed with the client throughout.

What the team can do now

The pilot met its objectives. Work that normally arrives as a service request and waits in a queue was instead handled in minutes, with a person approving anything that mattered.

23 / 23
test cases passed, zero disruptions, every action verified
8.7h
total engine time across all 23 operational tasks
23 min
average per task, the longest being 47 minutes
10
Terraform pull requests, codifying the estate for the first time

Against the baseline of 20 hours per service request agreed with the client for this assessment, that represents an estimated 451 hours of engineering effort avoided, roughly 56 working days, or a 53 times reduction. The 8.7 hours is measured. The 451 is what it compares to under that assumed baseline.

Estimated hours avoided by category, against that 20 hour baseline. Operations covers provisioning, incidents and patching, and accounts for nine of the 23 tasks.

Operations, 9 tasks
Reporting, 3 tasks
Databases, 3 tasks
Infrastructure as code, 3 tasks
Networking, 2 tasks
Software installation, 2 tasks
Human approved migration, 1 task

Beyond the figures above:

  • Engineers step back from the queue. The routine work that was consuming the team is exactly what the engine is suited to absorb, which puts senior people back on the platform improvements they were hired for.
  • Governance that existed only on paper is now enforced. A tagging standard was applied across the sandbox using the client’s own rules. The environment came under infrastructure as code from the outset, submitted as pull requests for the team to review rather than as changes made behind their back.
  • Cost decisions now have evidence behind them. A read only report identified machines running twenty four hours a day for twelve hours of actual use, worth over 50% on those machines if they were shut down when idle. Nothing was changed. The opportunities were ranked and handed over.
  • There is a safe route to widen scope. These results came from one new, non-production sandbox. The next step is the client’s own disaster recovery estate in Azure and its wider footprint of up to 500 machines, with the same engine, skills and guardrails, and humans in control at every step.

I love the tool when it says, hey, this machine and this machine are only operating 12 hours per day. Potential savings are over 50% if it’s turned off when it’s not used.

The client’s IT lead, on the cost report

The efficiency figures compare measured engine time against a baseline of 20 hours per service request, agreed with the client for this assessment. That baseline is an assumption, not an observation. Category figures are estimates derived from the same baseline.

See more case studies

  • ~24x less cloud operations effort across 36 tests, zero disruptions

    Autonomous cloud operations across a Nordic enterprise AWS estate. 36 of 36 test cases passed, zero disruptions, estate-wide changes in minutes.

    • 36 of 36 test cases passed across two phases, with zero disruptions
    • One security rule applied and verified across 21 security groups in 13 regions in 15 minutes
    • A 3-node Elasticsearch rolling patch in 21 minutes with zero data loss
    Learn more
  • 22% off the cloud bill, proved on one account before scaling.

    How autonomous cloud cost optimisation cut a Nordic firm's AWS bill by 22% in a single dev and QA account, every figure cross-verified against the live estate.

    • 22% annual cloud cost reduction, cross-verified on a single AWS dev and QA account
    • Nearly half of the saving from a single idle database
    • Continuous FinOps discipline, not a one-off audit
    Learn more
  • A decade of dormant AWS risk, triaged in two months.

    A decade of dormant AWS security risk at a Nordic digital marketing firm, triaged and resolved in a single engine run with the most urgent exposure closed under control.

    • Around 10,000 AWS security findings triaged and resolved in a single engine run, with no added headcount
    • Publicly accessible S3 bucket remediated autonomously
    • Higher-risk changes held for human approval by design
    Learn more

View all case studies

Talk to Firemind

Could your team get its time back?

See what the Firemind IT Operations Engine could do across your Azure estate.

Your benefits:

  • Outcome-driven - measurable business impact
  • Expert-led - hands-on delivery from senior practitioners
  • Secure by design - your data and compliance first
  • No lock-in - free to exit after the pilot

No obligation - just a focused 30-minute discussion about your goals.

We'll only use your details to respond to your enquiry. No newsletters unless you ask for them.