When you cannot staff the estate you already have
The client is a regulated Nordic enterprise. Its disaster recovery estate runs in Azure, a lift and shift replica of its on-premise data centre built primarily through Azure Site Recovery. Roughly 25 to 30 server instances in a single region: Windows and Linux servers, virtualised desktop workloads, integration servers, print servers, file servers and domain controllers. It sits inside a wider estate of up to 500 machines.
Scope: this pilot ran on a single non-production Azure disaster recovery estate over an eight week window, in which the engine had access for four and a half weeks. It was not a production deployment and does not cover the client’s wider estate. Every action ran in human in the loop mode. The 20 hours is the per service request baseline agreed with the client for this assessment, not a measured before state.
The work of keeping that estate running never stops. Patching. Provisioning. Backup. Incident response. Security remediation. Tagging. Lifecycle and cost reporting. None of it is optional, and all of it competes for the same engineering hours.
This is definitely welcome. We are not able to manage the software with the people that we have without some AI help.
Challenge
The estate carried the kind of operational debt that manual effort never quite clears. Nothing about it was dangerous in isolation. Together it added up to risk, cost, and a permanent drag on a team with no spare capacity to absorb it.
- Routine work was consuming the capacity needed for everything else. Patching, provisioning and incident response arrived continuously. Each request was individually small and collectively enormous, and it fell to the same engineers responsible for improving the platform.
- The estate had never been codified. There was no infrastructure as code. Tagging standards existed as documentation rather than as enforcement. Several overlapping backup tools were in use at once. The gap between what the documentation described and what was actually running had to be closed by hand, one resource at a time.
- Nothing could be handed to software on trust. This is a governance-tight organisation that has to evidence every change it makes. Any approach had to come with hard guardrails, human sign off, and a full record of what was done and why.
The ask was specific. Prove that autonomous operations could carry out real operational work safely, in an environment where a mistake would not hurt, before any conversation about production.
Solution
Firemind deployed the IT Operations Engine against the client’s non-production Azure disaster recovery estate. The scope was deliberately narrow: a single region, roughly 25 to 30 servers, and human in the loop on every action. Firemind ran the deployment and verified the engine’s output. The engine did the operational work, and showed its reasoning before acting.
The pilot proved three things.
Proof arrived inside four and a half weeks, not eight
The pilot was scoped for eight weeks, but the engine spent the first two and a half of them locked out. The new Azure subscription landed under a tenant whose policy blocked the connectivity the engine needed, so no test cases could run. Once access cleared, all 23 test cases were completed in the time that was left.
It handled real operational work, not demonstrations
Across nine categories of cloud operations, the engine did the job an engineer would otherwise have picked up from a queue.
Safety was structural, not procedural
Three independent layers governed what the engine was permitted to do, and every action carried a full reasoning trace. When something unexpected happened, an unavailable server size, a policy blocking access, a failed patch run, the engine re-planned and continued rather than stopping.
What the second point looked like in practice:
- Estate-wide patching in nine minutes. Every Windows and Linux machine, honouring an explicit exclusion. When one patch failed, the engine investigated the cause across firewall rules, proxy configuration and patch distribution settings, then fixed it and verified the install.
- A live incident resolved before a ticket existed. Responding to a CPU alert, the engine identified a runaway process and terminated it, under a policy the client had pre-agreed for that class of action. Load average fell from 115 to 25 inside a minute.
- A tagging standard applied in twenty minutes. 29 resources across six resource types, preserving existing tags and resolving a case sensitivity conflict on the way.
- A security gap nobody knew about, closed. Asked to block outbound traffic to a malicious address across the estate, the engine found one machine whose network interface had no security group at all. That gap was closed as part of the same request.
- Monitoring stood up at scale. Servers provisioned complete with their monitoring stack in a single pass, and a full Zabbix deployment across six servers spanning two separate networks.
The control model is what made the pilot acceptable in an organisation that could not afford a surprise:
Skills
The client’s own process documents, including their tagging policy, are read by the engine at execution time. Operational standards are encoded without anyone writing code.
Permissible actions
Every action is risk rated. Reversible, low risk work runs on its own. Anything with data loss potential requires explicit human sign off.
Access boundaries
Contributor level, time boxed access via Privileged Identity Management forms the hard outer limit. The engine cannot act beyond what its role permits, whatever it is asked to do.
The clearest test of that model was the riskiest task in the pilot. Asked to migrate a database from SQL Server to PostgreSQL, the engine built a fourteen step plan and stopped at every step that would change the schema. Once approved, it resumed and completed the migration with type mapping, constraints and row counts all verified. It took 47 minutes, the longest of any task in the pilot, and it did not proceed a step further than it was permitted to.
Looks like you have very clear steps and you have been doing that multiple times, so nothing on my side. I’m fully trusting you that you are walking that through.
Control stayed with the client throughout.
What the team can do now
The pilot met its objectives. Work that normally arrives as a service request and waits in a queue was instead handled in minutes, with a person approving anything that mattered.
Against the baseline of 20 hours per service request agreed with the client for this assessment, that represents an estimated 451 hours of engineering effort avoided, roughly 56 working days, or a 53 times reduction. The 8.7 hours is measured. The 451 is what it compares to under that assumed baseline.
Estimated hours avoided by category, against that 20 hour baseline. Operations covers provisioning, incidents and patching, and accounts for nine of the 23 tasks.
Beyond the figures above:
- Engineers step back from the queue. The routine work that was consuming the team is exactly what the engine is suited to absorb, which puts senior people back on the platform improvements they were hired for.
- Governance that existed only on paper is now enforced. A tagging standard was applied across the estate using the client’s own rules. The disaster recovery estate came under infrastructure as code for the first time, submitted as pull requests for the team to review rather than as changes made behind their back.
- Cost decisions now have evidence behind them. A read only report identified machines running twenty four hours a day for twelve hours of actual use, worth over 50% on those machines if they were shut down when idle. Nothing was changed. The opportunities were ranked and handed over.
- There is a safe route to widen scope. These results came from one non-production estate. The same engine, skills and guardrails apply as more of the client’s 500 machines move to Azure, with humans in control at every step.
I love the tool when it says, hey, this machine and this machine are only operating 12 hours per day. Potential savings are over 50% if it’s turned off when it’s not used.
The efficiency figures compare measured engine time against a baseline of 20 hours per service request, agreed with the client for this assessment. That baseline is an assumption, not an observation. Category figures are estimates derived from the same baseline.



