When the dashboards say everything is fine
The client is a large, multinational services group running a substantial AWS footprint. Roughly 30 accounts, thousands of servers, hundreds of databases.
Their own ticketing data told a story before Firemind ever touched their cloud estate. Ticket volume across the group had roughly tripled in eighteen months. Nearly all of it arrived automatically, got routed automatically, and was flagged as the lowest priority. In the slice of that queue tied to this estate, only around 1 in 14 incidents ended in a recorded fix. The rest were opened, checked, and closed with nothing done, or closed with no record of what was done at all.
A previous review of the same estate had recommended deleting old backup snapshots as a cost saving. That recommendation was later withdrawn. In several accounts, those snapshots turned out to be the only protection that existed at all.
Challenge
The client needed an honest, evidenced picture of what their cloud estate actually cost and where it was at risk, before committing to anything bigger. Not another set of recommendations taken at face value.
- A previous cost-cutting call had backfired. The snapshot-deletion episode meant any future recommendation needed to survive real scrutiny, not just look plausible on a slide.
- Ticket volume had roughly tripled, with little to show for it. Almost all of it was automatic, low priority, and went nowhere. That is not a healthy signal on its own, but it also could not be trusted as a complete picture of what was actually wrong.
- Nobody could say with confidence what anything cost, or whether it was protected. Cost, ownership and backup status all depended on estate-wide discipline that had never been enforced.
The ask was a fully read-only assessment. Evidence first, decisions after, nothing touched until the client chose to act.
Solution
Firemind ran a read-only assessment across roughly 30 AWS accounts: cost, architecture, resource lifecycle, patch management, a dedicated tag audit, and a comparison against the client’s own ticketing history. The IT Operations Engine did the investigative work directly, not just producing a report but actively finding things, and in at least one case, choosing not to claim a finding as a saving because the evidence did not yet support it.
The assessment proved three things.
What the estate actually costs, evidenced rather than estimated
Three findings account for most of the identified savings: reclaiming close to 750 disks left detached from anything and still billed monthly, decommissioning servers switched off for years while still holding storage, and correcting a cloud pricing commitment that had lapsed. Together those three alone add up to close to $140,000 a year.
Findings no dashboard in place had ever shown as a problem
Backup jobs that ran on schedule, reported success, and protected nothing, because the vault they wrote to was empty. In one account, 92% of servers were correctly enrolled for patching, and thirteen patch schedules were configured and switched on. Not one had ever actually run: every schedule was looking for a label no server in the account carried. Fixing that single label brought patch visibility back across the fleet. A tagging standard existed, was well designed, and was met by almost nothing in the estate, which is the direct cause of the patch gap: without the right label, the patching infrastructure has nothing to attach to.
Judgement, not automatic savings-grabbing
In a separate account, the engine found close to $900 a month in further disused disks. It did not claim them. They were less than a week old and traced to an automated process that was still running, so it flagged the process for investigation instead. The client’s own cost-anomaly tooling had not caught this at all.
Alongside the cost and tagging work, the assessment also surfaced security exposures worth acting on regardless of cost: a production database reachable from the public internet, unencrypted, sitting alongside credentials stored in plain text; one environment that returned 62 critical findings in a single pass; and access hygiene materially behind good practice, including root accounts with no multi-factor authentication at all.
Nothing in the environment was changed. Testing remediation is the explicit subject of the next phase, and it runs entirely inside a purpose-built, anonymised replica of the estate. Anything copied into it is scrambled first, so the engine works with the shape of the estate and its defects, never real data, and never a live environment. One of the measures already agreed for that phase is simply how much of the remediation work the engine can complete on its own, against how much needs a person to approve it first.
What the assessment gives the client now
The client has an evidenced, prioritised business case instead of a set of assumptions. Every figure in it traces to a named finding, not a general recommendation.
The three largest identified savings, against that ~$350,000 a year total:
Beyond the figures above:
- Three findings no existing tool had ever flagged. Not a gap in effort. A gap in what the tools already in place were configured to look for.
- A tagging standard that existed only on paper now has a path to enforcement. The same missing-label pattern connects the patch gap, the cost-attribution gap, and the estate’s inability to reliably tell production from everything else.
- A safe, evidenced route to the next phase. Remediation gets tested inside an anonymised replica of the estate before anything touches production, with a person approving anything that matters.
Scope: this reflects a single, fully read-only assessment across roughly 30 AWS accounts. No changes were made to the client’s environment at any point. The savings figures are identified opportunities, evidenced against named findings, not delivered results. Testing remediation is the explicit subject of the next phase, and it is confined to an anonymised replica of the estate, not production.



