The cloud bill audit: 34% recovered in one week, monitoring added the next
Fintech back-office platform. A one-week cost audit plus structural fixes — right-sizing, lifecycle rules, transfer routing and budget alerts — recovering a third of spend, with monitoring and backup verification added immediately after.
At a glance
- 150+
- Clients served
- 1,550+
- Projects delivered
- 14
- Industries
- 9+
- Years in business
Timeline
4 weeks
Team
3 people
Stack
AWS · Terraform · Grafana · Node.js
The world this system had to work in
Cloud bills grow the way unused subscriptions do: slowly, invisibly, and with a plausible explanation for every line. For a SaaS platform without a named owner for infrastructure, forgotten environments, oversized instances and cross-region traffic can double the bill in a year without a single bad decision being made on purpose.
The challenge
A growing SaaS platform's cloud costs had doubled in a year. Nobody owned the bill: forgotten staging environments, oversized instances and cross-region chatter inflated it monthly.
What we found on day one
- Monthly cloud spend had doubled in twelve months and nobody owned the bill
- Staging environments from finished projects were still running at full size
- Instances were sized for a traffic spike that never repeated; average utilisation was under 20%
- Backups existed but had never been restore-tested; monitoring was limited to a health-check ping
The solution
We ran a one-week cost audit — inventorying every resource against an owner, checking 30-day utilisation and tracing data-transfer charges — then applied structural fixes: shut down orphaned environments, right-sized instances, storage lifecycle policies and rerouted cross-region service traffic. In the following weeks we codified the setup in Terraform, enforced tagging in CI, added budget alerts and built Grafana dashboards with on-call routing, and restore-tested every backup.
Cost audit & inventory
Every resource mapped to an owner and purpose; orphaned assets flagged and retired.
Right-sizing & lifecycle
Instances resized to real utilisation; hot, cool and archive tiers applied to storage.
Traffic re-routing
Cross-region service calls consolidated, removing the single largest fixable line item.
Infrastructure as code
Existing environment codified in Terraform incrementally, without downtime.
Monitoring & alerting
Grafana dashboards, latency and error-rate alerts routed to a named on-call engineer.
Backup verification
Scheduled restore drills with documented recovery times for every data store.
Integrated with
Our approach, step by step
Inventoried every resource with an owner; flagged the unowned ones and shut down the forgotten environments.
Right-sized instances against 30-day utilisation data and applied storage lifecycle policies.
Rerouted cross-region service traffic, the single largest fixable line item.
Made savings structural: enforced tagging in CI, budget alerts at 80%, and Grafana dashboards with on-call alerts.
The outcome
- Monthly cloud spend fell 34% within one week of the audit, before any architectural change.
- Every backup is now restore-verified on a schedule; recovery time is documented rather than assumed.
- Tagging enforced in CI and budget alerts at 80% keep the savings structural rather than a one-off cleanup.
- The recovered budget funded the monitoring and on-call setup the platform had been missing.
“They migrated our platform to the cloud and set up CI/CD without a single weekend of downtime. Deploys that used to be a fire drill are now a non-event.”
Priya Nair
CTO, SaaS platform
More engagements
All 12 case studiesFacing a similar challenge in saas / devops?
Book a free scoping call — a written rough estimate in 3 days, a fixed proposal in 7.
Book a scoping call with a software architect — not a sales bot.
You'll get a reply within one business day. We'll send a rough estimate in 3 days and a fixed proposal in 7.
Prefer to talk first? Phone, email and office address are on the contact page.
