Executive Summary
Radcliffe Group is a UK-based publishing and media organisation serving millions of readers through four public-facing platforms on AWS: an editorial site, a corporate site for investor relations, an online learning platform for schools and colleges, and a digital publishing platform. The platforms were growing quickly, and each had a different tolerance for downtime.
Epitechnic carried out a Well-Architected remediation across the production estate. It began with a business impact analysis that set recovery targets for each platform, then identified and treated the failure risks that stood between each platform and its target, and replaced manual changes with automated, reversible deployments.
All four platforms reached 99.9% uptime after remediation, incident-related downtime fell by 60%, and the learning platform's availability rose from an estimated 99.4% to 99.96%.
Business Challenge
The four platforms shared an AWS account structure but carried very different consequences when they failed. The education platform serves school and college students on subscription and supports exam and assignment submission windows, so an outage during an active exam period carries reputational and contractual risk. The publishing platform has critical windows around publication release dates. The corporate and editorial sites matter most around investor reporting and campaign launches.
The estate did not reflect those differences. Application servers, and the education platform's database, each ran in a single availability zone with no automatic failover. Backups had never been restored, so restore time and data integrity were unknown. There were no alarms, and incidents were reported by users before the team was aware of them.
Changes were applied by hand through the AWS console, with no version control and no way to roll back. Failed deployments had caused two extended outages in the six months before the engagement. AWS maintenance notices went to an account email address the team did not monitor, and on one occasion a scheduled database maintenance event restarted a platform during business hours.
Success Criteria
- Recovery targets for each platform, set by the business impact of its downtime
- Every material failure risk identified, scored and treated
- Platform resilience matched to each platform's recovery target
- Infrastructure and deployments that are version-controlled, reviewed and reversible
- Protection for the education platform through exam periods
- Incidents detected by monitoring before users report them
Epitechnic Approach
Epitechnic started with a business impact analysis across the four platforms, assessing the financial, operational and reputational impact of downtime for each. The analysis placed the platforms in three tiers and set a recovery time objective (how quickly a platform must be back) and a recovery point objective (how much recent data the business can afford to lose) for each.
Tier 1, education platform: recovery within 4 hours, with no more than 1 hour of data at risk.
Tier 2, publishing platform: recovery within 8 hours, with no more than 4 hours of data at risk.
Tier 3, corporate and editorial sites: recovery within 24 hours, with no more than 8 hours of data at risk.
With the targets agreed, every resilience decision could be tested against them. Epitechnic then ran a structured failure mode analysis across the estate, scoring each failure mode for severity, likelihood and ease of detection, and requiring a treatment action for every risk above the agreed threshold.
Solution
A risk register built from failure modes
The analysis identified ten material risks: single-zone application servers, a single-zone database on the education platform, no web application firewall, multi-factor authentication not enforced for privileged users, no audit logging of administrative activity, untested backups, no monitoring alerts, manual deployments with no rollback, unencrypted storage, and no threat detection.
The missing monitoring and threat detection scored highest and were treated first. The education platform's single-zone database was treated as a priority because of the exam-period risk. When the register was reassessed after remediation, no risk remained above the threshold.
Resilience built to each tier
The production network and application servers were rebuilt across two availability zones, with load balancer health checks replacing failed servers automatically. The education platform's database moved to a multi-zone cluster with automatic failover, tested against its recovery target. Quarterly restore tests now record actual restore times against each platform's recovery point objective, so recovery is a measured capability.
For exam periods, Epitechnic agreed an operating procedure with the platform team. The education platform is placed under a change freeze and held at pre-scaled capacity throughout the window, so it is ready for peak demand before students sign in.
Automated, reversible deployment
All infrastructure is now defined as code, stored in version control and deployed through a pipeline with peer review, automated validation and security scanning. Production releases use a blue/green cutover, in which the new version runs alongside the old. If error rates rise during a five-minute observation window after cutover, the pipeline restores the previous version automatically.
Awareness of provider events
AWS service health events, including scheduled maintenance, are now routed automatically to the platform team, and events affecting the two most critical platforms are raised as alerts. Epitechnic also briefed the platform team on the shared responsibility model, setting out which aspects of availability AWS provides and which remain Radcliffe's own, such as backup restore testing, disaster recovery activation, application retry logic and monitoring thresholds.
Capacity and cost
Traffic spikes had previously degraded performance. With automatic scaling behind the load balancers and edge caching for content, the platforms sustained 50% more concurrent users with no material increase in latency. Budgets, tagging standards and monthly optimisation reviews with platform owners reduced monthly AWS costs by 25%, and together with USD 20,000 in AWS credits this kept the remediation cost-neutral.
Handover to the platform team
Radcliffe's platform owners receive monthly cost and security reports, and the operations team was coached on change windows, blue/green cutovers and rollback, enabling it to run the platforms without Epitechnic on the critical path.
Outcomes
Availability: 99.9% uptime across all four platforms after remediation, with the education platform's availability rising from an estimated 99.4% to 99.96%.
Downtime: Incident-related downtime reduced by 60% across the estate.
Risk: All ten material failure risks treated, with none above the threshold on reassessment.
Change: Every infrastructure change version-controlled and peer-reviewed, with automatic rollback for failed releases.
Scale and cost: 50% more concurrent users sustained with no material latency increase, and monthly AWS costs reduced by 25%, with USD 20,000 in AWS credits keeping the remediation cost-neutral.
Why It Worked
Setting recovery targets from business impact first gave every later decision a measure to test against, so effort and cost went where downtime would do the most harm, with the education platform during exam periods at the top. Scoring failure modes made the order of remediation explicit and showed when the work was complete. Automated, reversible deployment removed the manual change process that had caused two extended outages in the previous six months.
Key Takeaways
Challenge: Four content platforms with different tolerances for downtime, running on a single-zone estate with untested backups, no alerting, and manual changes that had caused extended outages.
Approach: A business impact analysis that set recovery targets for each tier, a scored failure-mode risk register with every material risk treated, and infrastructure and deployment automated with rollback.
Outcomes: 99.9% uptime across all four platforms, 60% less incident-related downtime, and a cost-neutral remediation.
Lessons: Resilience spending can stay in proportion when each platform's recovery target is set by the impact of its downtime before any design work begins.
