Executive Summary
Radcliffe Group, a UK-based publishing and media organisation, runs four public-facing platforms on AWS: an editorial site, a corporate site, an online learning platform for schools and colleges, and a digital publishing platform. Identity controls were inconsistent, there was no audit trail of administrative activity or threat detection, and monitoring was fragmented, so incidents surfaced through user reports.
Following AWS Well-Architected Reviews across all four platforms, Epitechnic hardened identity, data and the network edge, built a shared monitoring account with dashboards for each audience, and restricted production monitoring changes to the deployment pipeline, with the restriction tested and audited.
Unauthorised access risk fell by 75%. On the education platform, mean time to detect incidents fell from 28 minutes to 3.2 and mean time to resolve fell from 72 minutes to 24, and incidents needing a top-priority response fell by 75%.
Business Challenge
Privileged users could sign in without multi-factor authentication, so a single compromised password could expose a platform. Web entry points had no firewall against injection and bot traffic. Administrative activity was not logged, which would have blocked any incident investigation, and there was no threat detection to reveal a compromise in progress. Storage volumes and buckets were not consistently encrypted.
Monitoring relied on default metrics with no alarms, and each platform sat in its own AWS account, so there was no single view of service health. On the education platform, incidents were detected through user reports after an average of 28 minutes. Radcliffe's technology leadership had no business-level view of availability or incidents, and there were no baselines against which to recognise abnormal behaviour.
Success Criteria
- Multi-factor authentication and least-privilege access for every privileged role
- Protection at the network edge against injection and bot traffic
- A complete audit trail and continuous threat detection
- One view of health across all four platforms, with baselines and anomaly alerts
- A business view of availability and incidents for technology leadership
- Production monitoring configuration changed only through governed, reviewed deployment
- Regular, evidenced review of who holds access
Epitechnic Approach
The work followed the findings of AWS Well-Architected Reviews across all four platforms, which identified the identity gaps, the missing detective controls, and the absence of workload baselines and business-view dashboards. Epitechnic treated access to the platforms and access to monitoring data as one design question. Each group of people should see what it needs to act on, change only what its role requires, and leave an audit record of what it did.
Solution
Hardening identity, data and the edge
Multi-factor authentication was enforced for all privileged users through policy, and non-compliant accounts were remediated immediately. Access policies were reduced to least privilege, managed-key encryption was applied to all storage volumes and buckets, and database credentials now rotate automatically.
The web entry points were placed behind a content delivery network and a web application firewall with managed rule sets for injection and malicious bots. Audit logging of all administrative activity across the organisation, continuous threat detection and a central security findings hub were enabled, with high-severity threats, access policy changes and root account sign-ins raised as immediate alerts. Platform owners received runbooks for responding to threat findings and adjusting firewall rules.
One view for each audience
A shared monitoring account brings together metrics, alarm states and log metadata from the four production accounts, so the operations team can see every platform from one place. Log content, which may contain personal or application-sensitive data, stays in each production account and is queried across accounts during an investigation. The sharing policy accepts connections only from Radcliffe's four production accounts.
Four dashboards serve four audiences:
- Operations: alarm state, detection and resolution times, active incidents and AWS health events across all four platforms.
- Business: availability against target for each platform, transaction success rates and active priority incidents, for the CTO and Head of Technology.
- Security posture: security findings, threat detections, access policy changes and configuration compliance scores.
- Cost and capacity: monthly cost per platform against baseline, the largest cost drivers and idle resources.
Synthetic tests run scripted user journeys, such as student sign-in and course loading on the education platform and article publishing on the publishing platform, so degradation can be detected before users report it. In the first weeks of operation, the sign-in test found a third-party identity provider returning intermittent errors while the platform's own infrastructure reported healthy, and surfaced the problem within two minutes of onset.
Monthly reports cover availability, incidents, detection and resolution trends, and compliance. Quarterly reports cover capacity and cost, and feed right-sizing decisions.
Pipeline-only change in production
Dashboards and alarms are defined as code and changed through the deployment pipeline with peer review. In production, two independent controls enforce this: an explicit deny on the engineering role, and an organisation-level policy that blocks dashboard changes from any identity other than the pipeline. The control was tested during the build, when an engineer attempted to change a production dashboard and was refused. Any change attempted outside the pipeline is logged, raises an alert, and is investigated within one business day.
Leadership access is read-only and protected by multi-factor authentication, with an explicit deny on reading log content and on changing any monitoring configuration.
Quarterly access reviews
Each quarter, the access review confirms every assignment against Radcliffe's current contacts and the current engineering team, checks that the monitoring connections and the pipeline-only policy are unchanged, retests the policy, and confirms that no person changed a production dashboard in the previous 90 days. Findings are tracked as named actions with owners and due dates, and the summary goes to Radcliffe's Head of Technology within five business days. The most recent review found no stale stakeholder access, removed one former engineer, and confirmed that every control was operating as designed.
Outcomes
Access risk: Unauthorised access risk reduced by 75%, with the web application firewall blocking around 2,000 malicious requests a day and threat findings handled 40% faster.
Detection: Mean time to detect incidents on the education platform reduced from 28 minutes to 3.2 minutes, an 89% reduction.
Resolution: Mean time to resolve incidents on the education platform reduced from 72 minutes to 24 minutes, a 67% reduction.
Priority incidents: Unplanned incidents needing a top-priority response reduced by 75%.
Compliance: Configuration compliance across all platforms measured for the first time and held at 97.8%, above the 95% target.
Change control: Production monitoring changed only through the pipeline, with the control retested each quarter and no change from outside the pipeline in the 90 days before the most recent review.
Why It Worked
Each control had a defined audience, a defined scope and an audit record, so the design could be checked as well as built. Enforcing pipeline-only change through two independent controls, and retesting it every quarter, gave Radcliffe evidence that the control works in practice. Synthetic tests of real user journeys, the main driver of faster detection, gave the operations team an early signal of problems in services beyond its own infrastructure.
Key Takeaways
Challenge: Four content platforms with inconsistent identity controls, no audit trail or threat detection, and fragmented monitoring that left users to report incidents.
Approach: Hardening of identity, data and the edge; a shared monitoring account with dashboards for operations, security, cost and leadership; pipeline-only change enforced by two independent controls; and quarterly access reviews.
Outcomes: 75% less unauthorised access risk, detection 89% faster and resolution 67% faster on the education platform, and 75% fewer top-priority incidents.
Lessons: Access controls can be shown to work when each one is tested, logged and reviewed on a fixed cycle.
