Access Control Outage
Updates
Post-Mortem: Access Control Outage — 21 August 2026
On Friday 21st August, a fault in our release process caused a brief outage across all services, followed by a 35-minute disruption to Access Control. We know this is the second major outage our customers have seen in a short space of time. That’s not the standard we hold ourselves to, and we’ve spent the days since working through both the immediate fix and the process changes needed to stop it recurring.
What happened
The root cause was human error in our development process. A fix addressing how one of our authentication components is loaded — the component used by every customer with Azure SSO — was merged into our main branch after three builds had already begun running. Those builds shipped without it.
The affected builds were 4.930.3010.817, 4.930.3011.817 and 4.930.3012.817. Every customer who experienced this incident was running one of these three versions.
We rolled the deployed versions back immediately, which restored access across all services. That rollback placed heavy load on the API application that underpins Access Control, and Access Control dropped out as a result. This was the disruption most customers actually felt.
We then applied in-flight fixes to the broken deployments and moved every affected customer onto a corrected version. After a 5–10 minute delay, error rates fell to zero and Access Control came back online.
Timeline (AEST)
8:35 — Some customers on affected builds begin seeing intermittent login failures.
8:45 — Affected customers are rolled back to the previous versions, resolving all login failures. The system is once again stable.
9:18 — The root cause is identified, and in-flight fixes are prepared for the broken versions.
9:30 — All clients are moved onto the fixed builds with no issues identified. The system remains stable for all customers, but the API application is placed under high load.
9:34 — Intermittent Access Control issues are identified; investigation begins.
9:58 — Remediation underway, with a significant drop in error rates across API and Access Control services.
10:10 — All API and Access Control systems back online; performance fully restored.
What we’re doing about it
This failure was preventable, and the fix belongs in our process rather than our infrastructure. Our team has been working through the weekend on it. Two changes:
- Load balancing across business applications: This is the largest reliability improvement we’ve undertaken. Over the past 18 months we’ve brought high availability to our public ingress points and core routing layers; the business applications layer is the final piece. It’s in limited testing now, and it’s the change we’re most looking forward to putting in front of customers.
- Build gating: No build will be able to proceed to release unless it verifiably contains every fix merged to the main branch at the time it started. A build that starts ahead of a required change will be blocked, not shipped.
- API Load Protection: We’re hardening the API application that supports Access Control so that a mass rollback cannot cascade into a service drop-out. To this end, our development team are working on API reliability and performance from multiple angles. Access Control should hold even when a recovery action puts the platform under sudden load.
We’re aware this post-mortem reached you later than it should have. We chose to publish once we had the full picture and the remediation work underway rather than send a partial account, but we recognise the wait.
Thank you for your patience,
The PerfectGym Team
All API and Access Control systems are back online, and performance has been fully restored across other services. We will continue to monitor the platform for any irregularities while we perform a root cause analysis on the issue.
API and Access Control issues are being addressed and we have observed a large drop in error rates across all services. Performance across other applications may have slowed for customers during remediation efforts. We thank you for your patience with this.
We’re aware of an issue intermittently affecting access control for customers. Our team is investigating the cause.
← Back
PerfectGym Status