Microsoft 365 Outage: A Bug in the System
On July 23, a massive outage of Microsoft 365 services occurred, affecting thousands of customers. The outage began at 10:44 AM ET and was caused by a bug in Microsoft's automated network maintenance request system.
The bug mistakenly removed IP routes from more devices than intended, disrupting Azure and Microsoft 365 services. The outage mostly affected customers accessing Microsoft 365 services through network infrastructure connected to Microsoft’s West US Azure region.
Services Affected by the Outage
- Microsoft OneDrive: Access to OneDrive was intermittent.
- SharePoint Online: Users received 'Something went wrong' errors.
- Microsoft Teams: Chat functionality was degraded, including images not loading.
- Microsoft 365 Admin Center: The Admin Center loaded slowly or not at all.
- Power Automate: Automate flows did not load.
- Copilot Chat: Users experienced intermittent delays or failures when performing actions and queries.
- Microsoft Loop: Users were unable to open or load Loop pages.
- Fabric and Power BI, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender were also affected.
Some Defender customers experienced delays receiving responses from Microsoft Defender Experts, while investigations, workflows, and remediation actions triggered through Threat Explorer and Advanced Hunting could fail.
Initial Response and Resolution
Microsoft initially attempted to mitigate the outage by rerouting traffic through alternate network paths, which helped customers, but many services continued to be affected. Before determining what caused the outage, Microsoft warned customers that they might need to review their business continuity and disaster recovery plans and take actions appropriate for their environments.
The company later identified a recent networking change as the cause and began reverting it. Microsoft completed the reversion at 2:26 PM ET and confirmed through service telemetry and customer reports that the Microsoft 365 incident had been resolved.
Maintenance Bug: The Root Cause
In a preliminary Post Incident Review for the Azure incident, Microsoft said the outage was triggered during routine device maintenance in its West US Azure region, where specific network paths were being isolated. Microsoft's maintenance process converts these types of requests into system-readable instructions and checks that at least one of two redundant paths remains healthy before the work begins.
However, a bug in the request conversion system incorrectly marked additional network devices as part of the maintenance event. As a result, IP routes were removed from more devices than intended between Microsoft's West US datacenter and its wide-area network.
The removed routes disrupted network traffic entering or leaving the West US region. However, Microsoft said traffic remaining entirely within the region was not affected.
Azure Incident: Widespread Impact
The Azure incident caused connectivity failures, increased latency, and problems accessing numerous cloud services, including Azure App Service, Application Gateway, Azure AD B2C, Azure AI Search, Azure API Management, Azure Cosmos DB, Azure Databricks, Azure Firewall, Azure Kubernetes Service, Azure Monitor, Azure Virtual Desktop, ExpressRoute, Log Analytics, Microsoft Graph, Microsoft Sentinel, Power BI Embedded, Virtual WAN, and VPN Gateway.
Microsoft said its engineers began investigating the issues immediately after the outage began at 10:44 AM ET. The problem initially presented itself as large-scale route churn in Microsoft's WAN. Engineers later traced the route removals to a datacenter in the West US region and linked them with the recent maintenance activity.
Microsoft initiated a rollback of the maintenance change at 1:45 PM ET, which was completed at 2:26 PM ET. The rollback restored the affected network infrastructure and allowed Microsoft 365 services to recover.
Post-Incident Review and Next Steps
Microsoft is now conducting a full internal review focused on the safety checks and automated processes used to execute maintenance requests. The company said it will publish a final Post Incident Review after completing its investigation, which is usually within 14 days.
Microsoft will be performing a full analysis focusing on safety checks, automated maintenance request change process, and more as they progress through their post-mitigation internal retrospective.
Test every layer before attackers do. Security teams log 54% of successful attacks and alert on just 14%. The rest move through your environment unseen.
Source: BleepingComputer