AWS Resolves Major Outage After Nearly 24 Hours of Service Disruption
Amazon Web Services (AWS) experienced a significant service disruption in its US-EAST-1 region, affecting over 140 services and causing issues globally. The outage lasted nearly 24 hours, beginning late on Thu, Oct 19, 2025, and was resolved by the…
Amazon Web Services (AWS) experienced a significant service disruption in its US-EAST-1 region, affecting over 140 services and causing issues globally. The outage lasted nearly 24 hours, beginning late on Thu, Oct 19, 2025, and was resolved by the afternoon of Fri, Oct 20, 2025.
The incident began at approximately 11:49 PM PDT on Oct 19, when AWS engineers detected increased error rates and latencies across multiple services in the US-EAST-1 region. At 12:26 AM on Oct 20, AWS identified the root cause as DNS resolution issues affecting regional DynamoDB service endpoints.
This initial problem led to a cascading failure impacting numerous other services. After resolving the DNS issue in DynamoDB at 2:24 AM, AWS faced further impairments in EC2’s internal subsystem, dependent on DynamoDB, which affected the launching of new instances.
The situation escalated when Network Load Balancer health checks were impaired, resulting in network connectivity problems across services like Lambda, DynamoDB, and CloudWatch. AWS temporarily throttled several operations, including EC2 instance launches, SQS queue processing via Lambda Event Source Mappings, and asynchronous Lambda invocations to manage the recovery process.
Amazon Web Services (AWS) experienced a significant service disruption in its US-EAST-1 region, affecting over 140 services and causing issues globally.
Engineers restored Network Load Balancer health checks by 9:38 AM PDT. Throughout the day, AWS gradually reduced throttling while addressing network connectivity issues. By 3:01 PM PDT on Oct 20, all AWS services resumed normal operations, although some services continued processing backlogs for several hours.
The outage particularly impacted global services relying on US-EAST-1 endpoints, including IAM authentication and DynamoDB Global Tables. Customers experienced EC2 instance launch failures, Lambda function invocation errors, and difficulties accessing storage and database services. The disruption also hindered the creation or updating of support cases during the peak of the incident.
AWS plans to share a detailed post-event summary to provide a comprehensive understanding of the incident and to outline measures to prevent similar occurrences. The company advises customers to configure Auto Scaling Groups across multiple Availability Zones and avoid targeting specific zones during instance launches to enhance resilience against regional issues.
Based on reporting by GBHackers.
