GitHub Outage Blamed on Rapid Growth Outstripping Capacity

0
68

GitHub has revealed the root cause behind its extensive outage on August 17th, attributing the widespread service disruption to its infrastructure capacity failing to keep pace with the platform’s rapidly escalating usage. The incident, which affected GitHub’s website, authentication, GitHub Actions, API, Pull Requests, and Issues, lasted a total of 7 hours and 47 minutes.

Infrastructure Strain Leads to Downtime

According to GitHub’s investigation, the massive outage was not triggered by any code or configuration changes. Instead, a historical peak in platform traffic on August 17th led to a capacity shortage in a critical infrastructure component within its US Central data center. This initial bottleneck cascaded, creating pressure on other systems and resulting in authentication failures and a broad spectrum of service malfunctions.

Adding to the complexity of the recovery, a bug in the restoration process caused some services to incorrectly trigger client retry mechanisms. This inadvertently amplified system traffic, prolonging the time it took for technical teams to fully restore all platform services.

Recurring Capacity Issues

This August 17th incident marks the second significant service disruption for GitHub this month. The platform experienced a separate outage affecting GitHub Actions on August 6th. GitHub has emphasized that both major incidents stemmed from insufficient infrastructure capacity, rather than issues with software code or configurations.

Explosive Growth on the Platform

The company highlighted a dramatic surge in platform usage recently. Since April, the monthly number of commits has nearly doubled, climbing from 1.4 billion to 2.9 billion. Concurrently, the volume of merged Pull Requests and newly created code repositories has also seen sustained growth.

Mitigation Efforts and Future Plans

In response to this escalating demand, GitHub has significantly bolstered its infrastructure this year. This includes the addition of over 3 million CPU cores, 120PB of high-speed storage, and enhanced network capacity. Furthermore, the company has accelerated the migration of workloads to Microsoft Azure. Currently, approximately 58% of GitHub’s platform load is handled by Azure, a substantial increase from the 12% in May. Roughly half of all Git operations are also now processed by Azure.

Despite these substantial optimizations, the platform still encountered a situation where user volume overwhelmed its services. Looking ahead, GitHub plans to implement further measures to enhance stability. These include isolating critical platform systems, reducing shared dependencies between different services, and establishing uniform retry limits and timeout mechanisms for inter-service communications. These steps aim to prevent a recurrence of cascading failures caused by excessive automated retries during system anomalies.

Source: https://www.ithome.com/0/992/844.htm

LEAVE A REPLY

Please enter your comment!
Please enter your name here