
Every few months, a major tech outage makes headlines. Systems go down, users are locked out, services grind to a halt, and the post-incident write-ups that follow tend to reveal something uncomfortable: the failure was not entirely unpredictable. Somewhere in the architecture, there was a gap. A configuration that had not been reviewed. An assumption that turned out to be wrong. A single point of failure that nobody had stress-tested in a while.
What is striking is not the outages themselves, but how consistent the underlying causes tend to be. Whether it is a cloud provider going dark or a critical API failing at the worst possible time, the pattern repeats. The gap existed before the incident. The question was always when, not if.
When businesses think about cybersecurity or infrastructure risk, they tend to focus on the obvious targets: firewalls, passwords, antivirus tools. These are important, of course, but they are also the areas that typically receive the most attention. The gaps that lead to real outages are often quieter, sitting in legacy integrations nobody has touched in years, in third-party dependencies that were reviewed once at onboarding and never again, or in processes that work perfectly well until they meet an edge case at scale.
The 2021 Facebook outage is a good example. A BGP (Border Gateway Protocol) configuration change during routine maintenance accidentally removed Facebook's IP address blocks from the global routing table. The result was a six-hour outage affecting roughly 3.5 billion users. The technical cause was specific, but the broader lesson was not: a routine operation, conducted without sufficient safeguards, created a catastrophic chain reaction.
This is exactly why vulnerability testing services exist; not to confirm that your obvious defences are working, but to probe the spaces in between. The integrations. The dependencies. The configurations that have drifted quietly over time.
One of the most common hidden gaps is third-party risk. Businesses rely on an ever-growing number of external vendors, platforms, and service providers. Each one represents a potential point of failure, not just from a cyberattack, but from an outage on their end that cascades into yours.
The 2021 Fastly outage lasted under an hour but took down significant portions of the internet with it, including major news sites, government portals, and e-commerce platforms. Most of the affected organisations had done nothing wrong. Their gap was an over-reliance on a single CDN provider without adequate fallback options.
Reviewing data breach case studies, such as the lessons from the Brightspeed breach, reinforces how frequently external vendors and third-party access points serve as the path of least resistance for both attackers and unexpected failures.
The fix is not to eliminate third-party dependencies, which is unrealistic. It is to map them properly, understand where a single provider going down would be catastrophic, and build redundancies accordingly.
It is tempting to frame tech outages as purely technical failures, but the human element is almost always somewhere in the story. A misconfigured setting or an untested update pushed to production on a Friday afternoon is the moment that technical teams dread and that post-mortems consistently surface.
The 2017 AWS S3 outage, which disrupted a significant portion of the internet, was traced back to a team member who had entered an incorrect input during a debugging exercise. The command removed more servers than intended. The system had not been tested at that failure scenario, so the recovery took far longer than it should have.
There is no technology that fully eliminates human error, but there are processes that reduce its blast radius: proper change management, mandatory peer review for high-risk configurations, staged rollouts, and regular drills that test how teams respond under pressure, not how they think they would.
Most organisations that have experienced a significant outage will conduct some form of post-incident review. The challenge is that these reviews tend to focus on what went wrong in the incident itself, rather than why the conditions for that incident existed in the first place.
A more useful question is: what would have caught this before it became an incident? Often, the answer involves testing and visibility that simply were not in place. Unmonitored systems. Untested recovery procedures. Dependencies that were assumed to be stable but never verified. Configuration drift that nobody noticed because nobody was looking.
Building a habit of proactive review is what separates organisations that learn from outages to organisations that repeat them.
The businesses that weather tech outages best are rarely the ones with the most sophisticated technology. They are the ones that have built a culture of questioning their own assumptions. They test their systems regularly, treat third-party dependencies as potential liabilities, and invest in visibility across their entire environment, not just the parts that feel most exposed.
That mindset extends to cybersecurity as much as it does to infrastructure. Gaps do not announce themselves. They sit quietly, waiting for the right combination of circumstances to surface. By the time an outage reveals them, the cost is already being counted.
If your organisation has not recently examined where its hidden gaps might be, now is the right time to start. Group8 specialises in helping businesses identify vulnerabilities before they become incidents. Reach out to the team at group8.co to find out how a proactive assessment could protect what you have built.