Say you completely ignore scaling. The two things you simply cannot replicate at that scale are redundancy and operational resource. AWS has their entire operations team working at all hours of the day and night supporting their infrastructure. They also offer some of the most highly redundant services in the world. There is simply no way you could ever dream of replicating those service levels with such a small operation, and if you were to even attempt it, it would require an absurd level of over provisioning. As I said, you’re completely misrepresenting what the actual trade offs are, and there’s no possible way your claims about replicating AWS service levels is even remotely plausible.
> AWS has their entire operations team working at all hours of the day and night supporting their infrastructure. They also offer some of the most highly redundant services in the world.
And yet, a couple times a year perhaps, we have discussions right here on HN about the latest AWS outage that took down half the Internet.
No group is infallible. If I thought about a world where cloud providers didn't exist (AWS or otherwise), where every company had to build and maintain all of their infrastructure themselves, and had to make a guess, I'd wager the combined occurrences of issues around availability, durability, etc. would far outpace what we have had.
That's not even considering the potential impact to software development and innovation that we get with commodity cloud services. This is hand-wavy of course but I'd stick to it.
You mean the incident where a small percentage of EC2 instances were unavailable for 30 minutes in a single AZ in US East 1? I see your definition of major incident is pretty loose. I remember that incident. I had services running there. It was so minor that my auto-scaling picked it up and my service impact was nothing.
AWS has amazing marketing the truth of the matter is AWS Region has worse downtime than a single top tier DC. Mainly due to nightmarish complexity of their control layer. They had outages that lasted many hours in a row multiple times. You need to carefully separate marketing claims from operational reality and actual track record. When US East has major issues there is not enough spare capacity to spin up everything that was running there in other regions.
with high availability you don’t wait for an outage to spin up new resources, at that point it’s too late. it’s by definition not highly available and if you build infrastructure this way then you can’t blame AWS for an outage
So you have say 3 Region deployment are you saying that you are running 50% more instances than you need in 2 regions that are not US East to make sure you will have capacity when US East goes down :) ? I somehow seriously doubt that.
so you’re suggesting entire regions go down at once or one of the AZs? An entire region doesn’t go down. So again, you are not building for high availability.
i don’t remember seeing us-east go down in its entirely. show me supporting evidence or this is FUD. multiple DCs physically separated, different flood planes make up a region. it’s not easy to down an entire region. the biggest event they had, the S3 one you’re talking about, effected only 2 AZs and didn’t allow new EC2 instances to start and iirc some EC2 instances failed as well. this is a far cry from the entirely of us-east having an outage.
Amazon has some of the best uptime in the world, especially for basic services like EC2, even if you’re only considering the least reliable regions like US East. Their last major event was in 2017. There are few providers in the world that can compete with them in that respect, and there’s nothing that you or I or anybody else could build with 6-10 servers that could come close. If you were planning to try exceed their service levels yourself, there is no conceivable use case where a small to medium sized company could justify the requisite expenses to provide the redundancy and operational coverage necessary. What you are actually talking about is that you can meet your own needs without AWS, which is entirely plausible, but completely different from the absurd claims that you can build a low budget infrastructure that exceeds their service levels. That claim is so ridiculous, you might as well be saying that you can make a car faster than Toyota can, or that you can run a two minute mile.
Our AWS monthly spend is deep 7 figures over the last 10 years our colo projects had better uptime than AWS US East. You keep living in the marketing bubble for AWS.
There’s a few confounding factors to address before you get to the believability of the claims (which are remarkably dubious). For starters they’ve only spoken about the AWS region which is least reliable by design (as it is the first to receive new features), and they’re talking about a 10 year time span (AWS in 2010 was much less reliable than it is today). It’s also not really clear what they’re talking about when they say AWS. If it’s just the core features you’d need to forklift an on-prem service into the cloud, then the claims are especially frivolous, but the SLA difference between EC2 and AWS Ground Station is more than an order of magnitude.
Even if their claims are true (which I certainly don’t believe they are), you’d be more likely to get better uptime than EC2 with a small on-prem setup through dumb luck rather than through deliberate planning. Something still has to go wrong for you to have an outage, and you’re more likely to get an incredible lucky streak than you are to outperform their entire AWS infrastructure capability with a few people and half a rack of servers.
There’s a terrible amount more information you’re missing here. What services were you using that went down? (not what services went down, how were you affected) What is your availability for your DC? it seems you’re being light in the details and perhaps there’s a reason why.