Most instances get sized once, on launch day, by someone making a guess, and nobody ever looks at them again.

Nobody has ever been paged because an instance was too big.

Think about that for a second, because it explains almost everything about why oversizing is so common. Undersize something and you find out fast — latency goes up, pods get OOM-killed, somebody's phone buzzes at 2 a.m. Oversize it and nothing happens. It just sits there at 12% CPU, month after month, and the only evidence is a number on the bill that everyone has stopped questioning.

So when launching something new and unsure whether it needs a 2xlarge or a 4xlarge, the instinct is to pick the bigger one. Going too big costs nothing anyone will personally notice. Going too small might cost a weekend. Then the workload changes, the person who sized it moves teams, and that guess quietly becomes permanent.

Fixing this isn't hard. It mostly comes down to not trusting the obvious numbers.

Step 1: Don't trust the CPU average.

Average CPU is the first thing everyone checks, and it lies more than any other metric. An instance can average 15% and still hit 85% every afternoon when the reports run. That one isn't oversized — it's sized for its peak, which is exactly what you want. Look at p95 or p99 instead. If the p99 over a month is sitting around 30%, that's a real candidate.

Memory is the bigger blind spot. EC2 doesn't send memory utilization to CloudWatch unless the CloudWatch agent is installed, and a lot of accounts never did. That means Compute Optimizer is often making recommendations with half the picture, and memory-heavy workloads are exactly the ones that break if you downsize on CPU data alone.

Give it enough history, too. Two weeks of data won't show a month-end batch job. If the workload has a monthly rhythm, wait for a full month before trusting any recommendation.

Step 2: Sort by money, not by percentage.

Pull up any rightsizing report and there will be dozens of instances sitting at 3-5% utilization. It looks alarming. Most of it doesn't matter — a t3.small doing nothing costs about what you'd spend on coffee in a month.

Sort by monthly cost first, then look at utilization. One r6i.8xlarge running at 25% memory is worth more attention than fifty idle micros. In most accounts, the top 10 or 20 instances by spend are where nearly all the real savings are. Start there and ignore the long tail for now.

Step 3: Change one thing at a time.

There are really three ways an instance can be wrong: the size, the generation, or the family — worth fixing them in that order, since that's roughly the order of how likely each is to surprise you.

Size is the safe one. Dropping an m6i.2xlarge to an m6i.xlarge keeps everything the same except capacity, and within most families each step down roughly halves the price. One step is already a real saving.

Generation is usually fine too. Moving from m5 to m7i generally gets better price-performance for the same workload, but test it anyway.

Family and architecture are where the big wins are, and also where things go sideways. Switching to a memory-optimized family, or moving to Graviton, can save a lot — but if anything in that stack has native dependencies or an old x86-only container image, that's where the surprise shows up.

Make one change, let it run through a full workload cycle, then consider the next. One practical note: changing the instance type on an EBS-backed EC2 instance needs a stop and start, so put it in a maintenance window instead of doing it live.

Step 4: Rightsize before you buy commitments, not after.

This is where it connects back to Issue #1.

Savings Plans and Reserved Instances lock you into a level of spend. Buy them based on the current footprint — oversized instances and all — and then rightsize, and usage drops below what was committed to. You end up paying for capacity you just worked hard to get rid of. You've basically pre-paid for the waste.

The order that actually works is boring but it matters: clean up the idle stuff first, then rightsize what's left, then commit to whatever the new baseline is. A surprising number of "why aren't our Savings Plans saving much?" conversations trace back to doing this backwards.

If you're on Kubernetes

On EKS, GKE, or AKS, the oversizing usually isn't at the node level — it's in the pod resource requests. Teams pad requests to be safe (same instinct as picking the bigger instance), the scheduler treats those requests as real, and the nodes look busy even though actual usage is way lower. Shrinking the nodes won't help until the requests are honest. Compare requested vs. actual CPU and memory per workload first — that's the Kubernetes equivalent of Step 1.

Not on AWS?

Same process. Azure Advisor and Google Cloud's machine type recommendations do the job Compute Optimizer does on AWS. The same questions apply: which metrics are they actually looking at, and does the lookback window cover the real peaks?

The actual takeaway: Oversizing isn't really a tooling problem. It's a default that nobody is forced to revisit. Take the 20 most expensive instances, look at a month of p95 CPU and memory, and step down one size where the numbers say it's safe. Do that before signing any commitment, not after.

What's the most oversized thing you've ever found in an account? Hit reply and tell me. I read every one, and the good stories might show up (anonymized) in a future issue.