What is an AWS cost spike, and how do I catch it?
read

Quick answer: An AWS cost spike is a sharp jump in spend against your own normal, within hours or a day. It is not a slow drift or a tax lump. You catch it early by comparing each hour and each day to your own baseline, on hourly Cost Explorer data, and sending the alert where your team already works.
Every AWS bill has a shape. A spike breaks that shape fast. Here is what counts as a spike, what does not, and how to hear about one while it is still small.
What counts as a spike, and what does not?
Three things make a bill bigger. Only one of them is a spike.
What it looks like | What causes it | What should happen | |
|---|---|---|---|
Spike | A sharp jump within hours or a day. Yesterday $80, today $400. | A runaway job. A retry storm. A leaked key. A forgotten GPU instance. | An alert within hours, to a human who can act. |
Drift | The floor moves up over weeks. Nobody can name the day it started. | Each deploy adds a little. Logs grow. Snapshots pile up. | A weekly review. A digest, not a page. |
Lump | Money booked all at once, on one day. | Tax on the 1st of the month. Credits as one row. A refund. | Nothing. It is not spend. Detection should skip it. |
A spike is the one that hurts. It grows while nobody looks. A drift costs more over a year, but you can find it at your own pace. A lump only looks like a spike. If your alerting fires on the 1st of every month, it is reading the tax row as spend.
Why does a fixed dollar threshold miss most spikes?
Because a threshold is a line you guess in advance. A spike does not care where you drew it.
Say your budget is $5,000 a month. A $300 runaway job starts on the 3rd and runs for a week. $300 out of $5,000 is 6%. Your total is still under the line. No alert. You find it on the invoice. Set the line low instead, and it fires on a normal Monday. After the third false alarm, everyone mutes the channel.
So stop guessing a number. Compare spend to your own normal. Your normal is last week, and what that hour cost yesterday. A jump against that is a spike, whatever the total.
Which two shapes are worth alerting on?
Two comparisons catch almost every real spike.
A day well above your 7-day average. Average the last seven days. If yesterday ran 30% above that, something changed. This works on daily data, so every account can use it.
An hour far above the same hour yesterday. Compare 3 PM today to 3 PM yesterday. If it costs three times as much, something is running that was not running before. This needs hourly data, and it is the fast one.
Both compare you to you. A big account and a small account get the same rule. Neither needs a budget number typed in.
Why does hourly data matter so much?
Because a spike is measured in hours, and daily data hides them.
A daily total is 24 hours mixed together. A job that starts at 2 PM and burns $500 by midnight is just a bigger day. Hourly data shows the 2 PM jump on its own.
Hourly data is the fastest signal AWS gives you about spend. It is not on by default. You turn on hourly granularity, the per-hour view, in Cost Explorer. AWS bills it at $0.01 per 1,000 usage records a month and keeps 14 days of it. Cost Explorer itself runs 12 to 48 hours behind, so no tool sees spend live. Hourly is as close as it gets.
How late are the tools that come with AWS?
They come with your account. They also run on their own clock.
AWS Budgets updates up to three times a day, 8 to 12 hours apart. It fires when spend crosses a line you set.
Cost Anomaly Detection begins within 24 hours of setup and checks about three times a day. It reads Cost Explorer, so it inherits the lag.
CloudWatch billing alarm fires only on the current estimated total. Not on shape. It is a threshold with a different name.
All three are yours to run. You pick the numbers. You create an SNS topic, AWS's own notification relay. You subscribe to it. You confirm. Slack needs a separate Amazon Q Developer setup. When normal drifts, you re-tune. Nothing creates or updates itself. Here is the full comparison with Budgets.
How do I keep spike alerts quiet?
Noise kills alerting faster than lag does. Five rules keep the channel useful.
Set a floor. Alert on dollars that matter.
Use a cooldown. A quiet period after each alert. One spike is one alert. Not one per hour it lingers.
Cap each rule. A flapping rule should throttle itself and leave the others alone.
Deduplicate tickets. The first alert opens the Jira issue. Later ones comment on it.
Skip the lumps. Tax, credits and refunds are not spend. Keep them out of detection.
Then send the alert to the person who can fix it. That is an engineer, and engineers live in Slack or Jira. A spike is a to-do, not a newsletter. Here is how to get it into Slack. When it lands, the cause is on a short list. Nine causes cover most spikes.
Who gets paged when your bill doubles?
Nobody at AWS. That is not a dig. It is how the incentives sit. AWS bills what you use. A quiet spike is your problem, and you meet it on the invoice.
How does watchmy.cloud help here?
Connect AWS. Read-only billing access. No keys. No access to what you run. One CloudFormation template, about two minutes.
Send it where work happens. Slack for fast triage. Jira for follow-up. API when you want control.
Stop worrying about it. We watch AWS spend for you, so you don't have to keep checking it.
Both shapes above are on from day one. One dial, 1 (relaxed) to 5 (strict), tunes them. Cooldowns, a cap of 24 alerts a day per rule, and ticket dedup are built in. Tax, credits and refunds stay out. We check every hour. A flat $49 a month. See how it works. We are a small team, and this is all we do.
Watching AWS spend is our full-time job. You focus on your product and sleep well.
FAQ
What is an AWS cost spike? A sharp jump in spend against your own normal, within hours or a day. A runaway job, a retry storm or a leaked key are typical causes. A drift is different. It moves the floor up over weeks.
How big does a jump have to be to count as a spike? There is no fixed dollar amount. Compare to your own baseline. A day 30% above your 7-day average, or an hour at three times what that hour normally costs, is a good starting point.
Why did my cost alert fire on the 1st of the month? AWS books the month's tax as one row on the 1st. A tool that treats that row as spend sees a spike. Good detection skips tax, credits and refunds.
Can AWS Budgets catch a cost spike? Only if the spike pushes your total over a line you set. A $300 runaway inside a $5,000 budget never crosses it. Budgets also updates up to three times a day, 8 to 12 hours apart.
Do I need hourly granularity to catch spikes? Not for daily spikes. For an hour-level spike you need hourly granularity in Cost Explorer. AWS bills it at $0.01 per 1,000 usage records a month.
How fast can I hear about an AWS cost spike? Cost Explorer runs 12 to 48 hours behind, so no tool sees spend live. With hourly data checked every hour, an alert lands about an hour after the spike shows up in the data. With AWS alone, expect the same day or the next one.
See it before you connect anything: live demo, no sign-up.
More




