The Problem
SES gives you sending quotas and reputation metrics, but it does not give you a kill switch that fires on its own. The usual failure looks like this:
- An IAM access key with SES permissions ends up somewhere it should not be, and someone starts using it to send spam.
- An application bug or a retry loop sends the same email over and over.
- An old or purchased list gets imported, and bounces and complaints climb quickly.
In every case, the damage grows with every minute nobody reacts. And when someone does react, the first question is always the same: which key, from which IP, and which sender address? Without logs, that question is very hard to answer after the fact.
What It Does
- Watches three signals: total sends per minute, bounce rate, and complaint rate in each region.
- Stops sending automatically: when any signal crosses its threshold, account level sending is disabled for that region.
- Logs every email event: send, reject, bounce, complaint, and delivery events go to CloudWatch Logs, with a retention of a year by default.
- Names the suspect: the alert includes the top senders from the last 30 minutes, broken down by IAM user, source IP, and from address.
- Alerts on several channels: Telegram for the team, PagerDuty for on call, and email as a backup.
- Watches itself: if the responder fails, a separate watchdog alarm says so, because a silent failure here means SES was not stopped.
The whole thing is one CloudFormation template plus a few shell scripts. It is deployed once in every region where SES is used, since SES sending, metrics, and quotas are all regional.
How the Pieces Connect
- Logging Every identity is pointed at a configuration set so its events are captured.
- A
ses-guardconfiguration set publishes email events to the default EventBridge bus. - An EventBridge rule forwards all
aws.sesevents to a CloudWatch log group. - The log group has a
Retaindeletion policy, so the evidence survives even if the stack is deleted. - A setup script attaches every SES identity to the configuration set. If an identity already has its own configuration set, the script leaves it alone and only adds the EventBridge destination to it.
- Detection Plain CloudWatch alarms on the metrics SES already publishes.
- Send spike: the sum of sends over one minute is above a per region limit, 500 by default.
- Bounce rate: the reputation bounce rate reaches 4%.
- Complaint rate: the reputation complaint rate reaches 0.08%.
- Missing data is treated as not breaching, so a quiet region never alarms by accident.
- Response All alarms publish to one SNS topic, which invokes a small Python Lambda.
- On an alarm, it calls
PutAccountSendingAttributeswith sending disabled. This is the actual stop. - It runs a CloudWatch Logs Insights query over the last 30 minutes of send events to find the top IAM users, IPs, and sender addresses.
- It sends the result to Telegram and triggers a PagerDuty incident with the same details.
- The Telegram message includes the exact commands to disable a suspicious access key and to resume sending, so nobody has to look them up during an incident.
- When an alarm returns to normal, it sends a short notice, and makes clear that SES stays stopped until a human resumes it.
- Watchdog A guard that can fail silently is worse than no guard, because it creates false confidence.
- A CloudWatch alarm watches the Lambda's error metric and notifies a separate SNS topic.
- If the stop call fails, the Lambda raises an error on purpose. That fires the watchdog and lets SNS retry the invocation.
- The watchdog topic goes to email and PagerDuty, not through the Lambda, so it still works when the Lambda is the thing that broke.
Incident Flow
A key leaks, or a job starts looping ↓ Sends per minute or bounce/complaint rate crosses a threshold ↓ CloudWatch alarm publishes to the SNS alert topic ↓ Lambda disables SES sending in that region ↓ Lambda queries the logs for the top senders in the last 30 minutes ↓ Telegram, PagerDuty, and email get the trigger, action taken, and suspects ↓ A human investigates, disables the key if needed, then resumes
The automatic part ends at the stop. Resuming is always a human decision, made after someone has looked at the logs.
The Control Script
Everything day to day goes through one script, ses-guard.sh. Each command accepts a region or all.
| Command | What it does |
|---|---|
status |
Shows whether auto stop is on, whether SES is sending, alarm states, and unconfirmed email subscriptions. |
top |
Lists the top senders by IAM user, IP, and from address for the last N minutes. |
logs |
Tails raw email events in a compact, readable format. |
stop |
Emergency stop. Asks you to type the region name to confirm. |
resume |
Turns sending back on and re-arms the alarms. |
deploy |
Deploys or updates the stack in every configured region, in live or test mode. |
setup |
Connects new SES identities to logging. |
Secrets such as the Telegram token and the PagerDuty key live in a local env file that is ignored by git, and are passed to CloudFormation as NoEcho parameters. The per region send limits live in the same file, so each region can have a threshold that matches its normal volume.
Testing the Guard
A safety system that has never fired is only an assumption. The script has three test commands, each with a different level of impact.
Alert test
Temporarily switches auto stop off, forces the send spike alarm into ALARM, waits for the responder, then restores auto stop. It checks that Telegram, PagerDuty, and email all arrive without touching real sending. A trap restores auto stop even if the script is interrupted halfway.
Stop test
Fires the alarm for real and checks that sending actually stopped. It asks you to type the region name first, then offers to resume right away.
Watchdog test
Forces the Lambda error alarm into ALARM and back, to confirm the "responder failed" path reaches a human.
There is also a test mode for the whole stack. Deploying with AutoStop=false keeps every alert but never stops SES, which is useful while tuning thresholds against real traffic for the first few days.
Design Decisions
Stop first, ask later
Pausing legitimate email for a few minutes is cheap. Letting a compromised key run for an hour is not. The guard is deliberately biased toward stopping.
Stop the region, not the key
Disabling sending for the whole region is blunt, but it works no matter how the emails are being sent, whether through SMTP credentials, an API key, or a role. Narrowing it down is the human's job, with the logs in hand.
Re-arm alarms on resume
CloudWatch alarm actions only run on a state change. If an alarm is left in ALARM after an incident, the next breach would not fire anything, and SES would be unguarded without anyone knowing. So resume sets every guard alarm back to OK. The reset uses a marked reason, and the Lambda ignores those transitions so it does not send a misleading "back to normal" message.
Keep the evidence
The log group is retained when the stack is deleted or replaced. Logs are the only way to answer "who sent this" after the fact.
Keep it small
One template, an inline Lambda, and plain shell scripts. No build step, no dependencies beyond the AWS CLI and jq. Something that runs rarely should be easy to read when it finally does.
Limitations
- Reputation metrics lag: bounce and complaint rates are computed by SES over time, so they react slower than the send spike alarm.
- A fixed limit is a blunt tool: a legitimate campaign above the per minute limit will stop sending too. The limit has to be set with real traffic in mind and raised before planned large sends.
- Only connected identities are logged: a new identity added without running setup still triggers the alarms, but its events will not appear in the top senders query.
- Per region: a region where the stack is not deployed is not guarded at all.
- Not a replacement for good IAM: the guard limits the damage of a leaked key. It does not prevent the leak. Scoped permissions and key rotation still matter.
Lessons
Build the response before the incident
The most useful part of the alert is not the warning, it is the list of suspects and the exact commands to run next. Deciding those in advance is what makes the response fast. This follows the same idea as rollback in A Practical Deployment Pipeline.
Guard the guard
Ask a simple question of every safety system: what happens if it fails? Without the watchdog, the answer here would be "nothing, silently".
Make testing a command
When testing is one command with a safe default, it actually gets done. Running the alert test after each deploy confirms every channel end to end, including email subscriptions that were never confirmed.
SES Guard is not a large system. It is a few alarms, one function, and some scripts. But it turns a problem that used to be discovered by angry recipients or an AWS review into one that is stopped within minutes and explained in a single message.