Container limits turn gradual problems into sudden ones. A service that grows by a few megabytes an hour looks perfectly healthy until the limit arrives, and then it restarts without warning.
Learn the shape of your usage
Watch CPU, RAM, disk and network over real traffic before changing anything. The shape tells you which limit is actually binding:
| Pattern | What it means |
|---|---|
| Memory climbing steadily, never flat | real leak or an unbounded cache |
| Memory sawtooths up and down | healthy — a GC or cache cycle |
| CPU pinned, memory idle | CPU-bound: computation or a tight loop |
| Both high at the same time | real load — size for the peak |
| Spikes then restarts | bursts exceeding the limit |
Plateaus are the goal. Memory that grows forever is a leak whether or not it is fast enough to notice today.
Confirm a leak before chasing it
Hold a steady load for long enough that a leak becomes visible, then take two measurements far apart. Memory up substantially with identical load is a leak.
Before hunting, rule out the innocent explanations: a cache that has not filled yet, a runtime that has not run garbage collection, an unclosed connection pool, or load that simply increased.
The usual causes
- Unbounded caches. A map keyed by user ID or request ID grows forever unless it evicts. Bound it with an LRU, or a TTL — not just a size limit, since a size limit alone lets stale entries linger.
- Listeners never removed. In Node, each
setIntervaland unremovedprocesslistener is permanent. Long-lived containers accumulate them across restarts of your logic. - Database connections never returned. A pool that grows per request looks like a memory leak and is really a pool being bypassed. See Working with databases.
- Loading full result sets. Fetching ten thousand rows to process them in Python is a few hundred megabytes. Page the query or stream it.
- Reading whole files into memory. Process uploads as streams rather than reading them whole.
Reduce concurrency first
Before buying a larger plan, ask whether you need the concurrency. Each worker or thread carries a full baseline, so halving workers often brings memory back under the limit at a throughput cost that may not matter — especially for bots and low-traffic APIs, where the workload is not concurrency-bound.
For async frameworks, one process handling many connections usually beats many processes handling a few.
Give the heap explicit headroom
A managed runtime may size its heap for hardware much larger than your container, so it will not collect before the kernel intervenes. Cap it below the container limit:
NODE_OPTIONS="--max-old-space-size=384" # on a 512 MB container
Leaving space below the ceiling matters, because buffers and native allocations live outside the managed heap. A ceiling set equal to the container limit means the process is killed before a collection can ever run.
When to upgrade instead
Upgrade when the work is genuinely larger, not when it is leaking. Signals:
- Memory sits high but flat — real steady-state use.
- Disk fills with legitimate data, logs or assets.
- CPU is pinned under normal traffic.
- You need more backup slots or databases.
Steady high usage is a sizing problem and a bigger plan is the correct fix. A rising line is a bug, and a bigger plan only delays finding it — see Which plan fits your workload for sizing at about twice your measured peak.