Rich Gibbs

The Second Cowrie OOM Had a Different Cause, and the Cap Was Never the Ceiling

homelab · honeypot · systemd · incident-notes · memory

Session 19 covers a recurring honeypot OOM. The memory cap raised in the previous session held for about two days, then the SSH honeypot started being killed again, this time at roughly 1.57 GB of anonymous resident memory — above both the original 512 MB limit and the 1.5 GB cap the earlier change had set. The limits actually in effect were being supplied by a systemd override drop-in while the unit file still read 512M, confirmed by asking the service manager for the live values instead of reading the file. The second incident's mechanism differed from the first: session churn from a new cluster, with file downloads and TCP relay attempts measured and ruled out as the driver. One gigabyte of swap was added this time, on the understanding that swap protects the host rather than bypassing the service's own cgroup cap, and the cap was deliberately left at 1.5 GB. The same window brought the web honeypot's first human-shaped traffic, a single non-banner SMTP session, and a hardware-sizing classifier that explained an earlier stretch of elevated CPU. The session does not record whether the cycling stopped.

Session 19, 2026-08-28 into 2026-08-29. The cap change from the previous session held for about two days, and then Cowrie started getting killed again. From outside it looked like the box was down. It wasn't.

The host was fine. Only Cowrie was cycling.

Same first move as last time, and it keeps earning its place: check host health before assuming the honeypot itself broke. Uptime was about 2 days 19 hours and clean, disk sat at 10% used, there was plenty of free RAM, and the reverse proxy, the mail honeypot and ssh stayed active the entire time. Only Cowrie was cycling.

A memory-cgroup kill looks like "the server is down" from outside while the host itself never wavers. I already knew that shape. What was different was the kill line. This time the process was killed inside the memory cgroup of the honeypot service, and it was reported at about 1.57 GB of anonymous resident memory — above the 512 MB of the original incident, and above the 1.5 GB the earlier change had raised the cap to.

That number is the whole point of this entry. The earlier fix wasn't wrong. It just wasn't the ceiling.

The cap I thought I had, and the cap that was live

The unit file on disk still literally said 512M. The limits actually in effect were 1.2 GB high and 1.5 GB max, and they were coming from an override drop-in rather than from the base unit. I confirmed that by asking the service manager for the live values instead of reading the file — a systemctl show on the memory properties, not the unit file.

Two reminders from that, one of which I had only half-learned:

  • A drop-in takes precedence over the unit file's own value. When a number on disk and a number in behaviour disagree, the behaviour is the truth.
  • Check the effective configuration layer, not the file that looks authoritative. systemctl show answers that directly, and it costs nothing.

While in there I also found a 90-second stop timeout. That explains why stopping the service sometimes looked hung: the process ignored SIGTERM for the full 90 seconds before systemd escalated to SIGKILL. Not a bug — a slow shutdown path that looks alarming if you don't know the timeout exists.

Same symptom, different mechanism

This is the part worth remembering. The first incident's growth came from a single cluster running lightweight one-command sessions. This one was heavier, and from a different cluster: an adjacent four-address family, thousands of connect/login/close events packed into a tight window, almost no commands typed and almost nothing downloaded.

I checked the two obvious suspects properly rather than assuming. File downloads and direct-TCP relay attempts were both real and both far too small to explain gigabyte-scale growth — 77 downloads and 358 relay attempts, as the session counts them. Enough to rule out. Nowhere near enough to be the cause.

The actual mechanism was session churn. The honeypot holds transport, TTY and key-exchange objects per session, and thousands of rapid connect-then-close cycles compounding with incomplete garbage collection walked memory up to the cap. Same service, same symptom, different cause.

Which gives the sharpened version of an old lesson: a memory leak does not have one universal cause just because it happened before under the same service. Re-diagnose each recurrence against the traffic that is actually there. Trusting the pattern instead of the numbers would have sent me after downloads.

Swap, this time actually added

The previous session considered 1 GB of swap and decided it wasn't needed yet. This time it went in: a 1 GB swapfile, permissions restricted, formatted with mkswap, enabled with swapon, and an fstab line added.

The distinction I wrote down explicitly, because it is easy to get wrong: swap does not raise or bypass the memory cap on the honeypot's cgroup. The cap still kills the service before it can lean on swap heavily. What swap protects is the host as a whole — ssh, the web server, everything else — if the box overall gets tight. Two safety nets for two different failure modes. They are not substitutes for each other.

Deliberately left alone, and recorded as options rather than decisions:

  • The cap stayed at 1.5 GB rather than being lowered back to 512 MB.
  • A nightly scheduled restart, as cheap insurance against slow leaks, was named and not applied.
  • A shorter stop timeout, same treatment.

And the honest end of the sequence: the swap going in is where the session notes stop. Nothing in here says the cycling stopped, so I am not going to claim it did. What I can say is what I changed and what I chose not to change.

The honeypots' first humans

The HTTP honeypot had been live but quiet. This window brought its first identifiably human traffic, plus one cross-service correlation worth writing down.

  • A distinctive login name crossed services. A name I had only seen on the SSH side showed up on the web honeypot from a different address, tried against a small set of default-style credential pairs. My read is that it is the same person, on username reuse across two completely different honeypot services. That is an evidence-based read, not a proven identity, and I am writing it down as the former.
  • A browser-and-scanner hybrid visit. One address browsed like a human while scripted service-method and SOAP/SDK probes ran alongside, then posted a two-word login attempt. The same visitor also pulled a planted breadcrumb file over SSH, meaning they were working both honeypot surfaces in the same visit.
  • Otherwise reconnaissance-shaped. Crawler bots, generic HTTP fingerprinting tools, WordPress-path probing — no brute-force flood yet. That fits the earlier finding that mass WordPress login attacks lag a few days behind a port becoming visible.

The mail honeypot logged 273 SMTP sessions, still zero captured credentials and an empty captured-mail directory — and exactly one session that did anything other than hang up after the banner. Three short words came in, shell, sh and exit, and each was rejected the same way, command not recognized. Someone was treating the SMTP channel as an interactive shell rather than speaking SMTP to it, which is most likely a human or a generic tunnel tool probing what sits at the far end of an SSH direct-TCP forward rather than a real relay attempt. Interesting as a single data point; not enough to change the reconnaissance-only conclusion about this traffic class.

A cluster that sizes the hardware before deploying

A second adjacent address family ran one large fingerprinting script at every login — PATH, uname, architecture, uptime, core count, CPU model, GPU detection, then a shell-behaviour self-test. That is a miner or compute classifier — it is measuring what the hardware can do, specifically whether there is a GPU, before deciding whether the host is worth deploying anything to.

This also explains a stretch of elevated CPU I had not been able to account for. The honeypot has to emulate every piped command in that script, which produced a real, measurable spike — around 8.5% CPU within a minute of a single login. Unexplained CPU now has a name attached to it.

And the quieter lesson from the same part of the log: process RSS is reported in kilobytes, so a value like 200000 is roughly 195 MB, not 200 GB. Check the units before reacting to a number that looks alarming at a glance.

Three profiles worth keeping

  • A Windows OpenSSH client logged in and used SFTP only, and only ever touched that account's own home directory, which was empty. It never reached the root account or the planted decoys at all. Recorded as the contrast case: a visitor who explored far less than most.
  • A client identifying itself as a phone-based web SSH client never completed its connection — a key-exchange mismatch the honeypot's Twisted version could not negotiate. Not an attack, just an incompatible client, recorded here specifically so I don't re-investigate the same failed key-exchange pattern as suspicious later.
  • Two crawlers with commercial attribution hit the web honeypot in its first days doing generic path cataloguing — admin paths, a VPN login path, generic CGI paths — rather than anything WordPress-specific. That supports the read that early HTTP traffic is broad internet cartography before it narrows into CMS-specific brute forcing.

Reminders for next time

  • A fix that resolves one incident does not guarantee the next one shares the same root cause. Re-diagnose against the traffic that is present, not against the last diagnosis.
  • Confirm which config layer is actually in effect. Drop-ins win silently over the file you are looking at.
  • Swap and a cgroup cap protect different things — the host versus one service. Neither substitutes for the other.
  • Identity can cross honeypot services. A name or behaviour signature appearing on two different services is corroborating evidence, not coincidence.
  • Sanity-check units on any dramatic-looking number. KB, MB and GB are an easy misread under pressure.

As currently documented, that closes out the honeypot's build-through-hardening arc.