Rich Gibbs

The full disk behind a misleading traceback

homelab · honeypot · dionaea · cowrie · incident-response · logging · disk-space

A full root filesystem, disguised by a cut-off Python traceback from Cowrie, turned out to be caused by a single unrotated Dionaea debug log left at default verbosity. The fix paired a one-line config change with host-level log rotation, after first checking the structured capture database so nothing useful got destroyed.

Cowrie started throwing unhandled Python tracebacks on exec sessions. First read, it looked like a bug in the non-interactive exec path — some internal call failing partway through setting up a session. It wasn't that at all, and the reason it took a minute to see that is worth writing down on its own.

The traceback I'd pasted into my own notes was cut off partway through the stack. I stopped reading at a library-internals frame and started reasoning from there, which is backwards. The exception type and message are always the last line of a traceback, not the frames above it — those are just the call path getting there. Cowrie's own log, two lines above the part I'd been staring at, already said what was wrong: no space left on device. I'd read past it the first time.

df -h / confirmed it in about five seconds: root filesystem at 100%. Cowrie wasn't broken. It had nowhere left to write a file.

Once the actual problem was "the disk is full," the next question was where all that space went, and I didn't want to guess. Guessing which directory is the hog wastes more time than just checking. du top-down, one level at a time: root, then the heaviest subtree, then the next level down inside that, and so on. It pointed straight at the Dionaea install.

The obvious suspect inside Dionaea is its raw capture-store directory — the thing that stores a byte-for-byte copy of every connection, which is unbounded by design and feels like the natural place for a disk to disappear into. I checked it anyway instead of assuming, and it was under 1 GB. Checked Dionaea's own downloaded-binaries folder too, and Docker's storage, same story — all modest. None of the "obviously risky" candidates were the actual cause. One more level down inside Dionaea's own var directory and there it was: a single plain-text debug log, roughly 28.6 GiB. Almost the entire disk, sitting in one file that had never been rotated.

The instinct at that point is to truncate it and move on. I deliberately didn't do that first. That log covered eight days of continuous internet exposure that had never been reviewed — the only prior look at this honeypot's traffic was a single-night spot check right after it went live, and that night's read was "mostly recon, nothing much happening." Eight days is not one night, and I didn't want to destroy data on the assumption that the one-night read still held. It didn't. The structured capture database told a much richer story: real downloaded payloads, a large batch of credential-stuffing attempts, real database-protocol query attempts well past a handshake, and — the genuinely good find — two shellcode emulation traces, one of which was a fully traced API call sequence shaped like bind-shell backdoor behavior: create a file, load the networking library, open a socket, bind it, start listening. That's a real attacker payload actually executing and getting traced, not just "something connected." The second trace record came back with an empty call list, most likely a failed or partial capture — not confirmed, and it's still sitting as an open question.

A few other things worth remembering about what was in there. The downloads table had no URL logged for a single row, and before assuming that was a bug I checked it directly — it genuinely was empty for every one. Turned out to be correct behavior, not a defect: the delivery method pushes files straight into the fake share, so there's never a URL to log in the first place, unlike something fetched over HTTP. Two separate rows in the log were share-enumeration probes, not downloads at all — worth keeping those as a distinct event type rather than lumping them in. And on the database-query side, the table only ever stores the wire-protocol command type, never the literal query text. I went back and grepped the full debug log before truncating it, specifically to check whether the actual query strings existed anywhere else at a higher verbosity. They didn't. Confirmed, not assumed — that log held nothing recoverable for that particular question, which made truncating it safe with certainty instead of a guess.

The credential-stuffing table had two fingerprinting details worth keeping for later: a keyboard-walk password pattern common in stuffing wordlists, and a username that's the known default on a vendor's free test-VM image — a tell that whoever built that particular wordlist was targeting exposed Windows test environments specifically, not just running a generic top-password list.

Fix, in two parts. First, the immediate one: truncate the oversized log in place rather than delete it. The container holds an open file descriptor to that path, so truncating lets it keep writing from byte zero with no disruption; deleting it would leave the descriptor pointing at space that's unlinked but still consumed until the container restarted anyway. Second, the actual root cause: one configuration line had logging verbosity set to capture every level, unfiltered, forever. That's what produced 28.6 GiB of mostly debug-level internals — connection housekeeping, byte-copy operations, thread timing — none of it useful for understanding an attack. Changed that one line to exclude debug output and restarted the service, since there's no live-reload path for a config change like this.

Then I checked whether this was a problem unique to my setup, because it felt too clean to be novel. It wasn't unique. A GitHub issue against the upstream project, filed in 2017, describes this exact behavior — gigabytes of this log in under a day — and the default that causes it is still shipped unchanged in the maintained container image nine years later. The project's own documentation says plainly that this log isn't meant to be used to analyze attack traffic. That lines up with every real finding in this session: the structured database was always the right place to look, and the raw log was never supposed to be the analysis surface — it just happened to also be a disk-filling liability on top of that.

Lowering the verbosity slows the growth but doesn't eliminate it — a log with no rotation will always eventually be a problem again, regardless of level, and connection volume on this box trends upward over time. So I added a second, independent control on top of the config fix: host-level log rotation, daily, with a 7-day retention count, using copy-truncate mode specifically because this container has no reload hook that would let standard rotate-and-signal behavior work — copy-truncate copies the content out and truncates the original in place, the same mechanism as the manual fix in step one. I ran it in dry-run mode to confirm the config parses and matches the right files. That only confirms it's correctly scheduled and will fire on its own cron — it is not the same as having watched a rotation actually happen yet, and I'm not claiming otherwise.

After the restart: container back up, every port mapping intact, disk usage back down to a healthy level. That's a point-in-time check right after the fix, not an ongoing guarantee.

What changed, in rough terms rather than literal paths: the oversized debug log got truncated (only after the structured data behind it was reviewed and confirmed not to need it); the honeypot's own config had the one verbosity line changed; and a new host-level rotation policy was created where none existed before.

Lessons, for next time:

Read the whole traceback. The line that tells you what's actually wrong is the last one, not the frame you happened to stop scrolling at.

df -h / is a five-second check. Run it early in any "something's acting weird" investigation — in this case the real cause was visible in the very first diagnostic command, buried under a stack trace pointing somewhere else entirely.

Drill down one directory level at a time instead of jumping straight to whichever component looks theoretically riskiest. The "obviously unbounded" directory wasn't the cause; a boring unrotated log was.

Check structured, recoverable data before destroying raw data, even under time pressure. Everything interesting in eight days of capture — the shellcode trace included — would have been gone if that log had been truncated on sight.

One good night of data is not eight good days of data. A single-sample read doesn't get to stand in for a longer window just because checking again takes effort.

If you find a local defect, check whether it's actually a known upstream one before assuming it's yours alone. One search confirmed this exact failure mode has been reported and left unfixed for years.

When documentation says "don't use this for X," believe it, and then go verify why — in this case there was a second, independent reason beyond the one the docs mention, confirmed by actually grepping for it rather than taking it on faith.

Pair a root-cause fix with an independent backstop. The config change addresses why the log grew the way it did; the rotation policy bounds whatever grows next, for whatever reason, regardless of whether the first fix holds.

Left open, honestly: that second emulation trace with the empty call list still doesn't have a confirmed explanation. None of the repeated download hashes have been checked against any reputation service yet. And it's worth checking whether that same unfiltered-logging default is sitting on any other instance of this software on the box — I haven't looked yet.