Roughly a week of captures from three co-located honeypot sensors, analysed in one pass and written up at the end of it. The parts worth keeping are not the malware. They are two things the analysis itself got wrong, and what each one cost.
For scale, the SSH/Telnet sensor logged 15,168 session connects, 19,477 commands, 606 downloads and 26 uploads over the window. Only 297 connects ever negotiated an interactive terminal. Everything else was a script talking to a script.
The categories only a full event-type count found
My first pass worked the event types I already knew how to read, the command stream and the download log, and it missed whole categories. What fixed that was building a complete event-type histogram over the entire capture before drilling into anything. That histogram is the master index of the dataset, and it surfaced two things the greps structurally could not.
The first was uploads. Twenty-six events, all payloads pushed to the sensor rather than fetched by it. Download-based analysis is blind to those by construction, because the attacker never runs the fetch, the file simply arrives over SCP. The week's most significant sample was sitting in that category, and it had been invisible.
The second was the true scale of relay abuse. Attackers who log in and, instead of running commands, ask the sensor to forward a connection outbound are using it as an anonymising relay. That is 1,024 requests, one of the largest categories on the sensor, and I had under-weighted it in the first draft. Six hundred and thirty of them aimed at mail submission or relay ports, with the single most-targeted destination a large consumer mail provider. The forwards themselves were discarded, so nothing actually transited, and what the sensor holds is the intent, the destination and the port. Some destinations were IP-reflection lookup services, which is the relay operator checking what address it appears to come from before trusting the path.
The general lesson here is cheap to state and easy to skip. Enumerate the shape of a dataset before interrogating a slice of it. Working from the event types I already understood is exactly what hid both categories.
The sensor that looked empty, read alone
The SMTP sink recorded 328 sessions and captured zero credentials. Every session was one outbound banner and a disconnect, with no authentication, no sender and no message data at all. Read on its own, that is a null result.
Read together with the relay destinations above, it is the same mail-abuse traffic seen from the opposite end. The abuse was never aimed at the sensor's own mail service. It was aimed at real external mail servers, relayed through the SSH pivot. One dataset said nothing happened. Two datasets said a major theme was sitting right there the whole time.
What the week actually contained
Worth recording briefly, because the taxonomy is the part that carries the method point.
Traffic sorted into four tiers: commodity IoT and botnet automation as the bulk of it, an automated multi-vector scanner, a skilled human operator, and an unskilled human working from a pasted guide. The separation is argued from behaviour rather than asserted, using whether a session negotiates a terminal size, how it paces itself, copy-paste artifacts, typos, and whether the operator troubleshoots when something fails.
The strongest human case is a complete Discovery to Credential Access to Lateral Movement chain. It spans a multi-hour session with a genuine three-hour gap in the middle, sweeps credential-shaped decoy files, then uses a value lifted from one of them to try to reach a second host. Destructive commands were run and executed harmlessly against the emulated filesystem. The gap is the evidence I trust most, because automation does not pause three hours and come back to finish a job.
The most significant capture was in the upload category, a seven-file multi-architecture cryptominer kit pushed by SCP and deployed twice, two days apart, byte-identical both times. Five CPU architectures, plus a cleanup script and a setup script. The scripts delete themselves after running, and the setup script clears the immutable and append-only flags on the SSH authorized-keys file before writing a persistence key, which defeats a common hardening measure. The family is XMRig-based Monero cryptojacking. The same traffic also showed anti-honeypot fingerprinting, a shell-realism test and an architecture probe, run before a payload was committed.
Attribution on the most widely seen downloaded artifact resolved to a long-running campaign, publicly documented since 2018, with a fixed marker string in its persistence comment. The useful part is not the name. It is that the identical hash was independently captured and written up by an external handler diary, which moves a local inference into something corroborated by someone else's research.
One more line for the low-count findings. A telnet authentication-bypass attempt class was seen three times, from three different addresses over three days, all using the same injection string against a known argument-injection flaw. The HTTP surface, meanwhile, was reconnaissance and secret-file hunting rather than exploitation, over a capture window much shorter than the SSH one, so those two volume figures are not comparable to each other.
The records outlived the samples
This is the mistake with the longer tail.
The download directory and the terminal-log directory were both empty at write-up time, despite 606 downloads and nearly eleven thousand terminal-log events in the records. Samples were being rotated off disk while the JSON event records persisted. The malware findings are now stands-on-records-and-hashes findings. The binaries are gone and can no longer be examined.
Logging an event is not retaining the evidence. Those are two different systems with two different retention windows, and the one holding the evidence was the one nobody was watching. The write-up always happens after the rotation, so the retention window has to be at least as long as the gap between a capture and sitting down to analyse it. That fix is written down as an open item rather than a completed one, which is what it honestly is.
Reminders for next time
- Count every event type before drilling into one. The histogram is the master index, and the categories you never think to grep for are the ones it finds.
- Read sensors together. An empty result on one sensor can be half of a large finding on another.
- Retention is part of the instrument. If samples rotate before write-up, the analysis rests on records and hashes, and the samples do not come back.
- Keep external corroboration and local inference as separate classes of evidence, and say which one you have.
None of this is about the malware being subtle. It is about the order I did things in, which sensor I read on its own, and what had already been deleted by the time I sat down.