Skip to the content
Georgi DimitrovdaTuzzo

The server compromise

A framework RCE, a botnet, three failed rebuilds and the outbound connection that found it

Role
Solo, incident response
Status
Shipped
Source
Private repository
Stack
Ubuntu 24.04Next.jsOpenSSHUFWfail2bansysctlssTailscale

In numbers

3

clean-ups and rebuilds re-infected before the root cause

4h

from first deploy to root cause

42s

to install the patched Next.js 16.0.11

18

crontab entries the malware wrote for persistence

10

hardening steps, then 10 verification checks

The problem

I deployed the ValoxVSL video tool to a fresh cloud server. About an hour in, the server was found running x86_64.kok, a Mirai variant disguised as systemd-logind and persisted through crontab (18 entries), rc.local, init.d and bashrc. A clean-up in rescue mode did not hold, and neither did two rebuilds, one with new SSH keys and one behind a cloud firewall that limited SSH to a single IP. Each time it was back within five minutes.

The approach

I stopped cleaning and looked at what the server was talking to. ss -tnp state established showed the Next.js process itself holding a connection to a command-and-control server, which moved the search from SSH to the application. The npm dependency tree checked out clean. The cause was React2Shell, an unauthenticated remote code execution in the React Server Components protocol, which the incident report records as CVE-2025-66478 (CVSS 10.0). The app ran Next.js 16.0.1, and the RondoDox botnet was exploiting it over ports 80 and 443; each rebuild had redeployed the same version. Upgrading to 16.0.11 took 42 seconds. The server was rebuilt once more, its secrets were rotated, and a 15-minute automated watch came back clean.

How it works

  1. Outbound connections as the signal

    Inbound firewall rules cannot see a connection the server opens itself. ss -tnp state established lists each established connection with the process that owns it, and an application process connected to an unknown host is where to start. It is now a line of the verification checklist, and the expected output is my own SSH session and nothing else.

  2. Two attacks at once

    auth.log showed SSH brute force, so three rounds of work went into SSH. That attack was real and concurrent; the persistent one came through the application's own ports, which have to stay open. Ubuntu 24.04 added noise: ssh.socket's default TriggerLimitBurst of 20 kills the SSH listener after 20 connection attempts, which is why SSH kept dying during the brute force.

  3. The playbook

    Written from the incident: ten steps, from disabling ssh.socket, through key-only SSH with Mozilla's modern ciphers, a non-root user and a locked root, UFW deny-incoming, fail2ban, unattended upgrades with automatic reboot off, sysctl hardening and a login banner, to egress blocks for known C2 addresses. Ten checks verify it afterwards. Its rule for a compromised server: do not clean it; dump the database, rebuild fresh and rotate every secret, since malware inside the app process can read its environment.

  4. A second lesson, later

    Later I locked myself out of the game server: UFW and the cloud firewall both pinned SSH to one IP, and the match failed. The playbook now says key-only SSH plus fail2ban, with no IP pinning in UFW. The game server was rebuilt to it almost line for line, down to the banner text. Later hosts start from it and adapt: admin surfaces bind to the tailnet or loopback, and the Biogard stack adds container isolation.

What I chose, and what lost

Chose

Patch the framework, rebuild once more and rotate every secret

Over

A fourth round of SSH hardening and clean-up

Each rebuild had redeployed the vulnerable Next.js, so the server was re-infected over HTTP whatever the SSH setup was.

Chose

Key-only SSH plus fail2ban, with no IP pinning in UFW

Over

Restricting SSH to one IP in UFW and the cloud firewall

That setup later locked me out of the game server.

Outcome

Resolved once the framework was patched. The report records the patched deployment clean on a 15-minute watch and after two hours of uptime, with the botnet's later attempts showing up in the logs as rejected Server Action calls. The post-mortem runs to 647 lines in 21 sections, with a narrative version beside it. The playbook is the baseline for new servers and is applied unevenly: the game server follows it almost exactly, and the others take parts of it and substitute container isolation and tailnet-only admin.

What comes next

Bring the oldest box in the fleet up to the playbook it inspired.