Lethio

The Wall of Shame

mar 2026 · security · self-hosting

Edited 16 August 2026. Two code samples below had drifted from the script actually running. Fetching the current bans used fail2ban-client banned and pulled every IPv4 address out of the reply with a regular expression, which was only correct while the server had exactly one jail. Adding a second would have quietly folded its bans into this page's figures, with nothing to indicate it. That regex also skipped IPv6 addresses entirely. Separately, the log parser read fail2ban.log and .log.1 only, ignoring the compressed rotations, so the page reported on a fraction of the history it appeared to cover. Both are corrected in place below.

The first time you check the logs on a freshly deployed server, it's a little jarring. You put a box on the internet maybe an hour ago and there are already hundreds of failed SSH login attempts. Nobody knows your server exists. Nobody cares. The bots found it anyway, because they weren't looking for your server. They're scanning every IP address, all the time, forever.

This isn't a targeted attack. It's the background radiation of the internet. Automated scripts cycling through default usernames and common passwords on every open port 22 they can find. Most of them come from a handful of countries, run on compromised machines or cheap VPS instances, and they never stop.

This guide covers how to catch them, ban them, figure out where they're coming from, and put that data on your website so you can watch it happen live. Which is honestly more entertaining than it sounds.

Installing fail2ban on Debian 12

Fail2ban watches your logs and bans IPs that keep failing authentication. It's been around forever because it works.

sudo apt update && sudo apt install fail2ban -y

One thing about Debian 12 that trips people up: it uses the systemd journal by default instead of writing to /var/log/auth.log. If you install fail2ban and it doesn't seem to be catching anything, this is probably why. The default backend looks for log files that don't exist.

The fix is to tell fail2ban to use the systemd backend. Create a local config override (never edit the files in /etc/fail2ban/jail.d/ directly, they get overwritten on updates):

sudo cp /etc/fail2ban/jail.conf /etc/fail2ban/jail.local

Then edit /etc/fail2ban/jail.local and set the backend:

[DEFAULT]
backend = systemd

Configuring the sshd jail

The sshd jail is the one you care about. Still in jail.local:

[sshd]
enabled  = true
port     = ssh
filter   = sshd
maxretry = 3
findtime = 600
bantime  = 3600

This bans any IP that fails three times in ten minutes, for one hour. Which is fine as a starting point, but the bots don't learn. They come back. So you want incremental bans: each repeat offense gets a longer timeout.

[DEFAULT]
bantime.increment    = true
bantime.factor       = 24
bantime.maxtime      = 4w
bantime.overalljails = true

With this, a first ban is one hour. If they come back, it goes up. The factor of 24 means the second ban is 24 hours, and it keeps climbing up to a maximum of four weeks. overalljails means repeat offenders get punished across all jails, not just the one they tripped.

Restart fail2ban and check it's running:

sudo systemctl restart fail2ban
sudo fail2ban-client status sshd

You should see the jail listed as active with the filter picking up connections from the systemd journal.

Geoip enrichment

Knowing that 218.92.0.107 got banned is useful. Knowing it came from China makes the data a lot more interesting to look at. GeoIP databases map IP addresses to countries:

sudo apt install geoip-bin geoip-database -y

Now you can look up any IP:

$ geoiplookup 218.92.0.107
GeoIP Country Edition: CN, China

The free GeoIP database isn't perfect. Some IPs come back as "Unknown." VPS provider IPs sometimes register as the wrong country. But for aggregate stats and a "where are the brute-force attempts coming from" overview, it's more than good enough.

The data generation script

The idea is simple: a script reads the fail2ban log, runs geoiplookup on each banned IP, and writes a JSON file with the results. A cron job runs it every minute. The website fetches the JSON and displays it.

The script is written in Python because it handles JSON natively and the only external calls are subprocess calls to geoiplookup and fail2ban-client. No pip dependencies. The key bits:

Parsing the log for ban events:

BAN_RE = re.compile(
    r"^(\d{4}-\d{2}-\d{2}\s\d{2}:\d{2}:\d{2}),\d+\s+"
    r"fail2ban\.actions\s+\[\d+\]:\s+NOTICE\s+"
    r"\[sshd\]\s+Ban\s+(\S+)"
)

Getting the currently banned IPs (not just historical bans):

result = subprocess.run(
    ["fail2ban-client", "get", JAIL, "banip"],
    capture_output=True, text=True, timeout=10
)
return list(dict.fromkeys(result.stdout.split()))

Ask about one jail by name, rather than asking for everything and filtering afterwards. fail2ban-client banned returns every jail at once as a Python repr blob, and the obvious move is to regex the addresses out of it, which is what this script did until August 2026. That is correct right up until the day you add a second jail, at which point its bans silently join this page's "currently banned" figure while the totals and the country breakdown, which are matched against [sshd], do not move. Nothing errors. The number is just wrong. Asking for a single jail cannot drift that way.

Splitting on whitespace also picks up IPv6. The old regular expression matched four dot-separated numbers, so an IPv6 ban never appeared here at all. Since sshd listens on IPv6, that was a matter of time rather than a hypothetical.

GeoIP lookups are cached per IP so the same address doesn't get looked up twice. The JSON output includes a summary, the last 50 ban events, top countries, and the list of currently banned IPs. The write is atomic (write to a temp file, then rename) so the frontend never reads a half-written file.

The script also handles log rotation, and this is worth getting right, because getting it wrong is invisible. It reads every /var/log/fail2ban.log* the rotation keeps, including the compressed ones, since logrotate gzips everything from .2 onward:

for log_path in sorted(glob.glob("/var/log/fail2ban.log*")):
    opener = gzip.open if log_path.endswith(".gz") else open
    with opener(log_path, "rt", errors="replace") as f:
        ...

Until August 2026 it read fail2ban.log and .log.1 and stopped there. That looked like it handled rotation, and it did survive a rotation cycle, but it meant the headline figure covered only as far back as the last uncompressed file, and reset every time the logs rotated. Reading the .gz files as well roughly doubled the count overnight, from a number that had been quietly wrong for months. If your counter looks suspiciously round or never seems to grow past a certain point, this is the first thing to check.

Ban events older than the log retention period are then discarded in the script as well, rather than trusted to be gone. An IP address is personal data, kept only for the security purpose that justifies keeping it, so a stray file that outlives a rotation must not put a month-old address onto a public page. The retention period is published in the JSON alongside the count, so the sentence on the front page says which window it means instead of implying a total since the beginning of time.

The frontend

The website just fetches the JSON file and renders it. No websockets, no server-sent events, no build tools. A static JSON file that updates every minute and a vanilla JS script that polls it.

Why not websockets? Because this data updates once a minute, not once a second. Websockets add server-side state, connection management, reconnection logic, and a reason for nginx to hold open persistent connections. For data that changes every 60 seconds, polling a static file is simpler, cheaper, and more resilient. If the server goes down, the frontend just shows the last good data until it comes back.

The fetch runs on page load and then every 60 seconds via setInterval. If it fails, it doesn't crash the page or show an error modal. It just quietly keeps the last good data in place.

The flags will not render for everyone

Country codes become flags by mapping each letter to a regional indicator symbol, which is the standard trick:

String.fromCodePoint(...[...code.toUpperCase()].map(
    c => 0x1F1E6 - 65 + c.charCodeAt(0)
))

Windows ships no flag glyphs at all. Its emoji font has never included them, so Chrome on Windows renders the two regional indicator letters instead. That is not a bug: it is literally what a flag emoji is made of underneath. Firefox bundles its own emoji font and draws real flags, on Windows too. macOS, Android and iOS are all fine. So whether your wall looks good depends on a browser and OS combination you probably are not testing on, and if you happen to use Firefox you may never notice. Microsoft has never explained the decision; the reason usually given is that drawing a flag means deciding what counts as a country, but that is inference rather than anything they have confirmed. Emojipedia tracks per-platform support.

You can fix it by self-hosting a flag font of around 55KB. Whether that is worth it depends on whether the flag is carrying information or decoration. Here the country code sits right next to it, so it is decoration, and this site shows a short note to affected readers instead. Detecting which readers those are is worth doing by measurement rather than the usual way:

probe.textContent = "\u{1F1F3}";          // one regional indicator
const single = probe.getBoundingClientRect().width;
probe.textContent = "\u{1F1F3}\u{1F1F1}"; // the NL pair
const pair = probe.getBoundingClientRect().width;

const drawsFlags = pair < single * 1.5;   // ~2x wide means no flag support

A supported flag is one glyph; an unsupported one is two letter glyphs, so width alone separates them. The common approach instead paints a flag to a <canvas> and reads the pixels back to see whether any colour appeared. That works, but getImageData is exactly the call that fingerprinting scripts use, so privacy tooling watches it and Firefox's resistFingerprinting can hand back spoofed pixels. Measuring a hidden span asks the browser nothing it would not tell any layout engine.

What the data shows

The majority of brute-force attempts come from a small number of countries. China, the United States, and a rotating cast of others consistently top the list. This doesn't mean people in those countries are personally attacking you. It means those countries have a lot of compromised machines and a lot of cheap VPS providers who don't police their customers.

You'll see the same IP ranges come back repeatedly. These are scanning farms, likely cycling through enormous IP lists. Some of them are remarkably persistent. Others try once, get banned, and never return.

A surprising number of source IPs belong to cloud providers. Hetzner, DigitalOcean, OVH, Alibaba Cloud. Which makes sense: if you're running a botnet, you want cheap compute with decent bandwidth, and you don't care if the account gets suspended because you'll spin up another one tomorrow.

None of this is cause for alarm. It's just the internet being the internet. The important thing is that fail2ban catches them before they can try more than a handful of passwords, and the incremental bans mean repeat offenders get locked out for weeks.

Before you publish the addresses

An IP address is personal data in the EU. That's been settled since the Breyer ruling in 2016, and dynamic addresses count too. If you're somewhere else, go and check what your own rules say, because a lot of jurisdictions have landed in the same place.

Keeping them is the easy part. You need the full address to block anything, it's already in your logs, and stopping an attack is about as clean a legitimate interest as you'll find. Nobody sensible objects to that.

Publishing them is a different question, and the test isn't whether you can. It's whether you need to. Blocking an attack doesn't require showing the address to your visitors.

There's a real difference between sharing and displaying. AbuseIPDB, Spamhaus, DShield and your national CERT all take reports of attacking addresses, and they exist so other people can block the same hosts before they get hit. That's the address doing actual work. Sending your bans there is useful and proportionate, and it's the case the legitimate interest argument was written for.

A list on your homepage is not that. It doesn't help anyone block anything. It looks cool, which is a real reason, just not one that outweighs someone else's personal data.

Then there's who you're actually naming. Plenty of these hosts are compromised. The address you're about to put on the internet under a heading about shame might belong to someone whose NAS got owned two years ago and who has no idea it's out there knocking on doors.

So do both properly: report the full addresses to somewhere that acts on them, and cut them down before they hit the page. The script on this site drops the last octet at the point the JSON gets written, so the full address never leaves the server and isn't sitting in the page source either. You keep the wall. You just stop publishing somebody's address to make it.

def mask_ip(ip: str) -> str:
    """Drop the identifying tail before publishing. IPv4 keeps the /24, IPv6 the /48."""
    if ":" in ip:
        groups = [g for g in ip.split(":") if g]
        return ":".join(groups[:3]) + ":x" if len(groups) >= 3 else "x"
    octets = ip.split(".")
    if len(octets) == 4:
        return ".".join(octets[:3]) + ".x"
    return "x"

Call it where you build the JSON, not anywhere earlier. The geoip lookup needs the real address to resolve a country, your unique-IP count needs it or two machines on one /24 collapse into one, and fail2ban obviously needs it to ban anything. Masking in the browser instead is pointless: the JSON goes over the network, so the real address would still be sitting in the network tab for anyone who opened it.

The next layer

Brute-force attempts are the first thing to deal with. The second is making sure your web server isn't leaking information or missing security headers. Harden Your Web Server covers the nginx side: CSP headers, HSTS, hiding your server version, and blocking access to files that shouldn't be public.