CCE

Limiting expensive to render nginx endpoints

Contents

So a Fun Thing about running a web server on the internet is that it gets scanned by a bunch of automated agents who don't even pay attention to robots.txt; so sometimes something more aggressive needs to happen.

Lately, the web scraping on my Gitea has become much worse with HTTP clients from hundreds of different IP addresses, sometimes residential, sometimes cloud, sometimes VPN exit providers, always with different forged User Agents of many different mainstream browsers. Just the meanest mothers. I'm not the only person dealing with this either right now. All of these useless requests for every commit's version of every file, and every git blame for every one of those, for every repo including my nixpkgs checkout and various other things I mirror for my own sake.

I'm lucky enough that it's not wholly grinding my system to a halt because My Homeserver is hugely over-provisioned for what I ask of it, but even 128kib/s 24/7 ends up being nearly a third of my monthly internet data cap! But perhaps I have a decent enough solution to limit these requests without deploying a web-application firewall, it at the very least limits the amount of traffic transiting to The Wobserver...

data/202407/12T175747.710960/Screenshot_20250319_221056.png

There's three layered rate-limiters in here that are applied to only certain URIs on my Edge Server before the traffic gets close to my Gitea instance:

  • One does a per-IP limit excluding my Tailscale network and some ASNs I connect from. Each IP can make one costly request per minute, otherwise receive a 503.

  • One tries to map certain cloud providers in to a single rate-limit key. Each group of cloud IPs can make one request per minute, otherwise receive a 503.

  • One puts a global limit on each "site feature" in Gitea.

nix:tangle ~/nix/nixos/gitea-ratelimit.nix:noweb yes
{ pkgs, config, ... }:

{
  services.nginx.commonHttpConfig = ''
    # defines a log format with variables from our limit rules' maps and geos for testing...
    log_format gitea
        '$host $remote_addr - $remote_user [$time_local] "$request" '
        '$status $body_bytes_sent '
        '"$http_user_agent" "$http_x_forwarded_for" '
        '"XXX: X:$internal_ips Y:$by_ip_fwd Z:$clouds A:$sitefeature"';

    <<perIP>>

    <<cloudBuckets>>

    limit_req_zone $sitefeature zone=limit_feature:1m rate=1r/s;
  '';

Those two bracketed noweb sections above are inserted from below so that I can explain them a bit better because these nginx maps and geos are a bit confusing.

There is a sidechannel data loss in doing the "LAN-or-not" filtering using the geo module at the top here -- namely what the actual IP they checked against is or where it came from. This would be fine, one can just reference $binary_remote_addr or whatever except that I have a chain of nginx proxies between The Wobserver's Edge Server and The Wobserver itself and so sometimes I need to pull the IP out of X-Forwarded-For via the nginx $http_x_forwarded_for variable.

When I am on my LAN I connect directly to the wobserver's nginx, and then that proxies to gitea, but when any other client connects, the edge server sets up a proxy pass; so either variable is in play.

You can't, it turns out, use a variable in the default value of a geo, it sets the value to the variable name string (as in, literally, $internal_ips), so you have to do this extra nonsense with a pair of maps to dig back to a ternary between those two variables' values.

This is sort of clunky to even reason about, but nginx uses these variable-mapping directives to transform request state in to variables that can be used in the configuration, rather than procedural if/then kind of code.

text#+name: perIP:noweb-ref perIP
# This goes in the top-level http{} block
geo $internal_ips {
    proxy 100.85.220.128/32;   # king-mountain tailscale
    proxy 127.0.0.1/32;
    proxy ::1;

    default 0;
    100.64.0.0/10 1; # Tailscale CGNAT
    73.0.0.0/8 1;    # AS7922 Comcast my choice in the local duopoly
    75.144.0.0/13 1; # AS7922 Comcast my employer's choice
}

That maps internal LAN IPs to 1, and everything else to 0, accounting for proxy directives.

If that address is 0, we try to map back to a proxied IP, otherwise store that false empty value as an invalid value; if the connection does not have a proxy IP, that will be set to "" empty value. And if $http_x_forwarded_for is unset, so will the mapped value(???).

text#+name: perIP:noweb-ref perIP
map $internal_ips $limit_xff {
    0     $http_x_forwarded_for; # came from web; try to extract proxy header
    1     "_"; # came from LAN
}

And then we have constructed a ternary operator. If the request came from the LAN, set the final rate limit key to empty. If the X-Forwarded-For header was set, set it to that; if not that then set it to the remote IP address.

text#+name: perIP:noweb-ref perIP
map $limit_xff $by_ip_fwd {
  default $limit_xff; # limit using the existing value
  "" $binary_remote_addr; # if there was no proxy header limit using remote addr
  "_" ""; # if it came from LAN don't limit.
}

limit_req_zone $by_ip_fwd zone=limit_external_ips:100m rate=1r/s;

Easy! Ha ha ha. So that's how we do LAN exclusion-from-rate-limiting while making sure it still works with the Edge Server.

There is a second filter, now. Alibaba cloud instances were just beating the heck out of my server, this is what caused me so much heart-ache this week that I went and overhauled all of this. So they get slow-rolled first and harder. Eventually I'll add all the clouds, there's very little reason for web traffic from Cloud instances to be worth enough to me to prioritize, but this is good enough to ignore for a while.

text#+name: cloudBuckets:noweb-ref cloudBuckets
# This goes in the top-level http{} block
geo $clouds {
  47.74.0.0/15 ALI;
  47.76.0.0/14 ALI;
  47.80.0.0/13 ALI;
  # add more as necessary
  # maybe with https://github.com/PodderApps/ipcat/blob/main/datacenters.csv
}

limit_req_zone $clouds zone=limit_clouds:100m rate=1r/m;

Hopefully those make enough sense to fix later on.

Two location blocks are used to shunt the more expensive traffic in to the rate limiters and not affect "normal" or at least reasonable viewing and crawling of my source code. The configuration for each location's proxy details need to be shared between the two location, though, because only one location block will be processed for each request. To understand what's happening here please take a look at the location directive's documentation and the generated configuration.

nix:tangle ~/nix/nixos/gitea-ratelimit.nix
  services.nginx.virtualHosts."code.rix.si" = let
    proxyConfig = ''
      proxy_set_header Host $host;
      proxy_set_header X-Real-IP $remote_addr;
      proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
      proxy_set_header X-Forwarded-Proto $scheme;
      proxy_set_header X-Forwarded-Host $http_host;
    '';
  in {
    extraConfig = ''
      # enable to shunt traffic in to its own log file...
      access_log /var/log/nginx/gitea.log gitea;
    '';
    useACMEHost = "fontkeming.fail";
    addSSL = true;

Okay so you can use regular expression capture groups in regexp-matching location directives, and then reference them as variables in your limit_req_zone below! this is rad, you could implement per-blog/per-endpoint rate limits this way, except managing the configurations would be a real pain in the butt.

I guess your nginx would probably fail to start if you added the =limit_req zone=limit_feature= directive to the / location? careful!

nix:tangle ~/nix/nixos/gitea-ratelimit.nix
    locations."~* (?<sitefeature>commit|blame|raw|issues|nixpkgs|watchers|forks)" = {
      proxyPass = config.services.nginx.virtualHosts."code.rix.si".locations."/".proxyPass;
      extraConfig = ''
        limit_req zone=limit_external_ips nodelay;
        limit_req zone=limit_clouds nodelay;
        limit_req zone=limit_feature nodelay;
        limit_req_log_level notice;
      '' + proxyConfig;
    };
    locations."/" = {
      proxyPass = "http://last-bank:80";
      extraConfig = ''
      '' + proxyConfig;
    };

You can't say I didn't warn you...:

nix:tangle ~/nix/nixos/gitea-ratelimit.nix
    locations."=/robots.txt".alias = pkgs.writeTextFile {
      name = "gitea-robots-txt";
      text = ''
        User-agent: *
        Disallow: *
        Disallow: /
        Disallow: /upstreams
        Disallow: /compost
        Disallow: /rrix/*/commit
        Disallow: /rrix/*/rss
        Disallow: /rrix/*/blame
        Disallow: /rrix/*/raw
        Disallow: /rrix/nixpkgs/
      '';
    };
  };
}

Apropos of nothing, the config file verification which nix build tries to do does not work very well, so i'd deploy these changes and the build would say "lol 0 errors" and then fail to start up on the server complaining that there was an illegal directive in the location block or what have you. So there is something in here for one of My Grymoires, a way to spit debug log information out of location headers with a special log_format entry for it... Debugging this is a PITA otherwise heh sorry future self.

See Also

Block AI scrapers with Anubis - Xe Iaso

References

I got tired of this and made a tool to stop them for good. I call it Anubis. Anubis weighs the soul of your connection using a sha256 proof-of-work challenge in order to protect upstream resources from scraper bots. It's a reverse proxy that requires browsers and bots to solve a proof-of-work challenge before they can access your site, just like Hashcash.

Please stop externalizing your costs directly into my face

References

All of my sysadmin friends are dealing with the same problems. I was asking one of them for feedback on a draft of this article and our discussion was interrupted to go deal with a new wave of LLM bots on their own server. Every time I sit down for beers or dinner or to socialize with my sysadmin friends it’s not long before we’re complaining about the bots and asking if the other has cracked the code to getting rid of them once and for all. The desperation in these conversations is palpable.

I fear for the unauthenticated web

References

How long until scrapers start hammering Mastodon servers? Individual websites? Are we going to have to require authentication or JavaScript challenges on every web page from here on out?

All this for what, shitty chat bots? What an awful thing that these companies are doing to the web.

Excerpt from a message I just posted in a #diaspora team internal forum

References

And the best thing of all: they crawl the stupidest pages possible. Recently, both ChatGPT and Amazon were - at the same time - crawling the entire edit history of the wiki. And I mean that - they indexed every single diff on every page for every change ever made. Frequently with spikes of more than 10req/s. Of course, this made MediaWiki and my database server very unhappy, causing load spikes, and effective downtime/slowness for the human users.

If you try to rate-limit them, they’ll just switch to other IPs all the time. If you try to block them by User Agent string, they’ll just switch to a non-bot UA string (no, really). This is literally a DDoS on the entire internet.