HAProxy Rate Limiting for AI Crawlers and Runaway Agents

Rate-limit AI crawlers, runaway agents and API keys with HAProxy stick tables. A tested config with tarpits, 429s and Retry-After headers.

HAProxy Rate Limiting for AI Crawlers and Runaway Agents

If you run anything public (a Gitea instance, a docs site, a wiki, a small API), you've probably watched this happen: CPU pinned, database connections maxed, and the access log is a wall of GPTBot, Bytespider and a dozen user agents you've never heard of, all politely requesting every commit diff you've ever pushed. Then, on the other side of your stack, an AI agent someone wired to your API gets a 404, retries, gets another 404, and keeps going all night.

Blocking all of it is tempting. It's also the wrong move for most people. Block every AI crawler and you vanish from the AI search answers your users now rely on. Block the agent and you break a customer's integration. What you actually want is a budget: everyone gets a fair number of requests, and whoever goes past it gets slowed down, not your database.

HAProxy does this with stick tables, and it does it before a single request reaches your app. Below is the config I tested on HAProxy 3.2.25, with four rules you can copy.

How Stick Tables Work (In One Minute)

A stick table is an in-memory key-value store inside HAProxy. The key is something about the request (a client IP, a hashed API key). The values are counters HAProxy maintains for you, like "requests in the last 10 seconds" or "4xx errors in the last minute".

You track a request into a table with track-sc0 / track-sc1 (two independent counters per request), then read the counters back with sc_http_req_rate(0) or sc_http_err_rate(0) and act on them. There are no external services and no Redis. Entries expire on their own when a client goes quiet.

The Config

Here's the complete haproxy.cfg. Adjust the numbers, the bind port and the backend address to match your setup, and leave the structure alone.

global
    stats socket /var/run/haproxy/admin.sock mode 660 level admin

defaults
    mode http
    timeout connect 5s
    timeout client  30s
    timeout server  30s
    timeout tarpit  15s

# One stick table per key type. These backends only hold tables.
backend st_ip
    stick-table type ipv6 size 1m expire 10m store http_req_rate(10s),http_err_rate(1m)

backend st_key
    stick-table type string len 64 size 100k expire 1h store http_req_rate(1m)

frontend web
    bind :8080

    # Behind a reverse proxy, src is the proxy. Trust its header, only from it.
    http-request set-src req.hdr_ip(X-Real-IP) if { src 172.16.0.0/12 127.0.0.1 } { req.hdr(X-Real-IP) -m found }

    acl ai_bot  req.fhdr(user-agent) -i -m sub GPTBot ClaudeBot CCBot Bytespider PerplexityBot Amazonbot meta-externalagent
    acl has_key req.hdr(Authorization) -m found

    http-request track-sc0 src table st_ip
    http-request track-sc1 req.hdr(Authorization),sha2(256),hex table st_key if has_key

    # 1. Known AI crawlers: 20 requests per 10s, then tarpit.
    http-request tarpit deny_status 429 if ai_bot { sc_http_req_rate(0) gt 20 }

    # 2. Runaway agents: 30+ errors a minute means it's looping. Cool it off.
    http-request return status 429 content-type application/json string '{"error":"too_many_errors"}' hdr Retry-After 60 if { sc_http_err_rate(0) gt 30 }

    # 3. API keys: 120 requests per minute each.
    http-request return status 429 content-type application/json string '{"error":"rate_limited"}' hdr Retry-After 30 if has_key { sc_http_req_rate(1) gt 120 }

    # 4. Hard ceiling for every IP, API key or not: 100 requests per 10s.
    http-request return status 429 content-type text/plain string "Slow down" hdr Retry-After 10 if { sc_http_req_rate(0) gt 100 }

    default_backend app

backend app
    server app1 172.17.0.1:3000 check

Before you reload, validate it with haproxy -c -f haproxy.cfg. Here's what each rule does:

RuleKeyed onBudgetOver budget
AI crawlersClient IP (bot user agents only)20 req / 10sHeld 15s, then 429
Looping agentsClient IP30 errors / min429, Retry-After 60
API keysSHA-256 of the key120 req / min429 JSON, Retry-After 30
Per-IP ceilingClient IP (everyone)100 req / 10s429, Retry-After 10

When I ran it with curl loops, it did exactly what the table says: 100 requests went through and request 101 got a 429, the 121st call with one API key (paced to stay under the IP ceiling) got {"error":"rate_limited"} with retry-after: 30, and GPTBot's 21st request sat for the full tarpit timeout before getting its 429.

The Details That Matter

The tarpit is for bots, not for people. http-request tarpit holds the connection open for timeout tarpit before answering. A well-behaved crawler slows down because every extra request now costs it 15 seconds. HAProxy holds those connections very cheaply, but they still count against maxconn, so don't tarpit general traffic.

Hash the API keys. Stick tables are readable through the runtime API. Tracking req.hdr(Authorization),sha2(256),hex means anyone who dumps the table sees hashes, not live credentials. Hashing the whole header also means Bearer abc and abc count as different keys, which is fine as long as your clients are consistent.

Your own 429s don't feed the error counter. In testing, the 429s from http-request return never showed up in http_err_rate, while real 404s from the app did. That's what you want: a looping agent gets cooled off, and the moment it stops hammering a missing endpoint, its error rate decays and it's back in business. There's no permanent ban to clean up by hand.

User-agent matching is an honor system. GPTBot and ClaudeBot identify themselves, and many scrapers don't. That's why rule 4 exists and applies to everyone: a bot lying about its user agent still hits the per-IP ceiling, and so does one that sends a made-up Authorization header hoping for a fresh key budget. HAProxy can't tell a real key from a fake one, so never let the header alone lift a limit. Your app still validates keys, and rule 3 only exists to keep one real key from taking more than its share. Note req.fhdr rather than req.hdr. User agents contain commas, and req.hdr splits on them.

Watching It Live

The stats socket line exposes HAProxy's runtime API. Add - ./run:/var/run/haproxy under the HAProxy service's volumes: in your Docker Compose file, and you can query it from the host (apt install socat first):

# Who's close to the limit right now?
echo "show table st_ip data.http_req_rate gt 50" | socat stdio ./run/admin.sock

# Unblock a false positive immediately
echo "clear table st_ip key 203.0.113.5" | socat stdio ./run/admin.sock

Running It on Elestio

On Elestio, HAProxy runs as a fully managed service starting at $11/month for the VM (HAProxy itself has no license fees), with updates, backups and monitoring handled for you. Like every Elestio service, it runs in Docker Compose under /opt/app/ on your VM. Edit haproxy.cfg there, validate it, then run docker compose restart from that folder.

Elestio's standard setup puts an Nginx reverse proxy in front of the container, and it passes the real client IP in X-Real-IP. That's the header the set-src line trusts. Without that line, every client would appear as the Docker bridge IP and your whole user base would share one budget. Send a few requests and run show table st_ip: real client IPs mean you're set.

Troubleshooting

Everyone gets rate-limited at once. Your src is a proxy. Check show table st_ip. If the only key is something like ::ffff:172.17.0.1, the set-src line isn't firing. Behind Cloudflare, use CF-Connecting-IP and trust Cloudflare's published IP ranges instead of 172.16.0.0/12.

unknown converter 'sha2' on startup. Your HAProxy build lacks OpenSSL. The official Docker images include it, but a minimal source build won't.

Keys look like ::ffff:203.0.113.5. That's normal. An ipv6 table stores IPv4 addresses in mapped form. clear table st_ip key 203.0.113.5 still works.

Permission denied on the admin socket. The official image runs as the haproxy user (UID 99), so run chown 99:99 ./run on the mounted folder.

Counters reset after a restart. Stick tables live in memory. If you run two HAProxy nodes or need counters to survive reloads, add a peers section so tables sync between processes.

Start with generous budgets, run show table st_ip data.http_req_rate gt 50 a few times over a day to see what normal looks like, then tighten. Real numbers from your own traffic beat any default in this post, mine included.

Thanks for reading ❤️ See you in the next one 👋