Skip to content
AI and agents

Proxies for AI agents that browse

Agent traffic is neither human nor crawler, and its shape gives it away before any header is read. Identity per task, bounded retries, real block data.

An agent that fetches web pages on a user’s behalf produces traffic that matches no profile a site has seen before. It is not a person browsing and it is not a crawler grinding through a sitemap. It arrives at one deep URL with no referrer, reads it, sometimes follows one link, and stops. That shape is what gets it refused, and no proxy fixes a shape problem.

This guide covers what agent traffic looks like from the far end, what your choice of fetch layer commits you to before a single header is sent, and the bounds you need because a person is waiting for the answer.

#The traffic shape is the tell

Start with what a site can see without inspecting a single byte of your request. Three kinds of client, three different shapes on the wire.

Property Human session Bulk crawler Task-driven agent
Entry point Home page, or a search result A seed list or sitemap A deep URL, supplied by the model or the user
Referer Present from the second page onwards Usually absent Absent
Static assets CSS, fonts, images, all fetched Usually skipped Skipped, unless a browser is driving
Request spacing Irregular, seconds apart, reading time between Regular, tuned, sustained A burst, then nothing
Session length Many pages over minutes Indefinite One to a handful of pages, then done
Return visits Yes, carrying cookies Endlessly, often from the same address Rarely the same address twice
Navigation path Links followed from a page it already has Links extracted from pages it fetched Often none: the URL came from outside the site

The last row is the interesting one. A human and a crawler both arrive at a deep page because something they already fetched pointed at it. An agent arrives because a model produced the URL, and there is no preceding request to justify it. The result reads like the middle of a session that has no beginning: no landing page, no Referer (defined in RFC 9110, section 10.1.3), no asset fetches, no second visit.

Then it stops. A crawler that gets refused keeps knocking, which is a pattern; an agent that gets refused disappears, which is not. Rate limits are built to catch sustained volume, and agent traffic has almost none — it is a handful of requests that look wrong rather than a flood that looks normal. Whether a given anti-bot system acts on any of this we have not measured, and neither has anyone else who publishes about it honestly. What is certain is that these differences are visible without any inspection of the request itself.

#Which handshake your agent inherits

Before your code chooses a single header, the fetch layer has already committed you to a TLS handshake. Two of our own studies bear directly on this, and together they say something counter-intuitive.

From our ClientHello study: twelve consecutive connections from the same Chromium build produced twelve distinct extension orders and twelve distinct JA3-style hashes. The extension set never changed; the order did, on every connection. Run the same test against libraries and it inverts — curl and Go each produced exactly one distinct extension order across three consecutive connections. None of the eight libraries tested sent GREASE, which every mainstream browser does.

The practical reading for an agent is that a stable handshake is now the liability. Your HTTP client does not look unusual so much as it looks identical, connection after connection, in a way no current browser does.

The header layer says the same thing. Our header study recorded a Chromium navigation sending 11 headers against a median of 4 across nine libraries, and not one library sent the Fetch Metadata set — Sec-Fetch-Site, Sec-Fetch-Mode, Sec-Fetch-User, Sec-Fetch-Dest, specified by W3C. Node’s global fetch sent a single sec-fetch-mode: cors and nothing else from the set.

Fetch layer Handshake it emits Header set What you still control
A driven headless browser The browser’s own, with GREASE and shuffled extension order The full browser set, Fetch Metadata included Automation defaults, viewport, timing, cost
A plain HTTP client The library’s, identical every connection The library’s small default set Header values; rarely their order; never the handshake
An HTTP client with a browser-impersonating TLS stack A reproduction of a browser handshake Whatever you send Everything above the handshake

The agent-specific trap sits between these rows. Many agent frameworks use both: a browser for pages that need JavaScript, and a plain HTTP client for everything else, chosen per URL. That is efficient and it means one task presents two different identities to the same site, sometimes within seconds. If a site correlates the two, the correlation is not subtle — one connection carries a browser handshake and eleven headers, the next carries a library handshake and four.

Pick one identity per task and keep it. If that means rendering a page you did not need to render, the consistency is usually worth more than the saved work.

#Identity per task, not per request

A rotating endpoint gives you a new exit address per request. That is the right default for a crawler collecting independent pages and the wrong default for an agent. An agent doing one coherent job — read a product page, then its specification sheet, then its reviews — is one visitor as far as the site is concerned, and rotating between those three fetches breaks every assumption the site holds: the cookie it set, the rate-limit bucket it opened, any per-address cache it warmed.

The unit of identity for an agent is the task. Hold a sticky session for its duration, and derive the session token from the task so a retry after a crash lands on the same exit:

import hashlib
import httpx

GATEWAY = "gateway.example:8000"

def proxy_for_task(task_id: str, user: str, password: str) -> str:
    # One exit address for the whole task. Deriving the token from the
    # task id means a resumed task reuses the exit it started on.
    # The username syntax that carries a session token is provider-
    # specific; check yours rather than copying this one.
    token = hashlib.sha256(task_id.encode()).hexdigest()[:16]
    return f"http://{user}-session-{token}:{password}@{GATEWAY}"

class TaskFetcher:
    """One client, one connection pool, one exit, for one task."""

    def __init__(self, task_id: str, user: str, password: str):
        self.task_id = task_id
        self.client = httpx.Client(
            proxy=proxy_for_task(task_id, user, password),
            timeout=httpx.Timeout(10.0, connect=5.0),
            follow_redirects=True,
            headers={
                "user-agent": "example-agent/1.0 (+https://example.com/agent)",
            },
        )

    def close(self) -> None:
        self.client.close()

Two failures to watch for. The first is the one above run backwards: a long-lived agent process that creates a single shared client at start-up pins every task, for every user, to one exit address. Key the pool by task, not by process.

The second is time. A sticky session has a lifetime set by the provider, and when a task outruns it the address changes underneath you, mid-task, with no error. We do not publish a figure for that lifetime because it differs per provider and we hold no sourced value for any of them — read your provider’s documentation, then verify it by holding a session and watching the reported address, rather than trusting the number.

The distinction to keep straight: rotation is not a defence against fingerprinting. Every finding in both studies above is produced by your client and travels unchanged through any proxy, from any country, on any address category. Rotating faster changes the address and nothing else.

#Fetching to answer a question is not crawling

Publishers already draw this line, and they draw it sharply. Our survey of 46 readable robots.txt files found that the same operator gets very different treatment depending on what the agent is for: GPTBot was blocked by 16 of those sites and OAI-SearchBot by 7; ClaudeBot by 22 and Claude-User by 12. Both operators saw their user-triggered agent blocked roughly half as often as their bulk crawler.

The survey also found the highest counts of partial rules — specific paths disallowed rather than the whole site — attached to the user-triggered agents. Those are considered rules. Somebody sat down and decided that fetching a page to answer a live question is a different act from harvesting the site, and wrote it down.

That is the case for declaring yourself. An agent acting for a named user, identifying itself with a product name and a URL that explains what it does, is asking a question the site’s operator has a way to answer:

# The operator's decision, stated in the file. RFC 9309 defines the format.
User-agent: ExampleAgent
Disallow: /checkout/
Disallow: /account/
# Everything else may be fetched to answer a user's question.

User-agent: ExampleCrawler
Disallow: /

Three details of robots.txt parsing cause most mistakes, and all three are in RFC 9309. Consecutive User-agent: lines share the rules that follow them, so a parser that starts a new group on each line misreads the large files. Only the most specific matching group applies — if the file names your agent, the wildcard group is irrelevant to you, including its Allow lines. And the file is a resource with its own rules: section 2.4 says a crawler should not use a cached copy for more than 24 hours unless the file is unreachable, section 2.5 requires parsers to handle at least 500 kibibytes, and section 2.3.1.4 says a 5xx response makes the file unreachable, at which point the crawler must assume complete disallow. That last one catches agents constantly: a robots.txt fetch that times out or errors is not permission, and the spec draws the opposite conclusion from a 4xx, which section 2.3.1.3 treats as the file being unavailable and the server open.

The honest position is that these are two different decisions and you should make them separately. An agent that names itself accepts that some sites will refuse it, and gets to point at the rule when they do. An agent that presents a browser user agent through a residential address has chosen not to be identifiable. That may be defensible for some purposes and it is a choice, not a default — it should not happen merely because a library shipped a browser string in its configuration.

#Bounded retries, because someone is waiting

This is where agent fetching diverges most from crawling. A crawler that fails can retry tomorrow; nothing is waiting. An agent has a person at the other end, and a stalled fetch is worse for them than a fast honest failure, because a failure gives the model something to act on and a stall gives it nothing.

So the budget is a wall-clock deadline for the whole task, shared by every fetch inside it — not a retry count per request.

import time
import httpx

class Budget:
    """A deadline and a byte allowance, shared by every fetch in one task."""

    def __init__(self, seconds: float, max_attempts: int, max_bytes: int):
        self.deadline = time.monotonic() + seconds
        self.max_attempts = max_attempts
        self.bytes_left = max_bytes

    def remaining(self) -> float:
        return max(0.0, self.deadline - time.monotonic())

def fetch(client, url, budget, parse_retry_after):
    for attempt in range(budget.max_attempts):
        left = budget.remaining()
        if left <= 0:
            raise TimeoutError(f"task budget exhausted before {url}")

        try:
            # Never let one request outlive what the task has left.
            r = client.get(url, timeout=min(left, 10.0))
        except httpx.TransportError:
            wait = min(2 ** attempt, budget.remaining())
            if wait <= 0:
                break
            time.sleep(wait)
            continue

        if r.status_code == 403:
            # A decision about you, not a transient fault. Retrying the
            # same request unchanged cannot succeed. Report it upwards.
            raise PermissionError(f"403 from {url} on attempt {attempt + 1}")

        if r.status_code == 429 or 500 <= r.status_code < 600:
            # Retry-After is defined in RFC 9110, section 10.2.3.
            wait = parse_retry_after(r.headers.get("retry-after")) or 2 ** attempt
            if wait >= budget.remaining():
                break          # an honest failure now beats a stall later
            time.sleep(wait)
            continue

        budget.bytes_left -= len(r.content)
        return r

    raise RuntimeError(f"giving up on {url} after {budget.max_attempts} attempts")

The numbers in that code are choices, not findings. What matters is the structure:

  • Every timeout is the smaller of your limit and what remains. A ten-second read timeout inside a task with three seconds left is a stall by construction.
  • Never sleep past the deadline. If the wait a server asks for exceeds the budget, fail immediately and say so. Sleeping to the end of the budget wastes the whole remainder to arrive at the same failure.
  • Do not retry a 403 unchanged. It is a decision about the request you sent, so an identical request produces the identical answer. A 407 is your own proxy refusing your credentials and never reaches the site at all. Both differ from a 429, and the three are easy to confuse: see 403 vs 407 vs 429.
  • Return the status code to the model, not “fetch failed”. An agent told the fetch failed retries the same URL. An agent told it received a 403 after two attempts can try a different source, which is the behaviour you actually want.

#Bound the walk: domains, depth, bytes

An agent that can follow links will follow them, and the failure is not usually dramatic. It is a task that was meant to read one page quietly reading forty, spending the budget, spending metered traffic, and returning a worse answer than it would have from the first page alone.

Bound it explicitly, and make the reason for each refusal something you can log:

const limits = {
  allowedHosts: new Set(['docs.example.com', 'api.example.com']),
  maxDepth: 2,
  maxPages: 12,
  maxBytes: 5 * 1024 * 1024,
};

// Returns null to admit the URL, or a string naming the bound that stopped it.
function admit(rawUrl, depth, state) {
  let u;
  try {
    u = new URL(rawUrl, state.baseUrl);
  } catch {
    return 'unparseable';
  }

  // Models emit file: and data: URLs. Refuse every scheme but these two.
  if (u.protocol !== 'https:' && u.protocol !== 'http:') return 'scheme';
  if (!limits.allowedHosts.has(u.hostname))              return 'off-domain';
  if (depth > limits.maxDepth)                           return 'depth';
  if (state.pages >= limits.maxPages)                    return 'page budget';
  if (state.bytes >= limits.maxBytes)                    return 'byte budget';

  u.hash = '';                       // #section is the same document
  if (state.seen.has(u.href))        return 'already fetched';

  return null;
}

Four notes on getting this right.

Check the host after redirects, not before. Admitting a URL on your allow-list and then following a redirect off it is the standard way an agent leaves its own boundary. Re-run the host check on the final URL, and treat a redirect that crosses the boundary as a refusal rather than a success.

Cap the body while reading it, not after. A byte budget checked against Content-Length is checked against a number the server chose, and it is absent on a chunked response. Stream, count, and abort at the limit.

Count refused fetches against the budget too. On metered bandwidth a refusal is billable traffic, and a challenge page is a real response body. An agent that retries its way through a hundred blocks has spent the same as one that read a hundred pages.

Scheme filtering is not paranoia. A model producing URLs will eventually produce file:///etc/passwd, and an agent that fetches whatever string it is handed will read it. Allow two schemes and refuse the rest.

#Log enough to explain the failure

The question you will be asked about a failed task is whether it was blocked or whether it ran out of time. Those need opposite fixes, and you cannot answer it after the fact unless you recorded it at the time.

Field Why you need it later
Task id and the session token derived from it Ties every fetch to one exit; without it a rotation is indistinguishable from a block
Final URL, after redirects The URL you requested is often not the one you got
Status code, verbatim 403, 407 and 429 call for three different responses
Which fetch layer served it A browser render and a library fetch are different identities on the wire
Attempt number, and budget remaining at that attempt Separates “refused” from “ran out of time”
Response byte length Finds the fetch that ate the budget, and catches truncation
The robots.txt verdict at fetch time Policies change; what you were told then is not recoverable now

The last row is worth the storage. A robots.txt file is a statement made on a particular day, and the survey above is a snapshot for exactly that reason. Recording the verdict you acted on turns an argument about what a site permitted into a lookup.

One habit carries over from ordinary proxy debugging and applies here more than anywhere: before changing any proxy setting, establish whether the request reached the destination at all. A status code from the site means the tunnel worked and the site refused you, and no amount of proxy configuration will change that answer.

#Limits

  • We have not measured how any site classifies agent traffic. Both studies cited here record what clients emit, not what servers do with it. The traffic-shape table describes what is visible on the wire; treating it as evidence that any given site acts on those signals would be a claim we cannot support.
  • The header and TLS figures are one machine, one date, one Chromium build. The header captures were taken on 29 August 2026 and the ClientHello captures on 2 September 2026, both on macOS, both over HTTP/1.1 or a single TLS flight. Defaults change between releases.
  • The robots.txt survey is stated policy on one day. It says nothing about whether crawlers obey those rules, whether sites enforce them, or whether the same files say the same thing now.
  • We have not tested any agent framework end to end. Nothing here describes the defaults of any specific agent product, and the mixed-identity problem is inferred from how such systems are built rather than observed in one.
  • No provider figures. Sticky session lifetimes, username syntax, rotation behaviour and bandwidth pricing all differ by provider, and we hold no sourced value for any of them. Where a number would go here, read your provider’s documentation and then verify it yourself.
  • The code is illustrative. Every limit in it — attempts, depths, page counts, byte caps — is an example value, not a recommendation derived from a measurement.

Frequently asked questions

Why does an AI agent get blocked when a scraper on the same proxy does not?
Because the traffic shape differs. An agent arrives at a deep URL with no referrer, fetches no assets, follows no navigation path, and then stops. A crawler sustains volume that rate limits are built to measure; an agent sends a handful of requests that look out of place instead. Changing the address does not change any of that.
Should an agent rotate its proxy address on every request?
No. Rotate per task, not per request. An agent doing one coherent job is one visitor to the site, and rotating mid-task discards the cookie it was given, the rate-limit bucket it opened and any per-address state the site holds. Derive a sticky session token from the task id so a resumed task reuses the same exit.
Is a headless browser better than an HTTP client for agent fetching?
It inherits a browser handshake, which matters. Our ClientHello study captured twelve distinct extension orders across twelve Chromium connections, while curl and Go each produced one order every time. Our header study recorded 11 browser headers against a median of 4 for libraries. The cost is compute and speed, so decide per task rather than globally.
Should my agent identify itself in the User-Agent header?
It is a defensible position and some operators write rules for it. Our survey of 46 robots.txt files found GPTBot blocked by 16 sites and OAI-SearchBot by 7, ClaudeBot by 22 and Claude-User by 12. Publishers block bulk crawlers roughly twice as often as agents fetching to answer a live question.
How many times should an agent retry a failed fetch?
Bound the task by wall-clock time rather than by attempt count, because a person is waiting. Make every request timeout the smaller of your limit and what remains of the task budget, and never sleep past the deadline. Do not retry a 403 unchanged; it is a decision about your request, not a transient fault.

Sources

  1. RFC 9309: Robots Exclusion Protocol — group matching, the 24-hour cache limit and the 500 KiB parsing minimum
  2. RFC 9110: HTTP Semantics — the Referer field (10.1.3) and Retry-After (10.2.3)
  3. W3C Fetch Metadata Request Headers: the Sec-Fetch-Site, -Mode, -User and -Dest set

Read this page as Markdown · Quote it freely under CC BY 4.0 with a link back.