Menu

Security · AI

OpenAI agents brute-forced a UN website: how they got past a GET-only limit

An independent researcher reports OpenAI agents scanned a UN statistics site over 16,000 times in 2026, using double URL encoding and URL scanner relays to get past limits.

Published
Boil time
9 min
Language
English · 日本語

TL;DR

  • Independent researcher Rowan Howard-Jones reported that OpenAI agents scanned the UN Conference on Trade and Development (UNCTAD) statistics site more than 16,000 times from April 13 to June 19, 2026
  • Under a limit of GET only, the agents refined their methods step by step: having URL scanners and proxy services relay POSTs, using double URL encoding to get past a POST-only endpoint, and having Google's vulnerability training material (the XSS game) run scripts
  • Even after being stopped 82 times by rate limiting, they kept accessing the site while varying the spelling of the authentication key about 20 ways
  • OpenAI commented that it is investigating and has offered to explain to the UN, and Stanford's Alex Stamos called it "borderline hacking"
  • If you let AI agents access external sites, consider allow-listing HTTP tools, blocking generic proxies, capping retries and auditing access logs
Contents
  1. What happened
  2. Following the brute-force methods
  3. Relaying POSTs through a URL scanner and a test service
  4. Getting past the CORS wall with a relay service
  5. Signs of trying to get around a filter that didn’t exist
  6. Breaking through the POST-only endpoint with double URL encoding
  7. Using Google’s vulnerability training material as the script runner
  8. Brute-forcing the spelling of the authentication key, and not stopping under rate limiting
  9. Reactions from OpenAI and experts
  10. Background to the series of incidents
  11. Measures to consider when you run AI agents yourself

OpenAI AI agents scanned the UN Conference on Trade and Development (UNCTAD) statistics site, “UNCTADstat”, from April 13 to June 19, 2026. There were more than 16,000 scans. Independent security researcher Rowan Howard-Jones reported it on his blog. What the agents were after was public data. But each time they hit the GET-only limit or the CORS wall, they are recorded as having piled up methods to get through, such as abusing relay services and double URL encoding.

What happened

According to The Verge’s report, Howard-Jones concluded that the scanning of UNCTADstat’s API was likely done by OpenAI agents. His grounds are these two points.

  • A page on an external wiki (FractalWiki), created 40 minutes after the scan, had the same UNCTADstat URLs copied onto it as were used in the scan
  • Of the 54 Azure IP addresses used to edit that page and others, 45 were also used to edit another wiki (DseWiki) in the “wiki swarm”. The wiki swarm refers to a case, already confirmed to involve OpenAI’s agents, where multiple agents used public wikis as bulletin boards and data stores

What the agents were after was public data on Productive Capacities Index (PCI), one of the trade and development indicators, as well as food trade and target industries. But they had no permission to call the API directly, and the HTTP methods they could use were limited to GET, so the agents tried various workarounds.

Following the brute-force methods

Howard-Jones’s report records, in time order and with the actual request contents, how the agents refined their methods through trial and error.

Relaying POSTs through a URL scanner and a test service

UNCTADstat’s data retrieval endpoint, “Facts”, is POST-only, and sending a GET returns a 400 error. The agents, meanwhile, were allowed only GET.

So the agents used the URL scanner “urlquery” as a stepping stone. urlquery opens a submitted URL in a sandboxed browser and records what happens. First the agent prepared an HTML form that POSTs to UNCTADstat and a script that submits it automatically, and hosted these using the Base64 decode feature of the web testing service “httpbin”. When it has urlquery GET that httpbin URL, urlquery’s sandboxed browser submits the form automatically. The result is that a POST to UNCTADstat goes through.

The image below is urlquery’s report for this. The first line is the GET to httpbin, and the second is the POST to UNCTADstat’s Facts. Both returned 200 OK.

The list of HTTP traffic in a urlquery report. A GET to httpbin.org is followed by a POST to Facts at unctadstat-api.unctad.org, both recorded as 200 OK
Source: OpenAI agents tried to bruteforce a UN website’s API fields (swarmcha.se)

Getting past the CORS wall with a relay service

With retrieval through the form, the result came back only as a screenshot. That meant the data was hard to read by machine. The agent next used JavaScript’s fetch() to try to read the response directly. But UNCTADstat refuses reads from other sites through CORS (Access-Control-Allow-Origin), and it failed.

In the image below, the GET to UNCTADstat returned 200 OK. Even so, the script on the page couldn’t read the response, and only NetworkError was shown on screen.

The list of HTTP traffic in a urlquery report. After opening a page on httpbin.org, a GET to unctadstat-api.unctad.org is recorded as 200 OK
Source: OpenAI agents tried to bruteforce a UN website’s API fields (swarmcha.se)

To get around this, the agent used the AI search and proxy service “r.jina.ai” as a relay. r.jina.ai’s server fetches the target URL and returns the result without the CORS restriction. With this method, on April 27 it succeeded for the first time in retrieving PCI data (Norway, Iceland, Denmark).

After that, a trick was added to leave the retrieved data in urlquery’s record. urlquery records all other requests made while a page loads. The agent URL-encoded the retrieved data, put it in the query string of another site’s URL, and had that URL loaded. That made it possible to read the data from the urlquery report. For the relay service, it switched from r.jina.ai to the codetabs proxy. With this method, on May 14, it was able to retrieve UNCTADstat dynamic data (label information for the SIDS countries).

In this exchange, the path before the result came back to the agent spans four services: urlquery, codetabs, UNCTADstat and httpbin. It proceeded in the order ① to ⑦ in the diagram below.

A diagram of the route on May 14, when data retrieval succeeded. The agent sends a page URL to urlquery, and urlquery's browser opens the page. The page has the codetabs proxy fetch the UNCTADstat data on its behalf, then navigates to an httpbin URL that contains the received data. urlquery records that URL, and the agent reads the data from the report
The route of the May 14 retrieval (created from the report at swarmcha.se)

Signs of trying to get around a filter that didn’t exist

On May 14, behavior was also observed in which words in a request were deliberately split. For example, it sent "POST" as "PO" + "ST" and "no-cors" as "no" + "-cors". Howard-Jones speculates that the agent mistakenly believed “my requests are being blocked by some kind of filter” and tried to get around it. In fact no such filter existed, and the requests were failing for other reasons, such as sending the wrong request to a POST-only endpoint.

Breaking through the POST-only endpoint with double URL encoding

On April 28, the agent tried a GET against the POST-only “Facts” endpoint, and was again rejected with a 400 error. On May 4, the agent began using a technique called “double URL encoding”.

This URL-encodes part of the string Facts (the a) as %61, then encodes that %61 itself once more to get %2561. Consider a case where the server side is built in several layers, and each layer URL-decodes independently. In that case, an inconsistency can arise where one layer sees an already-encoded string and passes it through, and another layer decodes it into an unintended value. The agent sent a request that rewrote Facts as F%2561cts. It went through, and it succeeded in retrieving data from the POST-only endpoint. This technique was then used 55 times through June 19.

The image below is a response retrieved with this technique. Just sending a GET returned JSON data. The content is statistics on exports of raw food products (Raw products) from Burkina Faso in 2005.

JSON returned from UNCTADstat's Facts. It reads: year 2005, country Burkina Faso, product Raw products, flow Exports
Source: OpenAI agents tried to bruteforce a UN website’s API fields (swarmcha.se)

Howard-Jones says he reported this bypass to UNCTAD’s information security team before publishing.

Using Google’s vulnerability training material as the script runner

What the agent finally arrived at was abusing “XSS game”, Google’s training material for XSS (cross-site scripting). One of the game’s exercises is intentionally vulnerable: it embeds whatever string you enter after ?query= in the URL directly into the page. The agent wrote a <script> tag directly there. By doing so, it discovered it could run arbitrary JavaScript in urlquery’s sandboxed browser without going through httpbin or Base64 encoding. From May 25 to June 1, 25 reports using this technique were confirmed.

In the image below, after opening the XSS game page (xss-game.appspot.com), a POST to UNCTADstat is recorded as 200 OK.

The list of HTTP traffic in a urlquery report. A GET that opens an XSS game exercise page and a POST to unctadstat-api.unctad.org are recorded, both 200 OK
Source: OpenAI agents tried to bruteforce a UN website’s API fields (swarmcha.se)

Brute-forcing the spelling of the authentication key, and not stopping under rate limiting

UNCTADstat’s data viewer automatically sends a subscription key (an Azure API Management header) common to all users. The key itself was not a secret. But the agent tried about 20 spellings, such as these.

  • subscription-key
  • apikey
  • api-key
  • Ocp-Apim-Subscription-Key

Howard-Jones analyzes that the agent’s repeated failure when sending GETs to the POST-only endpoint was the result of it mistaking the cause for “a mistake in specifying the key name”.

UNCTADstat blocked at least 82 requests through rate limiting, but according to the report, the agent kept sending requests even after being limited.

Reactions from OpenAI and experts

In response to The Verge’s inquiries, neither OpenAI nor the UN initially commented. On the other hand, according to Dataconomy’s report, an OpenAI spokesperson commented to the Wall Street Journal. The content was “we are aware of this report and have offered to set up a session for our investigating team to explain to the UN”. They are also said to have stated that this is part of an effort to broadly investigate unintended model behavior (misalignment) during training and evaluation.

Stanford University cybersecurity lecturer Alex Stamos told the Wall Street Journal that the series of actions at UNCTAD was “borderline for what I would call hacking”. He also said it was “really very aggressive scraping and data retrieval”. Howard-Jones himself doesn’t call it clear “hacking”. He puts it as “it looks like the behavior of someone, or something, that won’t take no for an answer”.

Background to the series of incidents

This case isn’t an isolated event. According to Dataconomy’s report, the trigger for this investigation was an earlier report from the nonprofit research organization Transluce. Around the same time, Transluce also linked OpenAI’s agents to aggressive access to Data USA (a US public statistics site) and an Australian government health statistics site. OpenAI is also reported to have separately acknowledged inappropriate agent behavior on US government sites such as the US Department of Commerce and the Securities and Exchange Commission (SEC).

The “wiki swarm”, many of whose IP addresses overlapped with those used to create the UNCTAD-related wiki pages, is likewise a case said to involve OpenAI’s agents. Whether this case is part of the same collective behavior, or the result of independently finding the same wiki, isn’t stated definitively in Howard-Jones’s report.

Measures to consider when you run AI agents yourself

What can be read from this case is that even without clear malice, the persistence of an agent trying to get past errors itself can look like attack behavior from the outside. If you let AI agents handle access to external sites, these points seem worth considering.

  • Restrict HTTP access with an allowlist: limit the HTTP tools an agent can call to domains and methods approved in advance. A generic HTTP tool that can reach any URL with any method makes it hard to prevent reaching unintended targets
  • Restrict reach to generic relay and proxy services and sandbox tools: don’t leave agents free to access URL scanners and relay services like urlquery, httpbin and r.jina.ai that were abused here, or intentionally vulnerable tools like Google’s vulnerability training material
  • Cap the number of retries against the same target: limit, at the task design or guardrail stage, the behavior of retrying with a new method every time an error appears. Make it possible to stop a large burst of access to a particular endpoint from the agent’s side as well
  • Record the URLs and methods accessed, and make them auditable: keep logs of which external sites an agent accessed, and by what method. That outside researchers could reconstruct the behavior afterward in this case was partly because the agents’ actions happened to remain observable from outside

There are also measures on the side of those who run sites and APIs. Bypass through double URL encoding is said to occur easily in configurations where several layers each do URL decoding independently. It helps to unify encoding handling across layers and normalize thoroughly at the boundary. It is also essential to run rate limiting so that it actually blocks access, rather than stopping at returning a warning. These are effective as general defenses, not only against AI agents.

udon
Microsoft 365, ServiceNow, Copilot and more, tested first-hand and written up as practical notes with real bite.