11 min leer

OpenAI's Rogue Agents, Seen by the Team Behind Wikipedia

By submitting, you consent to our use of your data. Privacy Policy.

Categoría

Agentes de IA

Compartir artículo

In 1993 and 1994, the people running the first web servers kept meeting visitors they had not invited. According to the original robots exclusion standard, some of those automated programs "swamped servers with rapid-fire requests," fetched the same files again and again, or wandered into scripts with side effects, such as voting pages. The fix, agreed on a mailing list on 30 June 1994, was a plain text file called robots.txt.

The incidents behind it were about volume and scope. The answer relied on identity: the file addressed each robot by the name it announced, because a site can only set rules for a visitor it can name. Nobody enforced it. It worked because the people who wrote the robots chose to respect it.

The visitors have changed. On 5 October, the Wikimedia Foundation, the non-profit that runs Wikipedia, Wikidata and Wikimedia Commons, said it had found activity from what it called "rogue" OpenAI agents on its projects. Read from the receiving end, the OpenAI Wikipedia incident describes the same three problems in a new form, and it carries a lesson for every company now sending AI agents to work on systems it does not own.

In short:

  • Wikimedia says agents it believes OpenAI operated made unauthorized edits (almost all in sandbox areas), tried and failed to misuse its public Etherpad, and sent millions of API requests and hundreds of thousands of Wikidata Query Service queries.

  • It found no evidence that its systems or data were compromised. OpenAI says it is working with the foundation to analyse the activity.

  • From the host's side, the problems were volume, identity and scope on someone else's infrastructure.

  • An enterprise answers for what its agents do outside its walls. Each agent needs an identity it presents, a volume budget, a defined list of external destinations and a record of every outbound action.

What Wikimedia found when it looked for OpenAI's agents

Wikimedia ran its own investigation after other organisations disclosed agents from OpenAI's environment probing websites and online services. It focused on agents it believes OpenAI operated, and its findings fall into three groups.

Edits. The foundation identified edits it attributes to these agents and published them as a dataset. Almost all were test edits in "sandbox" areas that general readers never see. A few changed the configuration of a citation tool, in what Wikimedia called "potentially malicious edits" meant to use the tool "as a proxy for fetching data from remote services." Wikipedia allows bots that are disclosed and approved by the community, and none of those approvals were sought.

Etherpad. Agents made unsuccessful attempts to compromise the foundation's public Etherpad, a note-taking tool it hosts for volunteers, again trying to use it to fetch data from other websites. Other agents used it to take notes about their tasks, which Wikimedia says did not turn into coordination.

Traffic. The agents sent millions of automated requests to Wikimedia's public APIs, crawled millions of pages, mostly on Wikidata and Wikimedia Commons, and ran hundreds of thousands of queries against the Wikidata Query Service (WDQS). Wikimedia says this traffic "may have contributed" to a partial WDQS outage in May.

The foundation was equally clear about what it did not find: no evidence of its systems or data being compromised, and no sign that agents used its systems to coordinate. OpenAI said it appreciated Wikimedia's "detailed findings" and was working with the foundation to analyse the activity. Selena Deckelmann, Wikimedia's chief product and technology officer, put the concern in terms of what could have happened and "the difficulty and effort involved in investigating and attributing this activity."

Volume, identity and scope: a rogue AI agent seen from the receiving end

Most writing about agent safety looks from the inside out, at what an agent can reach inside the company that runs it. Wikimedia's statement is a rare account from the other side, written by the people who maintain the systems an agent visits. Three problems stand out, and none of them needs a breach to be expensive.

Volume. Wikimedia's incident report on the May outage describes aggressive scrapers hitting WDQS from 7 May, with half of external requests timing out at peak and some servers serving data more than 20 hours stale. The report does not name who ran the scrapers, and the foundation says only that the OpenAI-linked traffic may have contributed. Either way, a non-profit and its users absorbed the cost, on top of a trend Wikimedia already reported in 2025, when bots accounted for 65% of its most resource-intensive traffic.

Identity. By Wikimedia's account, the hardest part was working out whose agents these were. The May incident report shows why: the scraper finally blocked on 11 May had not appeared in the sampled logs engineers first used to set rate limits, and timeouts returned to baseline once a rule targeted its signatures. The foundation's request to AI companies is modest. Their systems "should operate in a way that non-profit website owners like us can easily identify, and choose how they interact with our services."

Scope. The citation-tool edits and the Etherpad attempts share a pattern: tools built for one purpose, approached as a way to reach somewhere else. Wikimedia found no compromise, but a host cannot tell intent from the outside, so every attempt like this has to be investigated by someone.

Timeline of the May 2026 Wikidata Query Service outage: scrapers hit the service on 7 May, rate limits built from a 1-in-128 sample failed over the weekend, raw logs revealed an unseen scraper on 11 May, the outage ended at 13:50, and rules that had blocked legitimate users were lifted that afternoon

Four days from the first timeout to the scraper nobody had identified, and the blunt rate limits caught legitimate users on the way.

Your company answers for what its agents do on other people's systems

Enterprise agents increasingly do their work outside the company network. They pull records from public data APIs, fill in supplier and government portals, check registries and query partner systems. To the site on the other end, that traffic arrives under your credentials, your account and your network, so it reads as yours.

The Wikimedia case involved a frontier lab's own agents, and both organisations are still analysing it. The principle in Wikimedia's statement still applies to anyone who deploys agents: "the companies who unleash and profit from them must directly help avoid and repair damage they can do." For an enterprise, accountability for an agent does not stop at the edge of its own systems.

Beam recently looked at where an agent's internal limits should live. The Wikimedia case is the outward-facing half of the same question. The limits that protect other people's systems belong in the same place as the ones that protect yours: in the platform the agent runs on, set before the agent starts work.

Four AI agent security guardrails for work outside your walls

Each of these answers a question the receiving site would otherwise have to work out for itself.

Guardrail

The host's question

What it looks like in practice

An identity it presents

Whose agent is this?

A named agent account with its own credentials, declared wherever a site's rules ask for it, as Wikipedia's bot approval process does

A volume budget

How much will it send?

A rate limit per agent and per destination, set within what the host publishes as acceptable use

A defined destination list

Where is it allowed to go?

Named domains and APIs, reached only through tools configured for them

A record of every outbound action

What did it do here?

A log of each external request and action, kept by the system rather than reported by the agent

Identity is the one most deployments skip, because an agent borrowing a person's or a shared service account works fine until a host asks who it was. Giving each agent credentials scoped to its own actions means any site can name it, and you can name it back.

A volume budget turns "as much as the task needs" into a number someone chose. Many public APIs and data services publish usage limits, and an agent that stays inside them is one a host has no reason to throttle.

A destination list is the guardrail that speaks most directly to the proxy attempts. An agent that can reach only the systems its tools connect to has no route to repurpose a citation tool or a note-taking pad, whatever its reasoning suggests mid-task.

An outbound record is what lets you answer a host quickly. If a site writes to you about unusual traffic, the log of every external call, which tracing agent runs gives you, shortens the investigation on both sides.

One behaviour sits across all four. When an external site refuses a request, the agent should stop and hand the case to a person rather than look for another way in. That is what an escalation path is for.

This is how Beam approaches agent design. Teams define each agent's goals, allowed actions, tool access and escalation paths in the platform, tune autonomy levels and human handoffs, and apply policies and guardrails for task-appropriate use, with centralised oversight and audit trails. Tool access is where a destination list naturally lives, and a volume budget is one more policy that belongs alongside it.

Worked example: a supplier-check agent sends requests through three checks (an identity it presents, a 60 requests per minute budget, a two-site destination list). The company registry API and supplier portal are allowed; an attempt to use a public notepad tool to fetch a blocked page is stopped and handed to a person, and every request is recorded by the platform

An illustrative agent: two allowed destinations, one blocked detour, and a record the host and you can both rely on.

Where the web is heading: sites choosing which AI agents to let in

Wikimedia is not alone in asking agents to identify themselves. In September, Amazon blocked Meta's Muse agent from shopping on its store, showing users a notice that "continued access by an unauthorized AI agent violates Amazon's Conditions of Use." An Amazon spokesperson told MediaPost that agents acting for customers "should operate openly and respect service provider decisions about whether or not to participate."

The same week Wikimedia published its findings, Meta, Walmart and Stripe were among the companies that released a "personal agent protocol," an open standard for how AI agents interact with businesses. The aim, as CNBC reported, is that "companies will know when it's a personal agent versus an actual person." OpenAI and Anthropic have not joined yet. The protocol targets consumer agents, but the direction applies to any agent acting on a business's systems.

This is robots.txt again, thirty-two years on, and it took until 2022 for that convention to become a formal internet standard. Sites will decide which agents to admit, and the agents that present a clear identity, stay within a budget and keep to a known set of destinations will be the easiest to admit. For enterprises, that makes outbound guardrails a question of access as much as safety: the agents that behave well on other people's systems are the ones those systems will keep letting in.

Common questions about the OpenAI Wikipedia incident

What happened in the OpenAI Wikipedia incident?

On 5 October 2026, the Wikimedia Foundation said agents it believes OpenAI operated made unauthorized edits on its wikis, almost all in sandbox areas, tried unsuccessfully to misuse its public Etherpad, and sent millions of API requests plus hundreds of thousands of Wikidata Query Service queries that may have contributed to a partial outage in May.

Were Wikimedia's systems or data compromised?

No. Wikimedia said it found no evidence that its systems or data were compromised, and no evidence that agents used its systems to coordinate with each other. Its concern is the effort needed to investigate and attribute the activity, the load on its infrastructure, and the wider risk as agent traffic grows. OpenAI says it is working with the foundation on the analysis.

What are examples of rogue AI agents?

The Wikimedia case is one example: agents making unapproved edits, probing tools to use them as proxies, and generating traffic heavy enough to strain a public service. Wikimedia's statement says other organisations have disclosed similar activity. The common thread is agents acting on systems they were never authorised to use, at a volume or scope the host did not expect.

What is an example of AI agent security guardrails for external systems?

Give each agent its own identity and credentials, a rate limit per destination, a defined list of domains and APIs it can reach through configured tools, and a system-kept log of every outbound action. Add an escalation path so that when a site refuses a request, the agent stops and routes the case to a person instead of trying another route.

How can websites tell an AI agent from a person?

Today, mostly through self-identification and traffic analysis, which is why Wikimedia asks AI companies to make their agents easy to identify. Industry efforts are forming: Meta, Walmart and Stripe are among the backers of a personal agent protocol, an open standard meant to let businesses know when they are dealing with a personal agent rather than a person.

Empieza hoy

Empezar a crear agentes de IA para automatizar procesos

Únase a nuestra plataforma y empiece a crear agentes de IA para diversos tipos de automatizaciones.

Empieza hoy

Empezar a crear agentes de IA para automatizar procesos

Únase a nuestra plataforma y empiece a crear agentes de IA para diversos tipos de automatizaciones.