Rule by Default
The rules for where AI agents may go, and what they may do when they get there, are being set by presets. I measured both ends of the connection. One is closed by default. The other is open by default. Hardly anyone bound by either setting chose it.
Checked. On September 26, 2026, 28 of 92 large US retail storefronts did not serve their robots.txt to a scanner that said who it was. Twenty-two of them refused with a 403. Of the 64 that served the file, 16 name any AI crawler or agent.
On September 27, 2026, 288 of 384 actively maintained, network-facing MCP servers whose code takes a sensitive action showed no authentication check in the files I scan. That is 75 percent.
A refusal cannot say why. The server figure is a static reading of source code, a signal and not a confirmed vulnerability. No retailer is named from scan data. Two runs on the door side and eight snapshots on the server side are a short series, and I treat them as one.
- 01
The checkbox
A vendor announces a preset, with a date, for a large share of the web.
Document - 02
Ninety-two doors
Six plain questions, asked of 92 storefronts every Saturday, and what came back.
Data - 03
The rulebook and the bouncer
What a site writes down, and what happens at the door.
Data - 04
The unlocked room
Of 384 network-facing tool servers, 288 show no authentication check.
Data - 05
Closed by default, open by default
The two measurements, side by side.
Drawing - 06
Why presets stay put
The record on people, and what my own series says about code.
Data - 07
The default is the law in force
Whose idea this is, and the railroads that set the clocks in 1883.
Document - 08
When a vendor moved a market, and when one could not
Four settings that changed, and one that was announced and then kept.
Document - 09
The default-setters
Five kinds of holder, and a sixth party that holds nothing.
Drawing - 10
What a good default looks like
Four properties: readable, published, able to tell who is asking, reversible.
Document - 11
Where I could be wrong, and the scorecard
Five arguments against this booklet, and eight predictions with dates.
Scorecard
The checkbox


On July 1, 2026, Cloudflare published a blog post with a friendly title: "Your site, your rules." Inside it was a date.
From September 15, the company said, its default for new customers, for new sites of existing customers, and for existing free customers who had not changed their settings would be to block AI training crawlers and AI agents on pages that display ads. Search crawlers would stay allowed.
Look at how the post defines an agent: "automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now." That is not a scraper. That is your assistant, fetching a page because you asked it to.
The reasoning was stated in one line.
"...we treat human attention as the end goal, and keep away the bots..."
Cloudflare says more than 20 percent of web domains sit behind it. A site owner can switch the setting off in the dashboard. An owner who never opens the dashboard has a policy anyway.
I should say what I cannot show. I have looked for a Cloudflare page dated on or after September 15 confirming that the change went live, and I have not found one. The documentation still speaks in the future tense. The July press release said the defaults would be finalized "with a deadline of September 15, 2026." So I will claim only this much: a vendor announced a preset, with a date, for a large share of the web, and whether your site carries it depends on a setting you may never have seen.
I am not picking on Cloudflare. It published its reasoning at length, which most companies in its position do not. It is simply the clearest example I have of the thing this booklet is about.
In the agent economy, the default setting is the law in force.
For new customers and for new sites for existing customers, on September 15, 2026, the defaults will be set to allow for search but block training and agent use for pages with ads.
On September 15, 2026 these changes will also be made for all existing free customers that have not changed their settings by September 15, 2026 in their dashboard.Cloudflare · press release · July 1, 2026
Agent: automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now. This includes chat fetch bots (e.g., ChatGPT-User) and browser-use agents (e.g., Gemini or Claude driving Chrome).Cloudflare · Your site, your rules · July 1, 2026
The default, as announcedSeptember 15
The press release gives the scope and the blog gives the definition. The default applies on pages that display ads, and search crawlers stay allowed. I have not found a Cloudflare page dated on or after September 15 that confirms the change went live.
Cloudflare press release and blog · both dated July 1, 2026 · re-typeset, not a screenshot
Ninety-two doors

Every Saturday I ask 92 of the largest retail storefronts in America the same six questions. The sample is the NRF Top 100, minus the eight with no single storefront.
The questions are plain requests. Can I read your robots.txt? Do you publish an llms.txt, a commerce profile, an agent card? Will you serve me your homepage if I am not a browser? The scanner names Major Labs in every request and links to a page that explains itself. It waits 1.5 seconds between requests. If the door says no, it leaves and does not try again.
Here is what came back on September 26.
Twenty-eight of the 92 did not serve robots.txt. Twenty-two refused with a 403. Four never answered. One answered with a 418, the joke status code that means "I'm a teapot," and one served something that was not a robots file.
Of the 64 that did serve the file, 16 name any AI crawler or agent at all. Two block the kind of agent that acts for a live person.
Forty-nine served a homepage an agent can read. Sixteen publish an llms.txt. Three publish a Universal Commerce Protocol profile. None has an A2A agent card.
One limit belongs here and not in a footnote. My scanner is an unknown research tool that identifies itself. It is not a shopping agent from a company the retailer has heard of. What I measure is how a stranger who is honest about being software gets treated at the door.
92 doors, 28 shutrobots.txt
Each mark is one storefront. On September 26, 2026, 28 of 92 did not serve robots.txt to a scanner that said who it was: 22 refused with a 403, four never answered, one answered 418, and one served something that was not a robots file. Of the 64 that served the file, 16 name any AI crawler or agent. The order is shuffled, because I publish the count and not the names.
Major Labs scan · US cloud run 2026-09-26 · 92 storefronts
GET /robots.txt User-Agent: MajorLabsReadinessScan/0.1 (+https://majorlabs.co/about; independent research; read-only, a few requests per site)
The knock, and the usual answer403
The one file the web sets aside to tell a visitor the rules, and the most common answer among the 28 storefronts that did not serve it. I drew this exchange. It is not copied from any one site. The user agent is the one my scanner sends in every request.
Drawing · Major Labs · request as the scanner sends it · user agent string from scan.py
The rulebook and the bouncer

The web has exactly one door policy that is written down in a standard place. It is robots.txt, and the people who wrote the standard were clear about what it is not.
"These rules are not a form of access authorization."
That is RFC 9309, published in 2022. Martijn Koster said the same thing in the original 1994 note: "It is not enforced by anybody."
So there are two layers. The rulebook says what a site wants. Something else does the refusing. My data lets me hold one against the other.
Of the 64 storefronts with a readable robots.txt, 62 have written no rule that blocks a user-triggered agent. Eight of those 62 still refused or challenged my scanner at the homepage. Five more handed back an empty shell. One returned a page too large to judge.
Two storefronts publish an llms.txt, a guide written for AI systems, and turned the scanner away at the homepage regardless.
The rulebook and the bouncer are not reading from the same page.
Here is a detail I did not expect. Under the standard, when a request for robots.txt gets a status code in the 400 range, the file counts as unavailable, and "the crawler MAY access any resources on the server." Twenty-two storefronts answered 403. Read by the book, a refusal to show the rules tells a standards-following crawler that there are no rules. I doubt any of the 22 meant that. The standard is written for crawlers, and "may" is not "should," so I would not lean on it in court. It shows how far the enforcement layer has drifted from the written one.
Week to week, the written layer did not move at all. All 92 storefronts gave identical answers on every written measure: who they name, who they block, which files they publish. The small changes I saw were all at the edge.
Now the weakness, stated as plainly as I can. My scan sees a refusal. It cannot see why. These are some of the largest retailers in the country. They have security teams and enterprise contracts, and their bot rules may be tuned by hand. If so, the refusals are decisions.
Then they are decisions nobody wrote down. Sixteen storefronts name an AI agent in writing, and two block one that acts for a person. Far more than two are turning it away.
What I can say today is that the pattern is consistent with a preset. I cannot say it is set by one. The scanner does not yet record which vendor answered the door. It will, and if the vendor predicts the answer better than anything about the retailer, that is evidence for the default. If it does not, this chapter gets weaker and I will say so.
Written against enforced62 of 64
On September 26, 62 of the 64 storefronts with a readable robots.txt had written no rule that blocks a user-triggered agent. Eight of those 62 still refused or challenged my scanner at the homepage, five more returned an empty shell, and one returned a page too large to judge. My scanner identifies itself as a research tool, so this measures how an unknown agent is treated, and it cannot show whether a refusal was a preset or a decision.
Major Labs scan · US cloud run 2026-09-26 · of 64 readable files
These rules are not a form of access authorization.RFC 9309 · Robots Exclusion Protocol · 2022
It is not enforced by anybody.Martijn Koster · A Standard for Robot Exclusion · 1994
A sign, not a lock1994, 2022
The standard behind robots.txt says this about itself, and the original note said it 28 years earlier. The file states what a site wants. Something else does the refusing.
IETF, RFC 9309, section 1 · robotstxt.org, original note of 1994 · re-typeset
If a server status code indicates that the robots.txt file is unavailable to the crawler, then the crawler MAY access any resources on the server.
Read by the book400 range
Under the standard, a status code in the 400 range means the file counts as unavailable. Read by the book, a refusal to show the rules tells a standards-following crawler that there are no rules. I doubt any of the 22 meant that. The rule is written for crawlers, and may is not should.
IETF, RFC 9309, section 2.3.1.3 · re-typeset · count from the Major Labs scan of 2026-09-26
The unlocked room


Turn around and look at the other end of the connection.
When an agent does something, it usually does it by calling a tool server. The common way to build one is the Model Context Protocol. I scan the public MCP servers on GitHub and read their source code for two things: does this server take a sensitive action, and does anything check who is asking?
As of September 27, 384 actively maintained servers are network-facing and take a sensitive action. Writing files. Running shell commands. Changing a database. In 288 of them, 75 percent, I find no authentication check.
Treat that as a signal. It is a static reading of up to 12 source files per server, and a check could live somewhere I do not look. An independent academic study took a different route in May. It probed 7,973 live remote servers and found 40.55 percent exposing tools without authentication. That is a different population and a different method, so the two numbers cannot be compared. They point the same way.
Why would three in four have no lock? Start with the specification.
"Authorization is OPTIONAL for MCP implementations."
That sentence is in the current release and in the two before it.
Then follow the path a new developer walks. The official quickstart builds a weather server with no authentication step. That is reasonable: it runs over a local transport, and the specification says local servers should not use the network authorization flow. The Python SDK's README offers "a server in 15 lines." A few lines further down, one command serves that same starter over HTTP. Nothing is put in front of it.
Authentication is available. Both official SDKs document it on its own page, and the Python documentation is candid: "The SDK has no opinion about what a valid token looks like." It is something the developer adds. The first-run path does not include it.
Nobody did anything wrong here. A protocol that wants adoption keeps the first step easy. A tutorial that runs on your laptop does not need a lock. Then the laptop demo becomes a service, and the setting that made sense on day one is still the setting.
384 rooms, 288 with nobody at the desk
Of 384 actively maintained, network-facing MCP servers whose source code takes a sensitive action, 288 (75 percent) show no authentication check in the files I scan, as of September 27, 2026. This is a static reading of up to 12 source files per server. It is a signal, and it is not a confirmed vulnerability.
Major Labs scanner · snapshot 2026-09-27 · the order of the marks is shuffled
Authorization is OPTIONAL for MCP implementations.Model Context Protocol · Authorization · release 2026-07-28
Optional2026-07-28
The specification says authorization is optional. The official quickstart, the SDK README examples and the starter templates contain no authentication step, and they run over a local transport where the specification says the HTTP authorization flow should not be used. Authentication is documented and available in both official SDKs, as something the developer adds.
modelcontextprotocol.io · dated release 2026-07-28 · the same sentence is in the two releases before it · re-typeset
The README's own heading
A server in 15 lines
The 15 lines are not reproduced here.
Serve server.py over HTTP:
uv run mcp run server.py --transport streamable-http
A URL means Streamable HTTP, the transport you deploy.
modelcontextprotocol/python-sdk · README.mdOne command to a networkREADME
The README offers a server in 15 lines. A few lines further down, this one command serves that same starter over HTTP, and nothing is put in front of it. The word auth does not appear in the README. It is a local demo, and authentication is documented on its own page as something the developer adds.
github.com/modelcontextprotocol/python-sdk · README.md · commit of 2026-09-23 · re-typeset
Closed by default, open by default

Put the two measurements side by side.
At the storefront, an agent that says who it is gets refused before it can read the rules. At the tool server, a caller who says nothing can often write a file.
One side is shut and the other is open. The cause looks the same. In both places the outcome is consistent with a setting that was inherited, and in neither place did the party bound by it write anything down.
Each default is rational for whoever set it. A bot manager's preset was tuned years ago for scrapers and card testers, and for that job, refusing unknown software is correct. A protocol's first-run path was tuned for adoption, and for that job, asking nothing is correct. Neither was designed with a customer's agent in mind, because when they were designed there was no such customer.
So the person who sends an assistant to buy a coffee maker meets a locked door at the shop, and the anonymous caller at the workshop finds the tools laid out. I would not say the locks are on the wrong doors. I would say nobody decided where the locks go.
Closed by default at the door. Open by default at the server. One missing decision, two opposite outcomes.
missing
decision
The two defaults, side by side
Two measurements, taken one day apart. They come from different populations and different methods, so I set them side by side and do not add them up. In both places the outcome is consistent with a setting that was inherited.
Major Labs scans · doors 2026-09-26 · servers 2026-09-27 · drawing
Why presets stay put

There is a long record on what people do with a default. Mostly they leave it.
The famous case is organ donation. In 2003, Eric Johnson and Daniel Goldstein compared countries where you must opt in to be a donor with countries where you must opt out. On effective consent rates, "the two distributions have no overlap, and nearly 60 percentage points separate the two groups." Same question. Different starting position.
Money behaves the same way. Brigitte Madrian and Dennis Shea studied a large US company that began enrolling employees in its retirement plan automatically. Participation was significantly higher, and a substantial fraction of those enrolled stayed at the contribution rate and the fund the company had picked for them.
A federal court has looked at the commercial value of the position. In United States v. Google, the court found that Google paid $26.3 billion in 2021 in revenue share under its distribution contracts, and it wrote a sentence that belongs on a wall.
"The default is extremely valuable real estate."
All of that is about people. My data is about code and configuration, and it says the same thing.
A week apart, 89 of 92 storefronts gave the same answer on robots.txt. I count readable, refused and no answer as three different answers, which is the strict way to count. On the looser test of readable or not, it is 91.
On the server side, I scanned 1,934 servers in both June and September. Fourteen weeks later, 1,668 of them had exactly the same number of findings. That is 86 percent. In that same group, 12 servers put authentication in front of a sensitive action. Six took it away. And 430 were unauthenticated in June and unauthenticated in September.
Then there is the part I did on purpose. In June I flagged 47 high-risk servers for disclosure and wrote to the maintainers of nine. One fixed the problem. Across the whole 47, as of September 27, two are fixed, four have fewer findings, 34 are unchanged and seven are worse.
The numbers are small, and I will not claim that writing to people makes no difference. I will re-score all 47 on December 12 and publish what I find. What I will say is that a polite letter did not move the line, and I no longer expect a thousand polite letters to.
The line that did not move1,668 of 1,934
Of the 1,934 servers I scanned in both June and September, 1,668 had exactly the same number of findings 14 weeks later, 114 had fewer and 152 had more. In that group 12 put authentication in front of a sensitive action, six took it away, and 430 were unauthenticated at both dates. Of the 47 servers I flagged for disclosure in June, two are fixed, four have fewer findings, 34 are unchanged and seven are worse.
Major Labs scanner · snapshots 2026-06-18 and 2026-09-27 · static reading of source code
Same answer next week89 of 92
A week apart, 89 of 92 storefronts gave the same answer on robots.txt, counting readable, refused and no answer as three different answers. On the looser test of readable or not, it is 91. On every written measure, all 92 were identical. Which marks are raised is arbitrary. Two runs on the door side and eight snapshots on the server side are a short series, and I treat them as one.
Major Labs scan · US cloud runs 2026-09-20 and 2026-09-26
The default is extremely valuable real estate.United States v. Google LLC · Memorandum Opinion · August 5, 2024
Valuable real estatepage 2
In United States v. Google, the court found that Google paid $26.3 billion in 2021 in revenue share under its distribution contracts, and wrote this sentence on page 2 of its opinion. The figure is the total paid under all of those contracts. It is not the price of one default.
US District Court for the District of Columbia · case 20-cv-3010 · document 1033 · re-typeset
Opt in, opt out2003
A schematic, not a chart. Johnson and Goldstein compared effective consent rates in four opt-in countries and six opt-out countries, and wrote that the two distributions have no overlap and that nearly 60 percentage points separate the two groups. My sources here give the gap and not a figure for each country, so I have drawn no country values. These are consent rates, not transplant rates.
After Johnson and Goldstein, Do Defaults Save Lives?, Science, November 21, 2003 · redrawn as a schematic
The default is the law in force


None of this is my idea, and I want to be exact about whose it is.
Lawrence Lessig wrote in 2000 that software regulates whether or not anyone admits it: "The code regulates. It implements values, or not." Kevin Kelly put the point about presets in two sentences in 2009. "Most defaults are never altered," he wrote, and so "the privilege of establishing what value the default is set at is an act of power and influence."
Jay Kesan and Rajiv Shah made software defaults a subject for law in 2006. Cass Sunstein built a theory of default rules in 2013. In 2025, Luke Hogg and Tim Hwang argued in Tech Policy Press that a kill switch offered to every site on one network "is effectively centralizing decisions about who can crawl vast amounts of information." This July, Paul Schaus published a commentary for banks under a title close to mine, "Governed by Default."
So the claim is old. What I can add is a measurement: both ends of the agent connection, the same way every week, with the dates on it.
There is a precedent I keep coming back to. On November 18, 1883, the American railroads reset their clocks to a system of standard time zones. No law told them to. Towns that had kept their own noon found that the station clock had a different one. The first federal law on standard time passed on March 19, 1918. By then the country had been living on railroad time for 34 years.
Law is not absent from my subject, and I will not pretend it is. On August 4, 2026, the Ninth Circuit ruled in Amazon's case against Perplexity that "it is the user who 'accesses' Amazon's computers, with the help of the Assistant to carry out specific acts on Amazon.com." The ruling is narrow. It vacated a preliminary injunction on the record in front of the panel. Europe has legislated defaults for personal data since 2018, and the Cyber Resilience Act will require products to ship "with a secure by default configuration" when it applies in full on December 11, 2027.
Courts are deciding who the agent is. I have not found a statute or a regulation that says what a website's default treatment of a person's agent must be, or what a tool server must check before it acts. Until one does, the setting is the rule.
The hard truth, as any engineer will tell you, is that most defaults are never altered.Kevin Kelly · Triumph of the Default
Therefore the privilege of establishing what value the default is set at is an act of power and influence.The Technium · June 22, 2009
An act of power and influence2009
Kevin Kelly put the point about presets in two sentences in 2009, and I quote both in full. The claim is his. What I add is a measurement.
kk.org, The Technium · June 22, 2009 · re-typeset
When a vendor moved a market, and when one could not

If presets stay put, how does anything change? Look at the times it did.
In 2021, Apple made tracking on the iPhone something an app had to ask for. A working paper presented at the FTC measured what followed in the United States: the share of trackable Apple traffic fell "from 73% to 18%." The authors estimate a 21 percent fall in ad revenue from Apple users for publishers. Meta's finance chief told investors the iOS changes were a headwind "on the order of $10 billion" for 2022. That was the company's own estimate.
In April 2023, Amazon Web Services changed what a new storage bucket looks like. Public access is blocked unless the owner turns it on. Years of guidance had asked customers to lock their buckets. The default did it for every new one.
Defaults can open doors as well. In November 2025, AWS said its web firewall would let verified agents through: "Verified WBA bots will now be automatically allowed by default." In March 2026, Shopify told merchants, in an email reported by Modern Retail, that its agentic storefront channel "will launch by default for your store."
And sometimes the holder of a default announces a change and then does not make it. Google spent years preparing to change how Chrome handles third-party cookies. In April 2025 it said it had "made the decision to maintain our current approach."
I do not have a theory of why one of these stalled and the others went through. What I take from them is narrower. When the number moved, the party that held the default moved it, and it moved for everyone under that default at once.
Change arrives in rows.
| Date | Holder | The default | In the source |
|---|---|---|---|
| The holder moved the default | |||
| 2021 | Apple | Tracking on the iPhone became something an app had to ask for. | from 73% to 18%Share of trackable Apple traffic in the United States · Kraft, Skiera and Koschella, working paper, October 7, 2023 |
| April 2023 | Amazon Web Services | Public access is blocked on a new storage bucket unless the owner turns it on. | Amazon S3 now applies two security best practices to all new buckets by defaultTitle of the AWS notice · posted April 28, 2023 |
| November 2025 | Amazon Web Services | The web firewall lets verified agents through. | Verified WBA bots will now be automatically allowed by default.AWS notice · posted November 21, 2025 |
| March 2026 | Shopify | The agentic storefront channel is switched on for the store. | will launch by default for your storeShopify email to merchants, as reported by Modern Retail · March 12, 2026 |
| The holder announced a change and kept the default | |||
| April 2025 | Third-party cookies in Chrome stay as they were. | made the decision to maintain our current approachGoogle, Privacy Sandbox blog · April 22, 2025 | |
Four that moved, one that was kept2021 to 2026
Four times the holder of a default moved it, and once a holder announced a change and then did not make it. When the number moved, it moved for everyone under that default at once. The Apple figures come from a working paper presented at the FTC and cover the United States.
Quoted words exactly as they appear in each source · dates are the dates of the sources · re-typeset
The default-setters

If the setting is the rule, it is worth knowing who holds the settings. I count five kinds of holder, and I am not ranking them.
Edge and bot-management vendors. They sit in front of the website and answer the door. Their presets decide what an unknown visitor gets before the site's own code runs.
Protocol and SDK maintainers. They decide what a server does when the developer specifies nothing. "Optional" is a default.
Authors of starters. The quickstart, the README example, the template. Whatever is in the first 15 lines is in a great many servers, though I cannot yet tell you how many. I have no measurement of how many servers begin from a starter, and I am not going to guess.
Commerce platforms. They can switch a channel on for every store they host, as Shopify did.
AI companies. They choose how their agents announce themselves: a named user agent, a signed request, or a browser that looks like a person. That choice is a default too, and it shapes what every door on the street can do about them.
Then there is a sixth party, which holds nothing. The merchant who inherits the bot rule. The developer who inherits the starter. They are bound by the setting, and mostly they have not read it.
Five holders and a sixth party
Five kinds of party hold a setting that binds someone else. The sixth holds nothing: the merchant who inherits the bot rule, and the developer who inherits the starter. The order is the order of the chapter, and it is not a ranking.
Drawing · Major Labs · roles only, no company is placed on it
What a good default looks like

I am not arguing that these settings are wrong. A preset that refuses unknown software protects merchants from real abuse. An open first step is how protocols get used at all. My complaint is about legibility.
A good default has four properties.
It can be read. The bound party can find out what the setting is without filing a ticket.
It is published where it takes effect. If the door refuses agents, the rule that says so is at the door, in a form an agent can fetch.
It can tell who is asking. A signed request from a customer's assistant and an anonymous scraper should not look identical. The work on this is real but unfinished. AWS notes that Web Bot Auth "relies on two active IETF drafts."
It can be reversed. The opt-out exists and an ordinary owner can find it.
On the server side the ask is older and has a name. In 2023 the US cybersecurity agency CISA and its partners published guidance that says it in one sentence: "A secure configuration should be the default baseline." Applied here, that means the path that puts a server on a network includes the lock. The lock already exists in both SDKs. It needs to be in the first 15 lines that reach a network.
A secure configuration should be the default baseline.CISA and partners · announcement of April 13, 2023
The default baseline2023
From the announcement of the guide that CISA published with its partners in 2023. It is guidance from a government agency, and it is not regulation. Applied here, it means the path that puts a server on a network includes the lock.
cisa.gov · April 13, 2023 · re-typeset
Where I could be wrong, and the scorecard

Here are the five best arguments against this booklet.
The scan cannot tell a preset from a decision. This is the strongest one, and chapter 3 concedes it. The index also did not move when Cloudflare's date passed. Two runs on September 26 gave 29 and 28, which is noise. My headline example and my headline number are not causally linked, and I have not said they are.
Defaults slip when someone powerful wants them to. Lauren Willis showed this in "When Nudges Fail." Retailers have a revenue reason to open the door. I agree that the doors will move. My claim is about how: in steps, when a holder acts.
Law is already here. It is, in part. Chapter 7 says where.
The default may be right. It may. I have argued for legibility and not for a particular setting.
The server figure overstates. The all-server count includes local servers behaving exactly as the standard intends. That is why I lead with the network-facing 288 of 384.
A theory that cannot be wrong is not worth much, so here is how to mark mine. Each line can be scored from my own series or from the public record.
| # | Prediction | Baseline | Scored on |
|---|---|---|---|
| 01 | Doors move in steps. The count withholding robots.txt stays between 23 and 33 of 92, or any move outside that range happens within four weeks. | 28 | Run nearest March 27, 2027 |
| 02 | Written policy lags. Fewer than 26 of 92 name any AI agent in robots.txt. | 16 | Run nearest March 27, 2027 |
| 03 | The server line holds. The network-facing no-auth share is 70 percent or higher. | 75 | Scan nearest March 27, 2027 |
| 04 | Templates beat letters. If an official SDK puts authentication in its first-run network path, servers first seen after that date fall below 50 percent no-auth within six months. | n/a | Void if no such change by September 30, 2027 |
| 05 | Disclosure stays flat. Fewer than 10 of the 47 cohort servers are fixed or reduced. | 6 | December 12, 2026 |
| 06 | Protocols arrive through platforms. Fewer than 10 of 92 publish a UCP profile. | 3 | Run nearest March 27, 2027 |
| 07 | Another vendor moves. One more large edge provider announces a changed default for agent traffic. | n/a | By June 30, 2027 |
| 08 | No statute first. No US federal statute and no EU regulation in force sets the default treatment of a person's agent at a website. | n/a | By September 15, 2028 |
The scorecard
Eight predictions, each with a baseline and the date on which it is scored. Each line can be scored from my own series or from the public record.
Baselines from the door run of 2026-09-26 and the server snapshot of 2026-09-27
The last one is the most likely to be wrong. I would be glad to lose it.
Two dates, and a blank1883, 1918
The railroads reset their clocks on November 18, 1883. The first federal law on standard time passed on March 19, 1918, and by then the country had been living on railroad time for 34 years. The second line starts in 2026, and its far end is blank because I have not found a statute or a regulation that sets the default.
Dates from the Library of Congress guide, The Day of Two Noons · drawing
How the numbers were made
Two scanners, run the same way each time. Every limit I know of is written here.
The doors.
NRF Top 100 Retailers 2026, one consumer storefront each, 92 of 100. Six plain requests per site, 1.5 seconds apart, under a user agent that names Major Labs and links to it. No login, no form, no cart, no pretending to be a browser or another company's bot, no retry after a refusal. One US cloud address, every Saturday. Comparison runs from a second US data center and a UK home connection on September 20.
The servers.
Public MCP servers on GitHub, actively maintained. Static analysis of up to 12 source files per server for sensitive actions and for authentication checks. Per-server snapshots on eight dates between June 18 and September 27, 2026.
What I never publish.
Per-retailer rows. Per-server findings, which go to maintainers through coordinated disclosure.
Limits.
A refusal cannot say why. A timeout is no answer, not a refusal. A site that refuses the scanner cannot show it a file, so file counts are undercounts. Static analysis can miss a check that lives outside the scanned files. The door series has two weekly runs. The server series has eight snapshots.
Where to check me
- CloudflareYour site, your rules, new AI traffic options for all customers
- Cloudflarepress release, July 1, 2026
- TechCrunchCloudflare's new policy pushes AI companies to pay for publishers' content
- Major MattersThe Merchant Agent-Readiness Index
- IETFRFC 9309, Robots Exclusion Protocol
- Martijn KosterA Standard for Robot Exclusion, 1994
- Major LabsMCP security scoreboard
- Major LabsAgent Identity Tracker
- Model Context ProtocolAuthorization, specification 2026-07-28
- Model Context ProtocolBuild an MCP server
- Model Context ProtocolPython SDK README
- Zhou et al.A First Measurement Study on Authentication Security in Real-World Remote MCP Servers
- Johnson and GoldsteinDo Defaults Save Lives?
- Madrian and SheaThe Power of Suggestion
- United States v. Google LLCMemorandum Opinion, August 5, 2024
- Lawrence LessigCode Is Law, On Liberty in Cyberspace
- Kevin KellyTriumph of the Default
- Kesan and ShahSetting Software Defaults
- Cass SunsteinDeciding by Default
- Tech Policy PressCloudflare's Troubling Shift From Guardian to Gatekeeper
- CCG CatalystGoverned by Default
- Library of CongressThe Day of Two Noons
- Ninth CircuitAmazon.com Services v. Perplexity AI, August 4, 2026
- GDPRArticle 25
- EU Cyber Resilience ActRegulation (EU) 2024/2847
- Kraft, Skiera and KoschellaEconomic Impact of Opt-in versus Opt-out Requirements
- CNBCFacebook says Apple iOS privacy change will result in $10 billion revenue hit this year
- AWSAmazon S3 now applies two security best practices to all new buckets by default
- AWSAWS WAF announces Web Bot Auth support
- Modern RetailShopify says purchases are coming inside ChatGPT through agentic storefronts
- GoogleNext steps for Privacy Sandbox and tracking protections in Chrome
- CISASecure by design and default principles
- Lauren WillisWhen Nudges Fail, Slippery Defaults
If the setting is the rule, who is reading the settings?
Take the whole archive with you.
The booklet is this page set for print: every plate and document, the scorecard, the method and the sources. Subscribe to Major Labs Weekly, one email a week, and the PDF is yours to download.
Both series run again.
The figures in this booklet are readings from two series that continue. The door index and its method are at majormatters.co/trackers/merchant-readiness. The server scoreboard is at majorlabs.co/security. I will score the predictions on the dates printed beside them.
Use it. Cite it. Argue with the method. Every choice is written down, and I will change the ones that are wrong.
- Data plates and drawings10, with dates
- DocumentsNine, with sources
- Illustrations11, one per chapter
- ScorecardEight predictions, with dates
Major Labs Weekly is one email a week. Unsubscribe in one click. Privacy.
Built by Charlie Major. Major Labs is independent of Mastercard and operates separately from Major Matters. Any opinions are Charlie's own.
Format inspired by the research archive pages of Dami Lee and Nollimedia.