For most of the last decade, “autonomous vehicles are two years away” was a joke you could tell at any industry dinner and be sure of the laugh. That joke has expired. As of this month, Waymo runs fully driverless service in fourteen US cities, with more than half a million paid rides a week and a fleet of more than 4,000 vehicles. Denver, San Diego and Tampa came online on 1 September. The company’s stated target is a million rides a week by the end of the year.
And the safety record isn’t a hedge. Through March, Waymo had accumulated 220.6 million rider-only miles and reported 82 percent fewer injury-causing crashes than the human benchmark for the same areas, 94 percent fewer crashes involving serious injury or worse, and 93 percent fewer pedestrian crashes with injuries. Those are large numbers on a large denominator. Anyone still treating driverless cars as vaporware stopped paying attention somewhere around 2023.
So, the new comforting story writes itself: Autonomy is solved, and what remains is rollout. I’d like to look at where those miles actually are before agreeing. Of the 220.6 million rider-only miles, Phoenix accounts for 80.6 million, the San Francisco Bay Area for 67.1 million and Los Angeles for 51.8 million. That’s 199 million miles, roughly 90 percent of the entire evidence base, in three metros. Austin contributes 15.8 million and Atlanta 5.4 million. And the new cities aren’t arriving as metros but as carefully drawn polygons: Denver opened with about sixty square miles of service area, Tampa with fifty. None of that diminishes the achievement; it tells you what the achievement is, which is a different and more useful thing.
What shipped isn’t a driver; what shipped is an operational design domain (ODD): a bounded set of roads, speeds, weather conditions and traffic situations together with an accumulated body of evidence that the system performs safely inside that boundary. The ODD isn’t a limitation on the product; the ODD is the product. Once you see it that way, the shape of the last decade stops being a story about disappointment and becomes a story about a category error.
The industry priced autonomy as a perception problem. Get the sensing right, get the planner right and the driving follows. Enormous effort went into exactly that, and it largely worked: The perception stack isn’t what’s holding anyone back today. What the industry underpriced was the cost of demonstrating that the thing is safe. That cost doesn’t scale with how good your model is; it scales with how wide your domain is.
Rand Corporation saw this coming with unusual precision. In 2016, Nidhi Kalra and Susan Paddock published “Driving to safety,” which asked how many miles of test driving it would take to show statistically that an autonomous vehicle is safer than a human. The answer was hundreds of millions of miles, and for the rarest and most serious outcomes, hundreds of billions; tens to hundreds of years at realistic fleet sizes. Their conclusion was blunt: Test-driving alone can’t provide sufficient evidence for demonstrating autonomous vehicle safety.
Three responses were available: Simulate more, build scenario-based testing or shrink the domain until the evidence you can actually gather is enough to cover it. The first two help. The third is the one that shipped.
Freight is the confirming case, and the ordering there is more instructive than anything happening in robotaxis. The first place a heavy truck went genuinely driverless wasn’t an interstate. Kodiak removed the human from trucks operating off-highway in the Permian Basin in early 2025: private roads, industrial traffic, a customer who controls the entire environment. Meanwhile, on public highways, the picture is more measured than the headlines suggest: Aurora has been running commercial freight in Texas since May 2025 and has expanded to five lanes, but as of last month, those trucks still carry safety drivers, and Kodiak’s plan to remove the operator on highway routes is a target for the end of this year. California only lifted its ban on testing driverless vehicles over 10,000 pounds on public roads in April, and the Teamsters sued the DMV over it in August.
Read that sequence carefully. Driving in the Permian Basin isn’t easier than driving from Dallas to Houston. The terrain is worse and the equipment is heavier. What’s easier is bounding it. One customer, one site, known traffic, no members of the public. The evidence required to justify removing the driver is a fraction of what a public highway demands, and that (not capability) determined what went first.
For a startup, this reframes the founding question entirely. You’re not building a driver; you’re buying down the cost of evidence in a domain you’ve chosen, and the choice of domain is the whole business.
The right questions are therefore economic rather than technical. How tightly can you bound this domain? How much evidence does a regulator, an insurer and a customer require before the human comes out? Can the route economics carry that cost before your funding runs out? A fixed lane between two distribution centers, a port terminal, a mine, a campus, an agricultural field – these look unglamorous next to “drive anywhere” and they’re the only places where the arithmetic closes at seed-stage scale. The teams that promised general autonomy spent a decade discovering that the cost of evidence for an unbounded domain is effectively unbounded too.
The corollary is that the accumulated evidence is itself the moat, and it’s oddly non-portable. Waymo’s eighty million miles in Phoenix are worth a great deal in Phoenix and very little in Oslo, where the roads are narrower, the lane markings vanish under snow for four months a year and the human benchmark it would be measured against is entirely different. Expansion isn’t a software rollout. Every new domain is a fresh evidence-gathering exercise, which is precisely why the map fills in at the pace it does.
Incumbents face a version of this that I think is underappreciated in the automotive industry, and it’s not the usual build-buy-partner question. Consider what a carmaker actually sells: a vehicle that goes wherever its owner points it. That’s the widest conceivable operational design domain, and the manufacturer controls none of it. A fleet operator sells rides inside a polygon it drew itself, on roads it mapped, in weather it can decline to operate in. The fleet operator owns its ODD; the carmaker’s customer owns theirs.
This, more than any difference in engineering talent, is why robotaxis reached driverless operation before consumer vehicles did, and it suggests the value in autonomy may accrue to whoever controls the domain rather than whoever manufactures the machine. An automaker selling autonomy as a feature is promising performance in an environment it can’t bound, can’t instrument and can’t withdraw from. That’s a difficult product to underwrite and a worse one to insure.
For logistics incumbents, the same logic points somewhere more comfortable. If the ODD is the product, then a lane is an asset – and freight companies already own lanes, terminals, yards and the traffic data that comes with them. The strategic move is to identify which of your lanes is most boundable and instrument it now, rather than waiting for a vendor to arrive with a general solution that will never exist.
For society, the honest position here is more positive than the discourse usually allows, and more qualified than the industry’s own messaging. An 82 percent reduction in injury-causing crashes across 220 million miles isn’t a marketing claim; it’s one of the better-evidenced safety findings in modern transport, and it deserves to move the public conversation considerably further than it has. Road deaths are a mass-casualty problem we’ve collectively decided to treat as weather. If this technology holds its numbers at scale, the argument for deploying it faster is a public-health argument.
But the evidence is domain-bounded, and public debate isn’t. “Self-driving cars are safer than humans” is being repeated as a global claim when what has been demonstrated is a claim about particular vehicles on particular roads in particular cities in mostly good weather. Those aren’t the same statements, and conflating them will produce exactly the backlash the industry fears the first time a system is deployed somewhere its evidence doesn’t cover.
There’s also a distributional question nobody is asking. Denver’s service area is sixty square miles, drawn by a company, on commercial criteria. If driverless vehicles genuinely reduce injury crashes by four-fifths, then the boundary of that polygon is a line between people who get the safer roads and people who don’t. We’re used to arguing about where transport infrastructure goes, with all the equity fights that entails. We haven’t yet noticed that we’re now allocating safety the same way, through a map that no public body draws.
And the labor question is live rather than theoretical. The Teamsters’ suit in California is an early instance of what will be a long argument, and the freight case is where it bites hardest, because a fixed lane between two terminals is both the easiest domain to bound and the one with the most jobs attached.
One last observation, because it’s the cleanest instance of something I’ve been circling for months. Throughout this series, I’ve argued that the durable artifact in AI systems is becoming the contract a system must honor plus the continuous evidence that it does. Autonomous driving is the one field that already writes both down. The ODD is the contract: an explicit, auditable specification of conditions under which the system may operate. The safety case is the evidence, maintained continuously and reestablished for every expansion. Software engineering has been looking for a discipline for evaluating learning systems. This industry, under regulatory pressure and liability exposure, went and built one.
Which brings me back to the joke about two years away. Roy Amara, then president of the Institute for the Future, gave us the standard explanation in 1978: “We tend to overestimate the effect of a technology in the short run and underestimate the effect in the long run.” It’s usually quoted as a comment on human psychology, as though we’re simply bad at forecasting. The autonomous vehicle story suggests something more specific and more useful. We overestimated the short run because we priced the problem as perception, which was tractable and got solved roughly on schedule. We’re underestimating the long run because we’re still not pricing the real work, which is the slow, unglamorous, domain-by-domain accumulation of evidence. And evidence, unlike capability, doesn’t arrive in a step change.
Next in the series: Energy and the Grid — AI as the largest new load on the system and the best tool we have for running it.
Want to read more like this? Sign up for my newsletter at jan@janbosch.com or follow me on janbosch.com/blog, LinkedIn (linkedin.com/in/janbosch) or X (@JanBosch).