What happens when you take a “frontier” (read: general-purpose) AI model like Claude or GPT and ask it to get behind the wheel and drive a car? Sure, some self-driving systems utilize artificial intelligence in one sense or another, but they’re specifically built for the task. Researchers wanted to see how well your typical AI agent handles the task of driving, via Comma vision hardware and OpenPilot. The results surely won’t leave anyone racing to replace themselves in the driver’s seat anytime soon, but they are eye-opening.
Aditya Ramabadran, Simon Mahns, and Tobias Gessler created DrivingBench: A simple test for Claude, GPT, and Grok models to navigate a short point-to-point course around a parking lot with a few curves, outlined by small cones. Each agent was issued the same command, with the same guidelines and directions to share what they’re seeing and doing after every step, as well as full control of a Toyota Corolla. They regularly stop along the way, because the commands are being sent away to data farms over a wireless hotspot, and all of this happens within one continuous chat. They have three attempts to get it right.
What the DrivingBench crew discovered is that frontier models, frankly, are not good at this. Of the 11 runs conducted across four different agents, eight failed to complete more than 11% of the course, crashing out in the first turn. Each session has a video attached to it, with the chat receipts scrolling in real time so you can see the AI working through it—or, at least, trying to.

Grok initially thought a gap between the mini-cones forming the outer boundary was a gate it was meant to pass through; it drove straight off in its first attempt, ending after two commands. One GPT model failed in part because invented a rule that cones on one side of the course were all the same color, which the initial, human-written command explicitly warned was not the case. And many of the agents simply failed to turn enough, because they couldn’t work out how much steering angle was required to round the bends.
Reading through these one-sided conversations is fascinating, because as clever as these agents can be at certain tasks, they’re completely in over their heads when it comes to driving. All except one, anyway, because GPT-6 Astra was the only agent to complete the course fully, on its second attempt. The route it took was a little wonky as you can see below, heavily hugging the long right-hander before almost driving off the course on the outside just before the finish box, but a win’s a win.

Per DrivingBench, Astra’s successful run cost $7.74—nearly four times the charge (and way more tokens) than its prior attempt that made it half as far. Some say AI will kill us all; this experiment demonstrates that a very fast way to make that happen is to let the world’s “best” agents drive our cars.
Got a tip? Reach out to tips@thedrive.com
Loading comments…
Comments couldn’t be loaded. Please refresh the page.