Databricks gave 3,500 engineers GPT-6 Astra. The bill rose 60%
2026-09-17 · 4 min read
On Wednesday, Databricks co-founder Patrick Wendell posted the numbers most companies keep to themselves. @pwendell wrote that Databricks had rolled out GPT-6 Astra "to every engineer at Databricks (N=~3500)," then listed what the company learned. The post picked up more than 2,000 likes and 900 bookmarks in a day, which for a spend memo is a lot.
Two lines carry the post. Astra "unambiguously out performs" the models Databricks used before, Anthropic's Opus 5 and OpenAI's Sol 5.6, on highly complex work like system design and long tasks that cut across a codebase. And engineers who got it "increased overall coding spend by around 60%."
What Databricks found
Wendell's third point is the one I'd tape to the wall. Databricks could not tell that Astra was any better on medium and low complexity coding. The team's guess is that existing models already handle those tasks about as well as they can be handled, so paying more for them buys nothing.
So Databricks did not make Astra the default. It piloted with about 200 engineers first, using its own Unity Gateway product to run cohort experiments on quality and cost. Then it gave every engineer a separate sub-budget for Astra, inside an overall budget they can spend across tools and models. The intent, in Wendell's words, is to use Astra "selectively on complex tasks" and cheaper models for everything else. Budgets get revisited regularly.
One caveat he added himself: Databricks has no solid comparison against Claude Fable 5.1, because it hasn't rolled Fable out widely over data retention policies. And Unity Gateway is a Databricks product, so the post is also a sales pitch. It is still more detail than large companies usually share about what a rollout costs.
The other reality check
A day earlier, Dan Luu pulled an older admission back into view. @danluu wrote that it was "interesting to see Yegge say he never successfully built anything with Gas Town."
Gas Town was Steve Yegge's multi-agent workspace manager, released on New Year's Day, and it drew a crowd fast. In an August essay, Yegge wrote that it was meant to be reusable but he "only ever wound up using it to build itself." A newer Claude model kept wanting to tinker with Gas Town instead of doing the job, and it never settled down enough to do real work.
Luu's point is that this matches what he wrote back in June: the popular orchestrators looked heavily vibe coded and unreliable at finishing tasks, and autonomous loops lose productivity over time unless a person checks in. "Turns out the author of the most famous one had the same issue," he wrote.
The third reality check is smaller. @sama posted Wednesday that "the main thing i was excited about launching this week will be next week instead," quoting his own promise of a big ship ahead of OpenAI's DevDay on September 29. Even the lab's calendar slips.
My read
Put the first two together and you get a fair picture of where AI at work stands this month. The frontier model is real, and it is worth paying for on the hardest slice of your work. It is also, by Databricks' own measurement, a 60 percent price increase that buys nothing on the routine rest. Databricks got value because it noticed the difference and built a budget around it.
Yegge's story is the same lesson from the other side. Elaborate agent systems are fun to build and hard to get value from, because reliability, not intelligence, decides whether a task gets finished. Luu's own answer is a simple loop he checks on every day, and he says that still beats leaving anything to run on its own. I don't mean that as a knock on the tools. The person checking in is still where the work gets finished.
What I like about Wendell's approach is how boring it is. Pilot with a small group, measure quality and cost separately, give people a cheap default plus a budget for the expensive thing, and revisit. You don't need Databricks-scale tooling for any of that. You need to decide, per task, what good enough looks like, the same cost-per-task question Google's Flash launch raised two weeks ago.
What this means for your business
If you're paying for one AI subscription per person and letting everyone pick the biggest model, you're running the experiment Databricks ran, without the measurement. Most of what a small business hands to AI is routine, like drafting the follow-up or summarizing the call. Those are the saturated tasks. A cheaper model, or a plain automation with no model at all, will usually do them just as well.
Save the frontier model for the few jobs where the answer changes with a smarter model, and know which jobs those are before the invoice tells you.
New Face Design's free process audit sorts your recurring tasks into exactly those buckets: what needs a frontier model, what a cheap model handles, and what needs no AI at all. Start here.