← All posts

AI wrote 5,500 lines in 2 hours. It couldn't check its work

2026-08-02 · 4 min read

A ten dollar experiment

Over the weekend, Andrej Karpathy posted the result of a small test. He handed Claude Opus 5 the opening paragraph of The Lord of the Rings, a one million token budget worth roughly ten dollars, and one instruction: render it in Three.js.

The model ran for about two hours and wrote around 5,500 lines of code that procedurally built the story as a 3D scene. Karpathy called the result "kind of janky but fun," and later posted a playable version of it.

His framing is what makes the post worth reading. We are leaving the era, he says, of testing a model by asking it for an SVG of a pelican on a bicycle. The interesting property now is that LLMs have "all the stamina and patience in the world." Work nobody would ever pay a human to do has become a question of how many tokens you are willing to spend.

That is the actual headline. Not that the output was good. That it existed at all, for the price of lunch.

The part worth writing down

Then Karpathy named the catch, and it is the most useful line in the thread. The model could build the world, but it could not really look at the world it had built. LLMs, he wrote, are not able to "efficiently and natively perceive videos or play games."

So Opus 5 checked itself the only way available to it: taking screenshots at various points, slowly and painstakingly. It got things wrong along the way. It shipped jank it never saw.

Read that as an operator rather than an engineer. Producing the work was cheap and tireless. Verifying the work was awkward and unreliable. The bottleneck did not go away. It relocated.

My honest read: this is the single most underrated fact about AI in 2026. Every conversation about adoption is about output. How much can it write, draft, quote, answer, build. Almost none of it is about the harder question of who confirms the output was right, and how they do it fast enough to matter. Karpathy just demonstrated that limit at the frontier of the field, using the best coding model available and an unlimited budget by normal standards.

Notice what OpenAI is actually selling

Two weeks before that post, OpenAI announced Presence, its enterprise product for putting voice and chat agents into real company workflows. The pitch is that agents can answer questions, use company systems, "take approved actions, and escalate to people when needed."

Look at what is being sold there. Not the model. Everybody already has the model. Presence is the harness around it: written policies, permission controls that limit what systems an agent can touch, escalation rules for handing a conversation to a person, and simulations that test the agent against edge cases before it goes live. OpenAI says it runs its own English phone support on Presence and resolves about 75% of inbound issues without a human.

It is also not self-serve. Deployments run through OpenAI's forward deployed engineers and a short list of integrators, under a limited availability program. The most capable AI company on earth looked at deploying agents into live business workflows and concluded that somebody has to sit down and scope it.

What this means if you run a business here

The lesson transfers cleanly to a five person company in St. Charles or Geneva. AI will happily generate your quotes, your follow-up emails, your intake summaries, and your appointment confirmations. It will do it at three in the morning without complaining. What it will not do on its own is notice when it got the price wrong, or misread which job the customer was calling about.

So the value is not in the generating. It is in the checking loop you wrap around it:

  • Where the agent must stop and get a human yes before anything leaves the building
  • What it is allowed to read and touch, and what stays off limits
  • A clear handoff path the moment a conversation goes sideways
  • A record you can actually review at the end of the week

That is unglamorous work. It is also the entire difference between automation that saves you six hours a week and automation that quietly costs you a customer.

If you want to know which parts of your operation are safe to hand off and which parts need a person in the loop, that is what our free process audit sorts out. No pitch, no software to buy. We map how work moves through your business now, and tell you honestly where automation pays and where it does not.

Karpathy spent ten dollars proving that generating is the easy half. The other half is still yours.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere