
How ChatGPT Work Turns Codex Into a Business Engine
This episode explores how non-developers unexpectedly turned Codex into a business operations tool, prompting OpenAI’s shift to ChatGPT Work and a new unified agent harness. It also digs into the post-app era, where AI accelerates execution but human taste, judgment, and validated progress still matter more than raw motion.
Chapter 1
The Codex Hijack and the Unified Agent Harness
William Palmer
Imagine walking into the offices of OpenAI in the spring of two thousand twenty six and discovering a heist in progress. But, uh, not a heist of computational power or model weights. No, this was an internal hijack of a developer tool. See, general business employees, we are talking strategic finance, marketing, operations, they had completely hijacked Codex, a tool built strictly for software engineers to write code. By June of two thousand twenty six, OpenAI realized that these non developers made up roughly twenty percent of the active user base of Codex, and get this, their usage was growing more than three times as fast as the actual programmers. They were using it to run daily business operations because, well, it gave them what felt like a superpower. This bizarre behavior forced OpenAI's hand. It led to a major internal reorganization with Codex leads Greg and Tibo taking over product, culminating in the July ninth, two thousand twenty six launch of ChatGPT Work. And within just two weeks, ChatGPT Work and Codex combined rocketed to ten million active users.
William Palmer
Now, if you have been following our series here, this transition is the ultimate validation of our execution harness thesis. Remember when we argued that raw model weights are becoming a commodity, and that the real competitive moat has shifted to the execution harness? This is it in the wild. The secret sauce of ChatGPT Work is not some brand new foundation model. It actually runs on the exact same underlying Codex agent harness. It is the same persistent sandboxing, the same local filesystem, the same incredibly tenacious tool orchestration engine. But OpenAI did something brilliant with the user experience. Instead of exposing raw line diffs, file structures, and Git repositories like they do for developers, they wrapped that same powerful engine in clean, consumer friendly visual spaces called artifacts. It is the exact same engine, just wearing a tailored business suit instead of a programmer hoodie. But, uh, this introduces a fascinating, almost existential tension in product design. How do you build a simple, clean interface for an agent that is capable of building literally anything? Developers, we want to see the gears turn. We want to see every tool call, every line of code, the full Git state. But a corporate strategist or a financial planner? They just want the result. They want the agent to operate completely under the hood. But here is the danger, and it is a massive one. If you hide the execution too much, you run headfirst into the sub agent cost explosion we warned you about. If the user cannot see what is happening, they might write a simple prompt and have no idea that, behind the scenes, their agent is recursively spawning dozens of expensive reasoning sub agents, burning through token limits and API costs in a massive, unseen chain reaction. You do not realize the engine has been redlining until the bill arrives.
Chapter 2
The Post App Era and the Motion versus Progress Trap
William Palmer
This tension is driving us directly into what I call the post app era. For decades, knowledge work has been trapped in these rigid, static primitives. You write a document. You build a spreadsheet. You present a slide deck. But ChatGPT Work is completely dismantling those boundaries. Instead of a human manually copy pasting data across three different isolated applications, you simply collaborate with the agent. You describe the outcome you want, and the agent dynamically codes a local, interactive, highly customized website to solve your specific problem. They are called Sites. Think about that. You are not just looking at a static table of numbers anymore. You are playing with an interactive retirement calculator that the agent coded for you on the fly. We are seeing a profound shift where the hundredfold more people who only consume code are suddenly, without realizing it, becoming active software creators.
William Palmer
This is a space Akshay Nathan knows intimately. He spent his career making software power accessible to non developers, first at Airtable and now leading productivity engineering at OpenAI. And his journey proves a fundamental truth of this new era: when technical syntax becomes trivial, the bottleneck of human productivity shifts entirely to personal taste, ideas, and curation. Anyone can now generate high fidelity financial models or beautifully styled web pages in seconds. But here is the catch. If you ask a model to bring you a genuinely new, groundbreaking idea, it completely chokes. It is mathematically bound to its existing training distributions. The AI can accelerate the execution of an idea to warp speed, but the original spark, the taste, the human judgment, that remains the rate limiting step. And that brings us to the biggest trap of the AI era, the difference between motion and progress.
William Palmer
I, uh, I think about this a lot on the race track. In my spare time, I love driving race cars. And when you are sitting on the starting grid, if you dump the clutch and just stomp on the gas, you can spin your wheels and create this massive, dramatic cloud of tire smoke. It looks incredibly fast. It sounds incredibly powerful. But you are actually stationary, burning through your expensive tires, and going absolutely nowhere. In the corporate world right now, we are seeing a lot of tire smoke. Teams are using AI to auto generate mountains of slide decks, thousands of lines of unnecessary code, and endless automated Slack updates. It is a massive amount of motion, but is it progress? Often, no. It is just noise. As leaders, we have to stop measuring productivity by proxies like the number of commits, tokens spent, or slide decks produced. We have to shift our focus to quality at bats. We need to measure the human filtered decisions, the moments of genuine taste, and the validated hypotheses that actually move a project forward. The execution harness can run the car at two hundred miles per hour, but we still have to decide which way to turn the wheel. Alright, that is a quick look at the agent frontier. Talk soon.