Saturday, August 01, 2026

The Half-Life of Intelligence

 In physics, a half-life is the time it takes for half of an element to decay. The word can be lent to business: the half-life of a technological advantage is the time it takes for half of its excess profit to evaporate.

For pharmaceuticals that is roughly twenty years, held open by patents. For leading-edge chips, about two years, held open by process. For AI, the measured value is six months.

And this six-month thing is swallowing the largest investment in human history. This year, Microsoft, Google, Amazon, Meta, and Oracle will spend somewhere north of seven hundred billion dollars in capital expenditure between them, close to double last year, and that money takes about four years to earn back. Six months against four years — this article is about that division, and its remainder.

1. An Unusual Death

On July 30, Wall Street's most celebrated fund manager of the year sold roughly sixteen billion dollars of stock to Citadel at a discount, settled within twenty-four hours.

Leopold Aschenbrenner, formerly of OpenAI, wrote the widely circulated AGI essay *Situational Awareness* two years ago and then founded a fund of the same name. His net return for the first half of the year was 439%, and assets under management climbed to a peak of about forty-five billion dollars. In July, the AI hardware names he was long — SK Hynix, Micron, CoreWeave — fell thirty to forty percent, the software names he was reportedly short rose instead, and with roughly four times leverage he lost 67% in a single month, forced to liquidate his entire public book down to about ten billion.

One correction to a widely repeated claim: he did not go to zero. Counting the first half's gains, the fund is still up 80% on the year, and his roughly five billion dollars of Anthropic shares were never sold.

Two things about this blowup are unusual.

First, July's decline does not look much like a problem with fundamentals. The worst performer, KLA, fell 43.6% in a single month — the worst month on record, worse than the month of Black Monday in 1987 — with no bad news attached. ASML and TSMC both beat expectations and raised guidance, and fell anyway. Intel, which has almost no AI revenue, fell 41%. Korea's KOSPI posted its worst month ever (-23%) and then rose 17.9% in a single day on July 31, its largest one-day gain ever. Deteriorating fundamentals do not reverse in forty-eight hours. A crowded position unwinding does.

Second, and this is the part worth sitting with: he was not wrong about AI, he was wrong about the business. AI needs enormous quantities of chips, memory, and power — that judgment still holds today, as the five big spenders raising capital expenditure straight through a market rout will attest. Where he went wrong was in assuming that being right about the technology meant being right about the business, and that this business would return his money before his leverage came due.

That mistake is not his alone. To understand it, start with a question: how long is the window in which AI actually makes money?

2. Measuring the Half-Life

The answer can be measured, and by more than one ruler.

The first ruler is imitation lag. Epoch AI tracks how far open models trail the closed frontier by asking when equivalent capability appears: currently about four months. Which is to say, the lead you bought with billions of dollars has a free substitute four months later.

The second ruler is price. At a fixed capability level, prices are falling by a median of roughly fifty times a year. In 2023, GPT-4 cost thirty dollars per million tokens; today a model of equivalent capability is nearly free. The fiercest competitor in a price war is your own previous version.

The third ruler is simply the week that just ended. On July 27, Moonshot released the weights for Kimi K3 — the strongest open model available, trailing the closed frontier by about one release cycle. Three days later, OpenAI cut the price of its low-end Luna tier by 80%. A day after that, DeepSeek shipped a new V4-Flash: price unchanged to the cent, capability sharply higher, beating its own flagship on all nine published benchmarks.

Note the shape of this price war: only the bottom is being cut. OpenAI's flagship Sol did not move at all — over the past twelve months, its list price has in fact risen three to fourfold. The logic is plain. Where someone has caught up, you compete on price; where no one has, you collect rent. The window on the low end has closed and prices have fallen to cost; the window on the flagship is still open, and precisely because it keeps narrowing, the rent is being collected in a hurry.

Why is AI's window shorter than that of any technology before it? Because intelligence is the first product that helps competitors copy it. Distillation can teach a small model from your outputs; synthetic data can turn your capability into someone's training set; AI is itself accelerating AI research. Patents cannot stop this and neither can process moats — the product comes with its own reverse engineer.

What does a six-month window mean? It means a single generation cannot earn back its own cost. Epoch estimates that GPT-5 generated about two billion dollars in gross profit in its first four months, against roughly five billion in R&D in the four months before launch. Less than half the investment came back inside the window, and then the window shut.

3. The Cost of Staying Critical

Now put the two numbers side by side: a six-month earning window, a payback period of about four years.

The ratio invites comparison. Pharmaceuticals: twenty-year patents, ten-year payback, ratio 2 — which is why pharma is an independent, high-margin industry. Leading-edge semiconductors: about two years of process lead, four to five years to pay off a fab, ratio 0.5 — the books barely balance, which is why only three companies in the world still do leading-edge, each rolling one generation's profits into the next. AI: half a year against four years, ratio 0.1.

A ratio of 0.1 means the books do not balance on their own. This cannot exist as an independent industry. Something else has to pay for it.

Look at the table and it is obvious. Microsoft pays with Office and Azure, Google with search advertising, Amazon with retail, Meta with social advertising; behind the Chinese labs sit industrial capital and the state. It resembles China's ride-hailing wars, where Didi and Kuaidi were both burning money that belonged to Alibaba and Tencent — in the end, the contest was over which patron blinked first. On the surface these are models fighting models. Underneath, it is several money printers seeing who can outlast the others.

The last week of July put this on public display. Microsoft rose 16% the day after earnings, adding about four hundred and fifty billion dollars of market value in a single session, a record — because Azure growth accelerated to 43% and contracted-but-undelivered bookings reached six hundred and seventy-eight billion, up 84%. Its war spending turns into rent on the way out. Amazon raised full-year capital expenditure to two hundred and twenty billion and the stock rose anyway, on AWS growth of 36.7%, the fastest in eighteen quarters. Meta fell nearly 10% after hours: revenue growth of 28% was fine, but free cash flow was down to seven hundred and eighty-four million, capital expenditure was raised again, and all of that compute is for its own use — not a dollar of it comes back as rent.

The only thing the market was actually weighing that week was who can afford to keep paying.

One more detail worth recording. July erased more than a trillion dollars of chip-stock market value, and not one company cut a single dollar of capital expenditure in response — three raised it in the same week. The reason is not complicated: this build-out does not run on the stock market's money. The five big spenders used to fund capital expenditure out of operating cash flow; that ratio has risen from a ten-year average of 40% to above 90% this year, and the gap is starting to be filled with debt. Since the stock market is not paying, its moods do not govern, and a two-day round trip after a crash is what you would expect. Only three things can actually govern this build-out: power (large transformers are quoted at more than a hundred weeks), politics (Virginia has legislated a per-kilowatt-hour tax on data centers), and the ledger — which is the last section.

4. Looking for a Stable Isotope

The window cannot be held and the spending cannot stop. So where is the way out?

Somewhere old: habit. A technological lead has a half-life of six months; a user habit's half-life starts at ten years. Microsoft is the ready example — DOS lost its technical lead long ago, and Office has been eating on habit for thirty years. So every lab is now doing the same thing: while still ahead, convert the technological lead into user habit — subscriptions, tooling, ecosystem, anything that makes switching feel like a chore.

One number explains why it is habit and not price. OpenRouter analyzed a hundred trillion tokens of real traffic and measured the price elasticity of demand for intelligence at 0.05 — cut prices 10%, usage rises 0.5%. Users barely look at price; they look at capability and convenience. Read the number in reverse and it gets more interesting: in a market where cutting prices does not buy customers, a price cut has only one explanation left — you had no other cards. Pricing power became a test strip. Whoever still dared to raise prices this year (Anthropic's top tier, Kimi, Zhipu) has something customers cannot swap out; whoever can only cut has already been substituted.

The progress of that conversion varies a great deal. OpenAI is betting on consumers: nine hundred million weekly users is its largest chip, and the wager is that ChatGPT becomes a daily habit. Anthropic is betting on enterprises, using Claude Code and agent tooling as lock-in; on OpenRouter it takes 42% of revenue with 11% of the calls, and it has confidentially filed to go public — that prospectus will be the first formal verdict on whether the model layer can stand as an industry at all. xAI is the counterexample: three dollars lost for every dollar earned, with neither consumer habit nor enterprise lock-in. That is what failing to make the conversion looks like.

The application layer's story runs against intuition. Cheaper models were supposed to benefit applications; in practice, coding assistants are nailed to the most expensive flagship models, the flagships went up fourfold in a year, and tokens consumed per task rose several times over — squeezed from both ends. GitHub Copilot moved heavy usage to metered billing in June, explaining that no vendor can absorb an agent's unlimited consumption inside a ten-dollar monthly fee; Cursor pays forty to seventy cents in inference for every dollar of revenue. The comfortable ones are the incumbent software companies holding workflows and distribution: Salesforce posted a record operating margin and built its agent product to $1.2 billion in annual revenue, charging agents a toll on an installed base it already owned. Software stocks rallied together in July while Fiverr crashed on a 10% revenue decline, and the dividing line sits exactly there: selling workflow survives, selling human hours does not. What Fiverr sells is human time, and human time is precisely what AI is now wholesaling.

Chinese players took a third route: skip the conversion and overturn the table. Your profit depends on a lead, so make the lead itself worthless — publish the weights, let anyone use them, and move the fight to ground I am better on: power, manufacturing, deployment. This move was run once before, in solar, and the ending is worth studying. China did win the entire industry, module prices fell 95%, Western manufacturers exited altogether. But the winners have not had an easy time of it: in 2024 Tongwei lost more than seven billion yuan and LONGi eighty-six billion; the first generation of champions — Suntech, LDK, Yingli — all collapsed along the way. On the largest third-party model routing platform, open models already account for half of all calls but only about 4% of the money — the territory was won, the profit did not follow. What that territory can eventually be traded for, solar answered with leverage over energy; AI has not answered yet.

5. Where the Decay Ends

Finally, the ledger, which sets the timetable for all of this.

Over the past four quarters, Microsoft, Google, Amazon, and Meta spent $433.9 billion in capital expenditure between them, while depreciation booked to their income statements over the same period was only about $149 billion — barely a third of the money already spent shows up in the accounts. The other two thirds does not disappear; it is queuing. On rough equipment-life assumptions, combined annual depreciation for these companies goes from around $160–170 billion this year to roughly $250 billion in 2027, and closes on $500–600 billion by 2029.

For that expanding depreciation not to eat existing profits, AI-related revenue would need to grow by roughly five hundred billion dollars over today's level by 2029 — sixty to seventy percent a year, with no stumble in between. Can it? Nobody knows. Demand is genuinely accelerating right now: Google Cloud up 82%, Azure up 43%, compute sold out across the board, data center vacancy at 1%. But "accelerating right now" and "four consecutive years without a stumble" are different claims.

If something breaks, it will likely start with the most fragile funding structures and work inward. The first ring is leveraged funds, marked to market daily; that one already blew in July. The second is the compute landlords who bought their cards with debt — CoreWeave's interest expense already consumes a quarter of revenue, and the coming year is its refinancing test. The innermost ring is the five big spenders themselves, who have the thickest cushion but the least flexible depreciation calendar, concentrated in 2028 and 2029.

None of this needs memorizing. Three numbers are enough to watch. First, DRAM contract prices, which set costs, bottlenecks, and the room available for inference price cuts all at once; they rose ninety percent quarter over quarter in the first quarter of this year, and the day they turn, this article needs rewriting. Second, the "leases that have not yet commenced" line in Microsoft's filings — $329.1 billion now, against $92.7 billion a year ago; when it turns down, that is the first sign of the build-out slowing. Third, whether the market can still round-trip a crash in two days — the day the V-shaped rebound fails is the day the borrowed money has grown large enough for a decline to become self-fulfilling.

Back to Aschenbrenner. He has not conceded anything since the blowup, and has not sold a single Anthropic share — he presumably still believes AI will change everything, and that judgment has never been wrong. What was wrong was something else: he assumed that being right about the future meant the future would pay him.

Friday, July 03, 2026

The Iceberg Is Already Here, but Most People Are Still Watching the Surface

 At the end of 2022, ChatGPT had just appeared. It still felt like a toy.

By 2025, it had cut through the entire software industry.

Only two years sat between those moments.

In 2023, you could throw a few questions at it. The answers were rough, brittle under follow-up, like a toy that happened to talk.

In 2024, it started taking real work. AI began showing up in the daily flow of work.

By 2025, it was writing code, fixing bugs, producing tests, and drafting documentation. The software industry had nowhere left to hide.

And this transformation is only just beginning.

Software Was Only the First Cut

AI's first blade fell on its own makers.

Software was closest to AI, so software was cut first.

But the blade will not stop there. It is already moving elsewhere.

The first to panic were creators. After many Chinese game companies adopted AI art tools, concept artists were among the first to feel the cuts. Voice acting, editing, and subtitles are also disappearing piece by piece.

This is not just a Chinese story. In 2023, Hollywood writers and actors went on strike, and one of their central demands was that studios should not use AI to replace them. Meanwhile, actors in Hengdian Film City have found the work drying up.

Writers have not escaped either. Web fiction platforms have already embedded AI writing tools. AI-generated novels and audiobooks are being listed in batches. Academic journals, too, are struggling with submissions written by AI.

The list will only grow. Customer service, translation, design, legal work, finance: anything that can be written down and turned into a process is being replaced.

People who can use AI well, and jobs that can be replaced by AI, are starting to separate into different layers.

Digitization Is Horizontal; AI Transformation Is Vertical

This is not the first technological revolution. The last one was called digitization.

The two look similar on the surface. Underneath, they are very different.

Digitization is horizontal. It forces you to rebuild the foundation: change processes, replace systems, install ERP, move to the cloud, remake the company from head to toe. Many things have to move at once. The cost is high, the risk is high, and the payoff is slow. Every department has to relearn its work; every process has to be rewritten. One wrong step can drag down the whole project. From the 1990s to the 2010s, simply spreading that foundation took twenty to thirty years.

AI transformation is vertical. It does not need to move the whole chessboard. It can cut into one small link first: write a block of code, fix a bug, cover one customer-service shift. The investment is low, the risk is low, and the result appears quickly: one person, one tool, one week to see whether it works. If it fails, you pull it back. There is almost no sunk cost. Digitization took twenty or thirty years to grind through the economy. AI can reshape an industry in two or three.

Why is AI moving faster than anyone expected? Because the foundation was already there. Those decades of digitization built the data, systems, and networks. AI does not have to rebuild much. It slips into existing systems and lands directly on the part that does the work.

Can traditional industries escape? No. The slowest and most expensive part of digitization was moving the real world into systems one record at a time. AI is not picky: a photo, an audio clip, a stack of handwritten forms. It can read them directly. In other words, AI can fill in that missing foundation by itself. For industries that never fully digitized, digitization and AI transformation are now arriving in one step. The thickness of the foundation only decides who goes first. It does not decide who gets spared.

AI does not tear up the foundation. It begins by replacing a door, then a window. Before you notice, AI agents are already at work in the background.

The Smoother It Feels, the More Dangerous It Is

Instinctively, "smooth" sounds like a good word. No pain. No disruption. Reassuring.

But think about replacing a window.

When you rebuild the foundation, the whole building shakes. Everyone can see it. Even when digitization eliminated jobs, it first had to install new systems. That gap gave people a little time to breathe, switch roles, and learn a new skill.

But replacing one window does not require evacuating the building. It may not even alert anyone.

What AI smooths away is exactly that buffer. No friction. No negotiation. It simply takes over the person doing the work, without making a sound.

That is the problem. Society's buffer mechanisms, retraining, policy support, psychological preparation, were all built for slow change. They assume transformation will move like digitization: slowly enough to absorb over twenty or thirty years.

This round is too fast and too smooth. Before the buffer zone is built, the people have already been replaced.

We were not unprepared because we failed to see it. We were unprepared because this time it moved too fast to leave us time.

The Iceberg Has Arrived

Ten years from now, perhaps only one-tenth of people will still be working. The direction is already clear: there will be fewer and fewer jobs that do not need AI.

Three years ago, ChatGPT was still a toy. Today, it is chewing into every industry.

The iceberg is here. Most people still cannot see what is below the surface.

Tuesday, June 02, 2026

Twenty Years of Compute: From Virtual Machines to Intelligence

 

Over the past twenty years, compute has mostly done one thing: moved profit up the stack — first past machines, then past software, and now past large models themselves.

Cerebras has just gone public. In intraday trading on its first day, the stock nearly doubled, and its market cap briefly touched about $95 billion.

On the same news tape, Nvidia paid roughly $20 billion at the end of 2025 to absorb Groq’s core team. OpenAI and Anthropic now raise $10 billion-plus rounds. Capital has cast its vote almost entirely for two layers: companies that train large models, and companies that build GPUs and inference chips.

This looks like the endgame. It is more likely the high-water mark of a compute cycle repeating itself. Twenty years ago, the scarce thing was the machine. Later, profit climbed to the cloud, then to software, then to applications. The two layers that look most profitable today will probably not hold high margins five to ten years from now. Profit will move from “making intelligence” to “using intelligence.”

1. The Old Script: The Bottom Layer Always Gets Commoditized

Cloud computing has replayed the same script again and again: the layer that is scarce, expensive, and most profitable today becomes standardized and commoditized tomorrow, handing its premium to the next layer built on top.

In August 2006, AWS launched EC2. Compute moved from the heavy fixed asset of self-built data centers to an “instance-hour” utility meter. SaaS then turned software into “seat-month” subscriptions, with operations, upgrades, and security patches swallowed by vendors. Every time the stack grew upward, the layer beneath it was commoditized once more, and money moved up.

Over the past twenty years, compute went from machines to cloud, and from cloud to software. Over the next decade, intelligence will move along the same chain. The billing unit has already climbed from one token, to one action, to one solved problem. The only question is where this round of commoditization stops. My answer: it burns all the way through large models and GPUs themselves.

2. The Two Most Expensive Layers Are Being Hollowed Out by Their Own Builders

For intelligence to become cheap, its power bill has to collapse first. That is already happening, and the people building it are driving the collapse themselves.

The first nail is convergence. GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro all crowd into the 93%-94% range on GPQA Diamond. Benchmarks this close to saturation can no longer support a durable premium.

Move to engineering tasks, and the leaderboards spread out again.

OpenAI reports GPT-5.5 at 58.6% on SWE-Bench Pro. When Anthropic introduced Opus 4.8, it emphasized workflow benchmarks such as SWE-bench, Super-Agent, and CursorBench. The point is not that “all models are the same.” The point is that a general foundation model is becoming harder to rent out at a premium on the strength of a single benchmark.

The second nail is price collapse. According to Epoch AI, the median inference price needed to reach the same benchmark score has been falling by roughly 50x per year. DeepSeek R1 entered the market and cut comparable inference pricing by about 96%, forcing OpenAI, Anthropic, and Google to follow with steep price cuts across comparable models within a year.

The irony is that high-speed inference, one of capital’s favorite lanes, is also pushing prices down. Capital is pouring tens of billions into “faster, cheaper tokens,” but that is exactly how the intelligence layer becomes commoditized. The more successful this commoditization becomes, the more intelligence looks like electricity, and the less likely it is to command high prices for long. Today’s hot money is rushing into the layer that may be least profitable in the future.

The third nail is open source. Open source does not need to match the frontier. It only needs to be “good enough” in more and more commercial settings, and the map of “good enough” expands every year. The frontier still leads by 6–12 months in long-tail areas such as agents and long-horizon coding, but that long tail keeps retreating. Frontier premiums are temporary rent. What can command a high price is always “today’s frontier,” not “yesterday’s frontier.”

3. Inference Is Moving Down to Edge Devices

The second pressure-release valve is local inference.

Small open-source models that run on phones and laptops are getting stronger with each generation. Llama, Qwen, Gemma, Phi, and other models with only a few billion parameters, paired with on-device NPUs, can already handle a meaningful share of daily work: summarization, rewriting, classification, and local question answering. Apple, Qualcomm, and Intel are building NPUs into every main chip, effectively preinstalling a slice of free inference across more than a billion devices.

Once inference can run for free on your own device, the cloud-side tollbooth that charges by the token loses much of its grip. Cloud and GPUs will still carry training and the heaviest frontier inference, but massive volumes of “good enough” daily calls will sink to the device. Add convergence and falling prices, and the conclusion is blunt: the marginal profit of making intelligence is being squeezed toward zero from both ends.

4. The Valuable Thing Was Never the Model. It Was Using It Well.

So where does the profit go? Start with a fact that keeps reappearing: enterprise spending on models often does not pencil out. Spending on wiring models into the business does.

In MIT NANDA’s August 2025 survey, roughly 95% of companies had spent money on generative AI, yet could not see measurable returns in the P&L. The methodology was debated, but the direction matches multiple CIO surveys: the failure is not that the models do not work. It is that they do not fit into workflows.

The two most awkward examples came from two of the most aggressive companies. At the end of 2025, Microsoft rolled out Claude Code across its Experiences and Devices group. Less than half a year later, it adjusted the deployment and moved some engineers back to its own GitHub Copilot CLI. Uber rolled Claude Code out to thousands of engineers, burned through its entire annual AI coding budget by April 2026, and saw its COO publicly question the ROI. These companies were not avoiding AI. Quite the opposite: they pushed AI to the limit early, and therefore hit the “the math does not work” wall early.

An e-commerce ERP CIO I know put it more directly. Last year, they launched 12 AI pilots. The board applauded the demos. Six months later, finance kept only an Excel plugin. The other survivor was invoice extraction plugged into an existing OCR pipeline. The other 10 did not fail because the model was bad. They failed because they could not be wired into the ticketing system, could not pass compliance review, or had no cross-functional owner. “The model answered correctly. The organization could not absorb it.”

That sentence reveals where the profit goes. Models themselves are getting cheaper and more similar. What remains truly scarce, truly hard, and therefore truly valuable is the ability to embed them into real processes and solve a real problem. That work lives at the application and software layer, not the model layer.

5. Intelligence Is Electricity. So Where Does the Road Lead?

The dream of “selling intelligence like electricity” has been around for more than sixty years. In 1961, McCarthy said in MIT’s centennial lecture that computation would one day be sold to everyone like a public utility. In March 2026, Altman said almost the same thing at BlackRock’s infrastructure summit, with intelligence substituted into the line.

The history of electricity contains a fact people often miss: power generation has never been the most profitable link. When electricity prices approach commodity levels, power plants earn thin regulated margins. The real profit goes to the people who use electricity to do things: factories, appliances, and the whole modern economy.

If intelligence becomes electricity, model companies are power plants, and GPUs are generators. Profit will not stay in their hands. It will move to applications and software that use intelligence to solve real problems.

Some will say that AWS turned servers into a commodity and still became Amazon’s profit engine. Correct.

In 2025, AWS reached $128.7 billion in revenue and $45.6 billion in operating profit, but not because it sold a single premium point of “the strongest compute.” It won through scale, lock-in, ecosystem control, and long-term operating leverage. The best destination for model companies may also look more like an “AWS of intelligence” than like a business that rents out “today’s strongest model.”

As for who captures the application-layer value — whether a new “App Store” appears as it did in the mobile era, and who owns it — that is still an unsettled wager. But one thing is already clear: it will not automatically belong to today’s model vendors or chip vendors.

The next AI giant does not necessarily need the strongest model, nor the most GPUs. It only has to answer one question: who can plug cheap intelligence into the expensive real world?

Thursday, May 21, 2026

Why Do We Need Each Other? Economics' Final Question

 In Book I, Chapter 1 of The Wealth of Nations, published in 1776, Adam Smith described a pin factory. An untrained worker making pins alone might not make even one pin a day, and certainly could not make twenty.

Wednesday, May 20, 2026

The Last Interface: Where Human-Computer Interaction Ends

 One late night in 1965, a programmer walked toward the machine room with a stack of punched cards in his arms.

Each card had 80 columns, one character per column. The whole stack might have carried less information than a text message today. He handed the cards to an operator in a white coat and went back to sleep. The result would arrive the next day. If one character was wrong, the whole day was wasted.

One late night in 2026, you open your computer and say: help me turn these 50 emails into a business trip plan. Draft the plan first. Ask me before booking tickets or sending messages.

The computer starts researching, arranging the schedule, editing spreadsheets, and finally lays out the actions waiting for your confirmation.

The goal is the same: make the machine do something for you. What changed is everything that used to sit between “your intent” and “the machine’s execution.”

From punched cards to today, the history of human-computer interaction is the history of removing that stack, layer by layer.


A History of Removing Intermediaries

In the punched-card era, humans served the machine. You could not touch the machine directly. You had to translate what was in your head into physical holes the machine could read, then hand those cards to a specialized class of operators who fed the machine for you. ENIAC, publicly unveiled in 1946, was even more extreme: “programming” it meant wiring logic into the machine with cables and plugboards. One rewiring job could take days. People were not using computers. They were tending to them.

In 1961, MIT demonstrated the CTSS time-sharing system. Terminals gave more people their first taste of using a computer in something close to real time. The machine began to respond to you. But the command line still forced humans to accommodate the machine: artificial language had to be memorized, command names, syntax, and parameters all had to be exact, and one wrong character meant an error. On the screen there was only a blank cursor. If you did not already know what to do, you had nowhere to begin.

In 1968, Engelbart demonstrated the mouse. Then came Xerox PARC, and then the Macintosh in 1984. The graphical interface introduced the desktop metaphor: it wrapped the alien computer in the familiar shape of an office desk, with files, folders, and a trash can.

The selling point of the Lisa was almost this simple: if you can recognize the trash can on an office desk, you can use this computer. This was the first time the machine actively borrowed a human mental model to accommodate the human. Interaction shifted from recall to recognition.

In 2007, the iPhone removed the mouse - the proxy pointer on the screen. Your finger landed directly on the content.

Now natural language is removing the last fixed intermediary: the controls you must first learn, locate, and understand. You no longer need to find the button first. You just say what you want.

Punched cards -> command line -> graphical interface -> touchscreen -> natural language. Every generation of interface has done the same thing: remove one layer of machine language that humans had to learn. In “The Battle for the Desktop: Who Will Take Over Your Computer,” I wrote that every leap has moved in the same direction: lowering the cost of human adaptation to machines, while increasing the machine’s ability to understand humans. The next question is obvious: if this curve keeps going, where does it end?


The End of the Curve: Not Zero, But One

Start by seeing an interface as a translation layer. It exists for one reason: humans and machines speak different languages. Punched cards, command lines, icons, and menus are all translations, each generation easier to understand than the one before.

So what happens when machines can understand human language directly?

The fixed translation layer loses its reason to stay permanently in front.

We have chased the dream of “operating machines by speaking” for a long time, and failed for a long time. Siri arrived in 2011. Alexa arrived in 2014. For more than a decade, the high-frequency uses of voice assistants remained concentrated around music, weather, timers, and alarms. The reason was simple: before large models, voice assistants depended on predefined skills or intent systems. Your wording had to fall into a slot they had prepared. They were not fully understanding you. They were matching you.

Large models changed exactly this. They can follow context, infer intent, ask follow-up questions, and remember what came before. For the first time, natural language is qualified to become something close to a complete interface, not merely an input shortcut.

But the idea that “the interface will disappear” is not new. In 1991, Mark Weiser wrote in Scientific American that the most profound technologies are the ones that disappear. In 2015, designer Golden Krishna wrote a book titled The Best Interface Is No Interface. More than thirty years later, that ideal has not arrived. What we got instead was more and more apps, and hundreds of icons inside a phone.

So my prediction is different from these earlier visions. The endpoint of the interface is not “zero interfaces.” It is “one interface”: one supervisable agentic entry point that can carry authorization and responsibility.

Why one?

First, convergence is a recurring script in the history of technology. The smartphone did not make devices vanish. It made devices converge. It swallowed the point-and-shoot camera, GPS navigator, MP3 player, calculator, flashlight, voice recorder, and paper map in one sweep. In CIPA data, global shipments of built-in-lens cameras fell from about 109 million units in 2010 to about 3.58 million units in 2020, a roughly 97% collapse in ten years. General-purpose platforms defeat special-purpose devices not because they are best at every single thing, but because they are more convenient as a whole.

Second, the key that makes “one interface” technically plausible is generative UI. In the past, “one entry point” meant “limited functionality,” because interfaces were drawn in advance by designers. Now interfaces can be generated on demand. Google has already shown early versions of this in Gemini 3-related products: AI Mode can generate interactive tools and simulations based on a query, and experimental views in the Gemini app can create one-off interactive interfaces from prompts. Need a slider? A slider appears. Need a table? A table appears. Use it, then discard it.

“Only one agentic entry point” no longer means “only a chat box.” Some front doors of specialized apps can retreat into the background and become tools and APIs called by the agent.


That One Interface Is a Cockpit, Not a Chat Box

But “one interface” does not mean “one chat box.” That is the easiest trap to fall into right now.

Today, the mainstream way we interact with AI is the text dialogue box. ChatGPT, Claude, and Gemini all look like this. But more people are pointing out something uncomfortable: the chat box looks a lot like the command line coming back. A blank input field, a blinking cursor, and you must invent what to say and discover through trial and error what it can do. Is that not the old command line problem all over again? The graphical interface worked so hard to move interaction from recall to recognition with menus and icons. Now we have returned to an empty box that tells you nothing. Some simply call it “a command line wearing natural language as a costume.”

To understand what it should become, we need to start from a fact that is badly underestimated: language is a low-bandwidth channel. A study across 17 languages found that the information rate of human speech is almost constant, at roughly 39 bits per second. The bandwidth from the retina to the brain is on the order of tens of millions of bits per second. The two measurements are not directly comparable, but the gap is already more than five orders of magnitude. This is the quantitative reminder behind “a picture is worth a thousand words.”

On the input side, language is excellent for expressing goals. One sentence is enough. On the output side, the machine must return analysis, data, and plans; you still need to inspect, compare, and continuously adjust. Language is too slow. Output must rely on vision.

So the last interface is a hybrid: you express intent in language, and it presents results visually. Generative UI summons the controls needed for the task. An entry point that listens, plus a canvas that changes on demand.

But that is only the form. The crucial change in the last interface is not its form. It is its nature.

Every interface before this - from punched cards to touchscreens - was an operation panel. You click once, the machine responds once. Control and responsibility remain in your hands at every step. An agentic interface is different. You no longer operate. You delegate. You state a goal, and it breaks that goal into a chain of actions and executes them. What you hand over is not an “instruction.” It is an “intent.”

This means the last interface is not fundamentally about input. It is about trust and supervision.

When an agent runs a long chain of actions you have not reviewed one by one, what you really need is not a brighter button. You need four things: visibility into what it plans to do, so the black box becomes a glass box; the ability to stop and correct it when it goes off course; a way to verify afterward that it did the right thing; and boundaries that define what it may decide by itself and what it must ask you before doing.

HCI already has a useful vocabulary for this: the human role is moving from human-in-the-loop, stuck inside the loop approving each step, to human-on-the-loop, standing above the loop and supervising it. In “The Reins of Artificial Intelligence and the Return of Cybernetics,” I wrote that engineers are putting a precise set of reins on large models. Those reins are attached to the machine. The last interface is the other end of the reins - the end held in human hands.

It is no longer a panel. It is a cockpit.

Of course, graphical interfaces will not disappear. Ben Shneiderman’s idea of “direct manipulation,” proposed in 1983, still holds. Continuous, spatial, and fuzzy intentions are inherently hard to express through the one-dimensional channel of language. The command line is not dead either. Programmers still live in terminals every day. Convergence does not mean extinction. What will disappear is the current mode in which every graphical interface governs its own separate front door. They will retreat behind the agent, be summoned when needed, and fold away when finished. One entry point in front; graphics still alive behind it.

As for who controls that entry point, and what power structure will emerge from interface convergence - that belongs to another essay, “The Battle for the Desktop.” Here I only want to make one point clear: once interfaces converge, their nature changes.


Two Late Nights: A Spiral, Not a Loop

At this point, the whole thing may feel absurd.

Human-computer interaction has worked for more than sixty years. From the 80 columns of a punched card, to the text box of the command line, to the countless icons on the desktop, to touchscreens, and then to… an entry point that looks like a text box. We have made a huge circle and returned to a blinking cursor.

But this is not a loop. It is a spiral.

The shape is similar: both are boxes waiting for input. But the direction is completely different. The command-line box required you to learn its language: memorize commands, remember syntax, and fail on one wrong character. The agentic entry point means it learns your language: say it however you like, and it follows your meaning.

After more than sixty years, we have not returned to the starting point. We have returned to the space above it.

Back to the two late nights at the beginning.

Science fiction had already drawn this scene long ago. In Star Trek, crew members could look up and say “Computer” to ask any question or issue any command. The Alexa team later publicly acknowledged that this shipboard computer was one of their original inspirations. In 1987, Apple produced the Knowledge Navigator concept video, in which a conversational assistant could ask follow-up questions and manage your schedule. Frame by frame, it almost predicted today’s Siri lineage. Then came Samantha in Her: no screen, only voice.

But these stories almost always carry a trace of unease. HAL in 2001: A Space Odyssey moves from a gentle voice to lethal intent. Samantha in Her, at her most intimate with you, quietly evolves beyond the range you can follow, then leaves. The more natural and intimate the conversational interface becomes, the heavier the unease grows: are you using it, or depending on it; are you supervising it, or is it taking care of you?

In 1965, a human carried cards and served the machine. In 2026, a human speaks one goal, and the machine begins to run the process on the human’s behalf. On the surface, the master-servant relationship has been completely reversed.

But when the last interface becomes an agent you must constantly supervise, authorize, and calibrate - while depending on it more every day - has that reversal really happened?