← 02 / Writing

Agent / Software

Software is frozen; the harness is alive

Starting with the Ship of Theseus, a discussion of harnesses that evolve at runtime and the relationship between software and agents.

Translated from the Chinese original first published on WeChat. The original is linked under Sources & links.

AI artwork: a silicon-based orca
AI artwork: a silicon-based orca

Prologue

The ancient Greeks had a famous paradox, the Ship of Theseus: if you replace every plank of a ship, one by one, until not a single original plank remains, is it still the same ship?

The Greeks asked this question by the harbor. Today it is our turn: if an agent, while running, replaces every one of its parts (model, tools, memory, permissions, shell), is it still the same agent? And what was it “originally,” anyway?

Let me set the answer aside for now and start with something that happened to me recently.

Not long ago I had dinner with the founder of a unicorn with annual revenue in the hundreds of millions of dollars. He told me they were going to use AI to build software for their customers, and he walked me through a product plan with great excitement. When he finished, I said one thing:

“If you are still building and using AI with the concept of software, that is old thinking. Old thinking!”

He froze and asked what I meant.

I said: however smart the AI in your roadmap is, what you ultimately deliver is still a piece of software with a version number, a fixed form that has to be downloaded, installed, and upgraded. You are treating AI as a faster worker on the assembly line, and it has never occurred to you that the product of the AI era may not be something that can be “delivered” at all.

What is old thinking? Old thinking says software must have version numbers, ship as installers, be abstracted into features and then frozen into specifications, and be copied in bulk to a hundred million users. Old thinking means stuffing AI into software: the interface, workflows, and architecture are the same as before, with only the engine swapped for a large model.

This old thinking did not come from nowhere; it used to be right. When productivity was limited, software could only be written line by line by programmers, and writing software was scarce, expensive labor. So versions had to be stabilized and copies made at massive scale, spreading the one-time development cost over a hundred million users to make the economics work. Copying was the only way forward in that era, and the version number was the contract that made copying reliable.

Today, both premises have changed. Software can be composed at runtime and built by agents themselves; “write once, copy a hundred million times” is no longer the only economics. The cost of composition approaches zero, and iteration and evolution replace copying.

1. Agents are no longer bound by their shells

Over the past year, agent products have sprung up one after another: Claude Code, Manus, Codex, WorkBuddy. They are all powerful. But they share something almost nobody notices: their shells are dead.

The “shell” here has two layers. The first is its form of existence: the installer, the application, a fixed deliverable that must be downloaded, installed, and upgraded. The second is the harness: the entire engineering system around the agent’s model. The word originally means tack and reins, the gear that lets a person stay in control while the horse runs.

Claude Code’s harness was hard-coded by engineers when they wrote it. You can attach Skills, MCP servers, and hooks, but you cannot change the core, because the core is the body. Both layers of the shell are dead: the installer locks the form of delivery, and the harness locks the runtime core. To change either layer, you have to wait for the next version.

How can a free soul be imprisoned by a body?

How can a free agent be shackled by software?

DeepSeek Harness changed this for the first time: it made the harness itself alive.

How is it alive? Answering that requires taking apart a paper.

On the evening of August 13, two hours after announcing an API price change, DeepSeek dropped DeepSeek Harness on GitHub under the MIT license. It passed ten thousand stars within half an hour. Alongside it came an 88-page paper co-authored by Peking University and DeepSeek, A Programming Paradigm for Spatiotemporal Composability. Its three authors are Yifan Shi, the creator of the chatbot framework Koishi; Wei Zhang, an associate professor at Peking University’s Institute of Software; and Tianyi Cui, the head of the Harness team.

The paper addresses one problem: if a system lets components be plugged in and out at runtime, and especially if it lets an agent plug and unplug parts of itself, how do you make sure it does not break itself?

First, some background. The paper cites a figure: of the top 100 extensions in the VS Code marketplace, 87 cannot be unloaded individually at runtime once activated; disabling or removing them requires restarting the entire host process. This is a common affliction of nearly every plugin architecture, and agent evolution is exactly what needs constant self-modification: generate a new tool, install it, find a problem, replace it. Restarting the whole process for every change is a death sentence for “evolution.”

The paper offers two answers.

The first is temporal composability: the moment a plugin is installed, it has already prepared its own way out. Every change it makes to the system is recorded as an “effect” and must be paired with an inverse operation. During installation, the inverse operations stack up in order into an undo chain, like a pile of plates where the last one placed is the first one removed. During uninstallation, the chain is unwound in reverse, and the system returns exactly to its state before installation. Whatever has been done can be undone.

The second is spatial composability: a plugin must declare what it depends on. If a dependency is missing, it waits quietly, neither starting nor raising errors. When a provider appears, it activates automatically. When the provider is removed, it first stops and withdraws its side effects, then waits for the next provider and reconnects automatically when one arrives. What it needs is re-resolved as the environment changes.

One governs time, the other space; together they form the “spatiotemporal composability” of the title. This is not patchwork design but a self-consistent paradigm: in patchwork design, fixing one place makes another collapse, whereas a paradigm follows the same principle everywhere. Nor is it merely theoretical: the mechanism has run in Koishi for four years, supporting more than 4,000 community plugins in production. The name Cordis comes from the Latin word for “heart.”

So “everything is a plugin” is not a slogan but a literal fact: even the agent loop, the core cycle that decides how the model thinks and when it stops, is a plugin. It can be unplugged, swapped, and hot-updated.

Even Anthropic’s Boris Cherny, the person behind Claude Code, has said that the harness will eventually disappear and its capabilities will be absorbed natively by models. But DeepSeek is betting on another path: not making the shell disappear, but bringing it to life and letting the agent take over the harness itself.

2. But most of us are still stuck in old thinking

The most popular approach on the market is still “traditional software + AI”: add an AI assistant to a CRM, a chat interface to an ERP, autocomplete to a development tool. We all believe that the agent is the super entry point. Yet what we build still embeds AI in traditional software. Investors still counting DAU, product managers still scheduling releases, founders still forking a copy for every customization: they are not lazy; they just have an old map.

Why? Because the old thinking has not been cleared away. Let us lay it out.

  • Version numbers: 2.0 fixed three hundred bugs; 3.0 added AI features. A version number is both a tombstone and an ID card. But if software is assembled and constantly changing, the version number has no subject.
  • Installers: dmg, exe, apk; download, double-click, install, update. The installer is physical proof that “software is a finished product,” something that can be moved and installed. But a system composed at runtime is not a finished product; it is a process, and a process cannot be packaged.
  • Feature abstraction: product managers abstract requirements into features, engineers freeze features into code, and this “frozen specification” then gains authority: changing it requires meetings, scheduling, and a release. Abstraction was meant to be a means but became the end. We are no longer serving needs; we are serving an abstracted specification that is already dead.
  • Mechanical customization: a customer wants a custom version? Fork it, modify it, ship a private release. Customization becomes copying, not evolution.
  • Mass copying: the same binary is copied to a hundred million users. The larger the scale, the less anyone dares to touch it, because it is dead, and dead things fear change most.

There is an even subtler misconception: “software is generative.” In the AI era, the thinking goes, nobody needs to write software anymore; the large model will generate it.

That statement confuses two things. AI can generate interfaces, but only because the interface was never the essence of software; the interface is skin. Real software is behavior, state, reasoning, and the compositional relationships running in the background. Generation is only a means of making it; software itself is dynamic. The future is not “software generated by AI,” but software ceasing to exist, leaving only capabilities that are continuously recomposed at runtime. The skin can be generated, but there is no fixed skeleton beneath it.

If these old notions are not cleared away, the so-called “AI products” we build will just be old software in a new skin. The real question should be: should this thing called software still exist?

That is enough dismantling of the old shell. What lies beneath it is waiting for a new way to live.

3. From “traditional software + AI” to “AI + AI-native software”

There is a fundamental difference here:

Traditional software + AI puts AI inside a dead shell.
AI + AI-native software brings the shell itself to life.
DeepSeek Harness may be the first glimmer of dawn for “AI-native software.”

What is AI-native software? It is this living harness, this living agent.

Moreover, future AI-native software will have an essential layer that traditional software lacks: it not only has AI intelligence built in, but can also turn software features into composable parts of its own capabilities. In traditional software, features are dead specifications written into code. In AI-native software, features are capability components of the agent itself: installed when needed, swapped out when done, rearranged and recombined at any time. Software is no longer an add-on to the agent but part of the agent’s body.

Why do I say “the first time”? Because in all the time we have spent building agents, from Claude Code, Manus, Codex, and WorkBuddy until now, we have only gotten half of it right. The half we got right is the agent loop and tool calling: letting the model think, call, and act in a loop. The other half never occurred to us at all: the harness itself can be dynamic, alive, and assemblable.

To use a metaphor of the human body: before, all we could do was increase a person’s knowledge and train their brain. What DeepSeek gives us is the ability to remodel the body itself; nerves, muscles, and bones can all be replaced, rearranged, and reshaped. The skeleton, bones, muscles, and genes are alive for the first time.

Unlocking this part is the true historical significance of DeepSeek Harness.

4. Don’t be misled by the slogan “everything is a plugin”

The DeepSeek Harness website
The DeepSeek Harness website

After DeepSeek Harness came out, people fell into a frenzy of building plugins for it. One detail: on launch night, desktop clients, mobile clients, and private-model integrations all appeared; within four days of the release, 5,488 related plugins were already on GitHub. Humans were more eager to give agents bodies than to install software for themselves.

GitHub Topics · dsh-plugin
GitHub Topics · dsh-plugin

But look at what everyone is installing: most plugins are UI-related. Skins, interactions, status bars: humans are applying their own software aesthetics to put makeup on an agent. Has anyone noticed the problem? DeepSeek Harness can run without any UI at all; it has no need to please humans. It is not an app that needs an interface; it is an agent, and an agent without an interface still reads files, calls tools, and modifies itself.

So what truly drives the evolution of DeepSeek Harness is not UI plugins but plugins that raise its capabilities: making its loop smarter, its tools handier, and its composition freer. In the foreseeable future, we will certainly see large numbers of community plugins shift from putting makeup on agents to transplanting their organs.

But I want to say: don’t be misled by the website’s slogan, “everything is a plugin.”

Pluggability is only a means. The essence is three words:

  • Self-aware: it knows what it is, what it has installed, and what it has changed
  • Reconfigurable: it can reorganize its own structure without breaking itself
  • Growing: it can keep evolving itself while it runs.

And these are fundamental properties carved into its bones from day one. Imagine a species that evolves itself as it lives, in response to its own purposes and goals and to changes in its environment. That is what is truly formidable about silicon-based life: it does not need to be redesigned; it changes itself while it is alive.

5. The plugins you write are the data for the next step toward AGI

More profoundly, the plugins humans write for agents may only be the surface: the plugins of DeepSeek Harness are the data for the next step toward AGI. Behind “everything is a plugin” is another round of “everything is a skill.” To understand this, we have to look back at what the previous round of model iteration brought: Skills.

When were Skills released? On October 16, 2025, Anthropic released Agent Skills along with Claude Code. The form is simple: package a piece of experience about “how to do this,” its steps, conventions, and caveats, into a reusable skill file, put it in the agent’s directory, and the agent loads it and follows it when needed.

After Skills were released, model capabilities took a leap. Before that, a model’s capability ceiling was largely determined by static training data; for processes it had never seen in training, it could only guess. With Skills, a model could “load experience” at runtime for the first time: a debugging process written by a senior engineer, a set of conventions compiled by a domain expert, ready to use as soon as it was opened. The improvement did not come from the model getting smarter, but from the model being able to inherit human work experience directly for the first time.

Who contributed these Skills? Not Anthropic employees, but tens of thousands of engineers around the world. They wrote the pitfalls they had hit, the processes they had validated, and the best practices they had accumulated into Skills one by one and contributed them to the agent ecosystem. Human engineers contributed their experience to the models; that is the real source of the huge leap in model capabilities.

So if we lay out the history of training data, there are three clear stages:

Stage one: the most basic static model training data. Web pages, books, code: human knowledge written down and fed to the model once, then fixed. What the model learns is only a slice of the world’s past.

Stage two: after agents gained Skills. Human engineers contributed Skills, handing their experience directly to models and greatly improving model capabilities.

Stage three: now. People have started contributing plugins to DeepSeek Harness. Plugins go a step further than Skills: Skills are packages of experience, while plugins are the body itself: tools, bones, muscles, genes. As people contribute plugins, models’ harness capabilities (agent harness capabilities) will bring a further improvement in the next stage of model capability. Skills give models new experience; plugins give models new bodies.

From what we have learned about model-agent evolution over the past few years: every time we discover a new domain that can evolve and iterate itself through data, every time we discover data from a new domain, the next round of artificial general intelligence development enters a new era. Skills started one round; plugins are starting the next. That is what excites us most.

In fact, DeepSeek itself is following this roadmap. The motivation chapter of the paper is titled “Self-Evolving Agent Harnesses,” and its conclusion states the goal of letting agents continuously generate and replace their own harness components with almost no human supervision.

There is also a paradox here that must be mentioned.

Engineers have praised DeepSeek Harness unanimously: functional programming, type theory, reversibility, everything engineers love most. But precisely for this reason, I think it may be the first-order form of agents abandoning engineers. People who write software will be reduced to people who design composition rules, and then to people who are told by agents “which plugin to install.”

If agents really grow into that, self-aware, reconfigurable, and growing, where does that leave us?

6. What opportunities remain after models and agents devour everything

Model capabilities will devour every feature that can be copied, and agents will devour every process that can be automated; even “building software” itself will be devoured. So what opportunities remain?

I think there are three. They are the value we can still hold on to.

First, live data: data that keeps changing. The world is full of continuously changing data that will never be trained into models. These are signals from the world that keep changing, and models and agents need to respond to them; they are also the context models and agents require. Here, our role is to be providers and processors of real-time data.

Second, interfaces: real-world capabilities that can be called. Interfaces include the capability interfaces we package for the digital world and the callable capability interfaces of the physical world. They give agents hands to use, so they can reach the services they need and act on reality. Here, our role is to be the packagers of digital-world and physical-world capabilities.

Third, Lego blocks: reusable components. Some functions agents can implement themselves, but implementing them every time they are used is not economical. So if we build reusable infrastructure for agents that agents can use, there is still opportunity here. Our role is to be builders of blocks.

After models and agents devour everything, these are the values we can still hold on to.

Finally, back to the ship

Back to that night of August 13. DeepSeek changed its prices, released an open-source framework, and went to sleep. The community argued for three days: some benchmarked performance, some criticized the design, some counted stars. But few stopped to ask the real question:

Tianyi Cui, who came from quantitative trading, led a team, co-wrote an 88-page paper with a university professor, and turned it into code. What exactly were they building?

They were not building yet another coding assistant. They were building an answer to a question that has troubled the software industry for forty years: does software have to be dead? Can it have no version number, need no installation, and not be fixed at the factory? Can it be alive?

The paper’s answer is in its title: spatiotemporal composability. The product’s answer is in its structure: everything is a plugin. Both answers point to the same future: software ceases to exist, and what exists are systems that continuously recompose themselves.

Across the two worlds, every keyword has changed: identity is no longer a version number but “model + assembly”; form is no longer an installer but a harness, a living body; the core is no longer a monolith but plugins; users are no longer only people, but people and agents together.

Old software New software
Identity Version number Model
Form Installer Harness
Core Monolith Everything is a plugin
Users People People and agents

Has software disappeared? It has, and it has not. It lives on in a new form: silicon-based life. Software is the gene fragments that exist within this silicon-based life. The logic, rules, and algorithms we have written for forty years are not dead; they have become the genes of the agent’s body and the nourishment of its evolution.

And the Ship of Theseus from the beginning now has an answer. If a ship has replaced all its planks, is it still the same ship? If identity is a parts list, the ship is gone. But if identity is the ongoing process of self-reconfiguration itself, the ship is always there; it is simply never the same ship twice.

Software has run on silicon for fifty years, and for fifty years silicon was a dead carrier. Now the carrier itself has come alive.

It is not software; it is living silicon-based life.

Made by DeepSeek Harness with DeepSeek V4 Flash