1c-1800x600-no-logo

Maintainers of well-known open source projects are dealing with a new kind of inbox. Pull requests still arrive, often in higher volume than ever, and many are clean, well-formatted, and plausible on first read. What changed is that one person can now produce a hundred of them in an afternoon, and the maintainer still has to read every one to find the three that are real.

That inversion is the most consequential thing to happen to open source in the last two years, and it has almost nothing to do with code quality. For thirty years, the cost of contributing is what made contribution meaningful. It worked like proof of work: sustained contribution was expensive, so it was a reliable signal of commitment, and projects paid for it with influence. Agents have driven that cost close to zero. The signal is gone, and everything built on top of it is now unstable.

Open-Weight Models Are Not Open Source Software

One clarification first: none of this applies to open-weight models. I do not think “open source” is the right term for them, because they are a different species. A model is foundational intelligence, closer to a production tool than to a program.

Nobody is reproducing a model on the scale of DeepSeek V4 or Kimi K3 at home, no matter how many weights are published. Opening weights does not open the training process, and post-training gets more expensive as models grow. The real function of open weights is democratizing intelligence, not enabling community or governance. Treat a model as a powerful new production tool and the traditional software market still exists. What changes is the supply side, and that is where the consequences show up.

Software Is Not Going Away

Since ChatGPT, people have predicted software without interfaces and without discrete features. I do not think that happens, at least not radically.

I have spent my career on database software, and for query and retrieval, natural language is not as precise or as complete as SQL. Take a warehouse management system: when you need to display and update inventory, a table is the most direct form the information can take. Describing an inventory update in prose is painful and slow.

That is not an intelligence problem. It is an interface problem. Database queries have to be precise and unambiguous, and natural language is inherently ambiguous. No amount of model capability fixes a mismatch between the medium and the intent.

We already see the limits in coding. A prototype is fast to describe in natural language. The complex parts, and any precise change to existing behavior, get cumbersome quickly. There is an old line about software complexity: it never disappears, it only moves. Programming in natural language moves it somewhere you cannot see, which is exactly where software debt accumulates.

So we will still need databases, compilers, operating systems, and business systems. Where ambiguity is acceptable, they can stay ambiguous. Where things have to be exact, interaction design still has to be intuitive and correct. Which is why, even with the cost of producing software collapsing, the best software is still built by strong engineers and domain experts.

When Contribution Is Free, Merge Authority Becomes Scarce

Open source has been, in essence, a very low-cost distribution model. Wide distribution built an ecosystem, contributors accelerated the work, and the combination became a moat. Governance, meanwhile, rested on contribution as proof of commitment.

Now that anyone can generate a large volume of contributions in an afternoon, that measure no longer means what it meant, and allocating the benefits that used to follow from it becomes a real problem. This is why more and more leading projects are closing external contribution channels. I think that is a natural response, not a betrayal of anything.

The side effect is that maintainer power keeps growing while upstream iteration may actually slow down. When supply is effectively infinite, the scarce resource is the decision about what gets merged, and facing a flood of noisy PRs, maintainers get more conservative rather than less.

Gradually these projects come to resemble curation. DHH’s Omarchy is a good representative, and the name is not an accident: it is built on omakase, chef’s choice. Curation shifts community discussion away from technical detail and toward politics. In the past, requests were rejected upstream because they were genuinely infeasible. Today a downstream contributor can implement most of them in a fork over a weekend, so whether upstream merges something is decided less by difficulty and more by the judgment of whoever holds merge rights.

The Silent Fork

The second trend is quieter and worse. Traditional enterprise users are drifting out of public discussion and governance entirely. Maintaining a private fork is now usually cheaper than convincing upstream to merge your requirement, and an agent’s understanding of a popular project is already good enough to support most of what those companies need.

Call it the silent fork: serious production users solving their problems downstream, in private, and never reporting back. Upstream loses real enterprise feedback, which is dangerous for systems software specifically, because production feedback is the only way some classes of bugs are ever found. Over time upstream slows from a lack of signal and an excess of caution, while each fork repeats its own low-quality iteration. Since nearly everyone uses similar models and similar harnesses, those forks tend to produce similar hallucinations, so the duplicated work is not even diverse.

Distribution is going the same way. Software choices used to be made by engineers, and that judgment was often flawed but it was theirs. Increasingly the choice is made by an agent. Ask one to build a small website today and it will reach for a default stack, usually Postgres for the database, because that is what appears most often in its training data. That is statistics, not evaluation, and options that might fit better are less likely to appear as candidates at all. Open source looks vibrant on the surface while distribution quietly centralizes.

Why Open Core Stops Working

Commercially, the simple operations business was the first to go. Open core is going next, and I have never been fond of it anyway. The model was: open source the core, sell the supporting pieces closed, visualization, operations tooling, encryption, permission management, packaged as an enterprise edition.

That no longer holds. A nicer dashboard is trivial for an agent. An operations tool is a day of work. Any peripheral feature with a low implementation barrier can be rebuilt by the customer on demand, so the nice-to-have features that used to create lock-in are losing their stickiness fast.

There is a subtler version: hosted cloud services for open source software. A lot of it is open core wearing a SaaS costume, with the operations layer moved from the customer’s hardware to yours. The problem is that you do not control the underlying resources, and the infrastructure cost to the customer does not change much either way. Unless you have a serious oversubscription strategy or an unusually good discount from your cloud and model providers, you are collecting money from the customer with one hand and passing it to the upstream platform with the other.

So the most dangerous position right now is the thin-feature SaaS in the middle: businesses that manufacture a hosting premium out of deployment hassle, data connectors, forms, and dashboards. What holds up better is storage systems broadly defined, payment and settlement networks, platforms that hold persistent state and enable multi-party collaboration, and any product willing to take responsibility for a business outcome.

What Enterprises Actually Buy Is Responsibility

The real problem with open core was never how hard the dashboard was to build. What makes a large enterprise choose a vendor comes down to one word: responsibility. Who handles it when something breaks. Who responds when there is a new requirement. Whether there is a contract that obligates someone to keep maintaining this software.

A long-term SLA on critical software is something enterprises are genuinely willing to pay for. It is also true that big customers often buy a dashboard so their finance department has something concrete to book, because many finance departments do not accept that an intangible guarantee can be expensive.

None of this is new, and that is the point. Complex software still has a complex production side, and domain expertise is still scarce. The delivery contract is still where the value exchange happens. What disappears is convenience of hosting as a value proposition. I should be direct about my own position here: TiDB has been open source under Apache 2.0 since the first commit, and PingCAP sells a managed service on top of it. These are the constraints my own company operates under, not an observation about someone else’s business.

The past two years created an illusion that traditional software was being systematically overvalued. I started out with the same optimism, and I have come around to thinking that complex systems remain complex systems. Building an industrial-grade operating system or a distributed database is not much easier today than it was ten years ago. The time spent typing code goes down, and anyone who has built one of these knows that typing code was always a small fraction of the work.

This is also why the one-person company and the independent forward-deployed engineer do not hold up in the enterprise market yet. Delivering software is easy now. Expecting one person to carry sustained responsibility for years is not realistic, institutionally or humanly, without a backing guarantee, platform-level verification, and insurance behind them.

Agents are making the small software that nobody needs to be accountable for free. In doing so they are making the complex software that requires an accountable party more expensive, while cutting off the path junior engineers used to take to become that party.

Ownership Is the Right to Modify

So do we still need open source? Yes.

The easy reason first: if your software is not open source, it is unlikely to end up in future models’ training data. The deeper reason is that the industry’s prosperity rests on the democratization of technology, with or without AI. When development was expensive, closed source had real legitimacy. Read Bill Gates’s 1976 open letter to hobbyists and the logic is reasonable for its time, because the production side genuinely was doing enormous labor.

That assumption has changed, and the right actually at stake is ownership. Coding agents solve the ability to modify. They do not solve permission to modify. The difference between open and closed source is whether the user can take over the system in front of them without the vendor’s consent. With closed software, every change has to come from the vendor, and even when they are willing to talk to you, the conversation is bottlenecked by human speed.

Which means the software you bought is not software you own. Ownership still sits with the vendor and you hold a right to use. True ownership is expressed through the right to modify, and open source is what grants it. That right used to be theoretical for most users because exercising it required scarce skill. For the first time, it is becoming real.

There is a social argument too. Even assuming agents can eventually rewrite anything, rewriting is not inheriting. A rewrite does not inherit the tacit knowledge accumulated over years, which in complex systems lives mostly in the experience of a small number of core developers. Picture a world where agents are extremely capable and all software is closed. It would not lack software. It would have an overabundance of low-quality private software alongside extreme poverty of high-quality public software, with every company’s agent reinventing the same wheel and the lessons trapped in private silos forever. Agents would accelerate the production of software while slowing the cumulative progress of software civilization.

What Has to Stay Open

Four categories, in my view, must be open and probably will be:

  • Software expected to stay in use for a long time.
  • Software that many parties jointly depend on, such as protocols and data formats.
  • Software where the consequences of failure are severe.
  • Software carrying historical data, large amounts of state, and a great deal of tacit expert knowledge. This does not mean the data itself has to be open.

Concretely: programming languages, compilers, databases, operating systems, network protocols. Iteration in these areas will speed up. Rewriting them will stay hard, and not only for technical reasons. Too much tacit knowledge is embedded in old codebases and in people’s heads, and rebuilding the social trust around them is not something vibe coding can do.

Three Strategies That Still Work

If open source capability now belongs to society, open source companies stop making money by gatekeeping capability and start making it by running things well and taking responsibility for outcomes. For anything aiming to become a standard or an agent’s default choice, open source is still mandatory. The strategy around it shifts.

  • Open what you want the whole industry to adopt. Protocol specifications, data formats, anything with a shot at becoming shared infrastructure. Promote it with no self-interest attached. Anthropic’s MCP is a good example, and the detail that matters is that it lives at modelcontextprotocol.io rather than on an anthropic.com subdomain.
  • Know whether the value is the software or the process of building it. For some software, the construction process and the know-how embedded in its evals are worth far more than the artifact. Test environments, test data, and unusual harnesses can matter more than the code an agent emits at the end.
  • Monetize what requires ongoing human accountability. Managed operations, continuous updates, large-scale hosting, compliance audits, disaster recovery, SLA guarantees and incident response, insurance and claims.

MCP makes the point. The protocol and its SDK should be fully open. An enterprise-grade hosted MCP service with a real reliability guarantee should be paid. If your business is building custom MCP services, the delivery itself is not worth much, while clients who need you to keep updating and guaranteeing what you built are worth a great deal. The same logic replaces seat-based rent with what I would call bounded delegated authorization: instead of selling an operations platform and leaving the customer to run production, you deliver the operational SLA itself. Pricing that is genuinely hard, because you have to know what the guarantee is worth and whether you can absorb the cost of failure.

One more thing: not every open source project should become a company. The playbook has been to use open distribution to build broad dependency, then commercialize through a license change once dependency is deep enough. The more extreme version builds a public-interest image early and later re-encloses what has already become public infrastructure. That is bad practice, and there is no shortage of recent examples.

The value chain may fragment into five roles: developers who define and build the product, cloud platforms that host and run it, agent platforms that distribute it, certification bodies that verify generated software, and insurers who underwrite its outcomes. The first two exist. Agent platforms are just emerging. Verification and insurance for agent-generated software do not exist at all.

Where I Might Be Wrong

Three places, at least. The first is timing. I have already been wrong in this direction once, having started out convinced that agents were about to compress complex systems work. If capability keeps compounding, the gap between rewriting and inheriting could narrow faster than I expect.

The second is the value chain. Verification and insurance are load-bearing pieces of the structure I just described, and neither exists. They might never materialize, in which case responsibility stays concentrated in a few large vendors and the market looks a lot like it did in 2015.

The third is that I may be overweighting systems software because that is what I build. The argument is strongest for databases, compilers, and operating systems. It is weakest for the large middle of application software, where rewriting may genuinely become cheaper than inheriting.

The principle I keep coming back to is this: capital should be able to earn a return from producing public capability, but not from confiscating it.

Agents made producing software almost free. They did not make being responsible for it any cheaper.


Experience modern data infrastructure firsthand.

Start for Free

Have questions? Let us know how we can help.

Contact Us

TiDB Cloud Dedicated

A fully-managed cloud DBaaS for predictable workloads

TiDB Cloud Starter

A fully-managed cloud DBaaS for auto-scaling workloads