Blog
June 8, 2026

AI Agents Have a Monoculture Problem

In June 2024, in From AI Bubble to AI Superstructures, we argued that the bubble debate was aimed at the wrong target. Whether any single large language model was overhyped mattered less than what would get built once many models were orchestrated into systems none of them could match alone. That is where we said the value would form. Two years later, the industry has moved decisively in that direction. Orchestration became the industry’s center of gravity, with its own tooling, engineering discipline, and economics. The next leap now looks less like continuing incremental improvement in yet larger language models and more like heterogeneous model families plus a true conductor layer.

Agentic platforms went from demo to default in the AI application layer. Coding agents, computer use agents, research agents, support agents, data analysis agents, and a lengthening list of others are now built from harnesses, loops, tool use, skills, memory, scheduled runs, and handoffs between models. All of that machinery is orchestration: the layer that coordinates many models and tools to do a job none could do alone, routing the work, holding the context, checking the outputs, and recovering when something fails. There is interesting engineering all over AI, including in the frontier models themselves, but it is this layer that is producing the most powerful applications in use today. The model became a component, and the assembly became the product.

This is the fiber story from the first piece playing out on schedule. The fiber overbuilt in the dot com years looked like ruinous excess until streaming and the cloud arrived to consume it. The compute being overbuilt now may be consumed the same way, by an agentic layer that barely existed as a commercial category when the clusters were financed.

As impressive as agentic AI has become, it is at most halfway to what we described, and probably much less. The orchestration arrived; the orchestra did not. Today’s agentic systems are overwhelmingly language models orchestrating other language models, with text as the glue. The stack is, for the most part, talking to itself. The conducting got remarkably good. The instruments barely changed.

The language models did get more varied and multimodal along the way. They see and create images, they write code, they use tools, and they increasingly differ by cost, speed, context length, reasoning depth, and specialization. A well designed agentic system sends small fast models to do the muscle work, reserves the costliest frontier models for the deepest reasoning, and uses models in between based on their strengths. That is real engineering and real economics, but it is variation within a single species. Today’s ensemble differs mainly in size, speed, and training emphasis, not in kind. It is a monoculture, and not only in the loose sense. Agents built on the same family share the same blind spots, so they can fail together while appearing to check one another. That can mean correlated hallucination, shared benchmark blind spots, or one agent “verifying” another while relying on the same underlying reasoning style. That is redundancy without independence. It is like building a house from one material. You can do it, and it may even stand, but a real house works because concrete, glass, copper, steel, insulation, wiring, lumber, and stone each do something the others cannot.

That is a long way from the system we described in 2024: heterogeneous compositions of frontier reasoning models, small specialized models, world models that predict how physical environments behave, vision and audio models, models that read instruments and control processes, models that design molecules and forecast demand, each contributing what it is genuinely best at. We are not there because the ensemble is too uniform. Language became the universal interface between models, and nothing comparable exists yet between unlike families. MCP and A2A are real progress, but their center of gravity is still LLM-centric: tools, resources, messages, agent cards, tasks, and handoffs. They can carry structured and multimodal payloads, but they do not yet give unlike model families a common way to exchange rich latent state, uncertainty, physical predictions, control constraints, or domain-native representations.

The missing instruments are no longer hypothetical. Over the past year world models moved from research papers into commercial developer platforms for robotics, autonomous vehicles, and simulation, yet they live in parallel verticals, rarely wired into the same loops as the agents writing code and running research. Specialized models for forecasting, optimization, control, science, simulation, and other domains already exist in varying degrees of sophistication, but they are not yet in the orchestra.

What changes this is not a standard alone, but a change in the whole engineering pattern: more heterogeneous model families, more engineers who know how to compose them, and standards that let unlike systems pass more than text back and forth. The shipping container did not make ships faster; it made every ship, truck, and crane interchangeable parts of one freight system, and the cost of moving goods collapsed. Interfaces that let any of these hand their output to a reasoning model as naturally as text now passes between chatbots would do the same for AI. That is what we watch for most closely, along with a variety of model types crossing into general agentic loops and agents reaching the physical world through sensing and action.

A two panel diagram titled From Monoculture to Orchestra. The left panel shows today's agentic stack: a single orchestrator coordinating a row of similar language models into a useful result, with shared blind spots, correlated hallucination, and verification without independence noted alongside. The right panel shows the next leap: a purpose-trained conductor model routing among heterogeneous model families including reasoning, vision, forecasting, optimization, simulation, and scientific models, with additional models discovered through a standard protocol, producing a dynamic adaptive system output.
From monoculture to orchestra. Today's stack coordinates one family of models; the next leap routes a purpose-trained conductor across heterogeneous model families, each contributing what it does best.

One more thought, offered as conjecture. The language model is unlikely to be the highest form of AI model we ever build. Something different may simply be discovered, in the sense mathematicians use the word, already there and waiting. Or the language model may evolve under its own pressures: models that spend most of their time talking to other models have little reason to keep using human language between themselves and every reason to compress it, and a protocol many times denser than our language would eventually produce systems that no longer think in anything we would recognize as language. If models begin training their own successors, how fast would recursive self-improvement get there?

The 2024 question, whether AI was overhyped, now sounds antique. The 2026 question is what it takes to assemble the rest of the orchestra, and we think the answer has two parts. The other families of models need the kind of sustained investment language models have enjoyed, so the agentic whole has more to draw on. And orchestration deserves a model of its own, trained from the beginning to conduct, to know which player to call and when to overrule it, rather than a language model repurposed as the leader of a pack of its own kind.

The human brain, the only general intelligence anyone has met, is itself a composition of specialized, semi autonomous parts working under orchestration. The path to broader machine intelligence may run through composition rather than scale. We expect purpose-trained orchestration models to emerge soon and would be surprised if the frontier labs were not already building them. The next leap is the orchestration of heterogeneous model families that cannot work together without it. The opportunity belongs to the platforms that let people combine those families into workflows no single model, and no monoculture of models, could produce alone.