Models Are Not the End: Agents and the Next Organizational Shift
The agent ecosystem is still waiting for an entry point that dramatically lowers the barrier to use. Models will absorb simple scaffolding, but not complex work—and the next advantage will come from task loops and intelligence flywheels.
Jonathan
Founder
The most visible AI companies today may not be the companies that ultimately define the AI-native era.
That is not a prediction about any company’s next quarter. It is a way of reading industrial history. In Zhang Xiaojun’s interview with Zeng Ming, Zeng describes the development of a general-purpose technology in three stages: infrastructure formation, an explosion of applications, and the arrival of truly native applications. Each transition moves the center of value and rewards a different kind of company.
If that framework is right, most foundation-model competition still belongs to the first stage. Models are turning intelligence into a standardized capability that can be purchased by the token. The next stage belongs to agents that can apply that capability to real tasks. Only after that might we see new entry points and platforms designed around agents themselves.
This produces several application-level judgments that matter more than asking which model is strongest. The agent ecosystem may still be waiting for its low-friction entry point. The current application downturn may be a contrarian window. Models will absorb simple scaffolding, but not every complex domain. And the moat for an application company will not be access to a model; it will be an intelligence flywheel built by repeatedly completing real work.
The change will not stop at the product layer. Once agents can deliver work reliably, the basic unit of a company begins to shift from the job to the task. Organizations must redesign context, permissions, verification, and feedback. Strategy can then move from a plan periodically authored by a few executives to a system that is continuously generated and corrected.
Models are not the end. They are the starting point for everything that comes next.
General-purpose technologies move through three centers of value
Zeng’s framework can be reduced to a simple progression:
Infrastructure matures → Applications proliferate → Native products and platforms emerge
The first stage answers whether a capability exists and can be supplied reliably. Electricity needed generation and grids. The internet needed connectivity and standards. Mobile computing needed smartphones and fast networks. For AI, this layer includes foundation models, compute, inference services, and token supply.
The second stage asks what the capability can do. Once the infrastructure becomes stable and affordable enough, builders take it into different domains. Many applications will not survive, but together they discover demand, product boundaries, and technical gaps. Their feedback also pushes the infrastructure forward.
Only in the third stage do products emerge that are native to the new technology. They do not merely insert a new capability into an old workflow. They rebuild interaction, business models, and organizational relationships around a new mode of production. Short-form video was not television compressed onto a phone screen. Mobile payments were not cash with a digital button added. A native application becomes visible when its product logic can no longer be adequately explained by the previous paradigm.
These stages overlap. Infrastructure, experimental applications, and native products can coexist for years. The framework is valuable not because it can identify an exact date for each transition, but because it reminds builders that different stages solve different problems—and reward different capabilities.
Leadership in one stage does not automatically transfer to the next
Infrastructure companies excel at expanding supply, lowering unit cost, and establishing technical standards. Application companies excel at understanding tasks, designing interactions, and building user relationships. Native platforms must define the rules of an entirely new ecosystem. These capabilities are related, but they are not automatically inherited.
The interview uses AOL, Yahoo, and BlackBerry to illustrate this mismatch. Each occupied a crucial position during a technology’s expansion and reached extraordinary scale. None secured permanent leadership in the following stage. The company that provides internet access does not necessarily create the dominant search or social platform. The company that produces energy rarely becomes the defining car manufacturer.
Zeng therefore compares foundation-model companies to the future cloud providers of AI and offers a deliberately provocative judgment: OpenAI and Anthropic may become essential infrastructure companies without necessarily becoming the final platform winners of the native-application era. The point is not that models are unimportant. It is that once model capability becomes standardized, competition moves to a different set of dimensions.
When several suppliers provide substitutable intelligence, prices become constrained by marginal cost. Applications can move between models, and underlying capability gradually shifts from scarcity to common supply. Value then migrates outside the model: toward whoever defines the task, owns the workflow, accumulates trustworthy execution history, and earns the user’s willingness to delegate increasingly important work.
Valuation, growth, and technical leadership demonstrate an advantage in the first stage. They do not prove that a company already possesses the product and organizational capabilities required by the next one.
The agent ecosystem is still waiting for a low-friction entry point
Zeng uses the early internet as an analogy for the current agent stage. The underlying capability is already exciting, but ordinary users still lack a sufficiently simple and unified way to access it. In the early 1990s, publishing a website and getting online remained an activity for enthusiasts. Once browsers dramatically reduced the barrier, websites, portals, and internet applications began to proliferate.
“Browser” here describes an industrial function. It is not a prediction that the next platform must literally be an “agent browser.” The important question is whether a product or protocol will hide the complexity of models, tools, and permissions so that many more people can directly use agent capabilities. It might take the form of an operating system, a work entry point, a protocol layer, or something that does not yet exist.
Agents face a similar discontinuity today. Models can use tools, write code, and execute multi-step tasks, but different agents do not express capabilities, permissions, state, and results in a common way. Users still need to understand models, prompts, APIs, MCP, workflows, and product boundaries before a system can reliably finish meaningful work.
So the arrival of an agent application boom should not be measured only by demos or funding. Three signals matter more:
- A dramatic reduction in the barrier to use: Can a nontechnical user express intent without assembling models and tools by hand?
- Delegation of important work: Will users entrust agents with tasks that last for hours and affect real system state?
- A standardized capability market: Can a platform discover, compare, and invoke different agents—and take responsibility for the result?
Before these signals appear, many products will resemble early website experiments more than mature native applications. The stage will be noisy, but it is also where the first forms of the next entry point are likely to emerge.
The application downturn may be the contrarian window
The interview is not pessimistic about agent startups. On the contrary, as investors become cautious after watching a generation of products get absorbed by model upgrades, applications may be entering a contrarian phase.
Many first-wave products wrapped temporary model limitations in SaaS. They split a task into fixed steps, chained several prompts, and used workflow code to compensate for model weaknesses. These systems could produce impressive demos quickly, but their moat was built around what the model could not yet do. When the model advanced, the distance between the application and the general-purpose model collapsed.
What the downturn invalidates is not the agent direction. It invalidates a product strategy that freezes temporary limitations into permanent architecture. The real application stage begins when users trust intelligence with serious work and builders choose problems that are complex, valuable, and capable of supporting long-term accumulation.
This gives teams a practical decision rule: stop building thicker scaffolding around simple tasks. Assume model capability will improve quickly, then choose a problem that is ten times more complex and produces a more valuable outcome. If the next model release can replace the entire product, the task boundary is probably too narrow.
An agent’s value is delivered capability, not generated text
The internet reduced the cost of publishing and retrieving information. Agents should reduce the cost of invoking capability.
In that sense, a chatbot is an interface; an agent is a unit of work. The user does not ultimately want another block of text. They want the research completed, the code to pass its tests, the ticket resolved, the customer risk identified, or a cross-system process brought to a verifiable end.
Durable agent products take responsibility for complex tasks that model providers cannot easily cover directly:
- The task crosses several systems or lasts too long for a single conversation.
- Execution depends on proprietary organizational context, not only public knowledge.
- Intermediate actions are constrained by permissions, budgets, compliance, or approvals.
- Completion can be verified through tests, system state, human acceptance, or business metrics.
- Each run produces feedback that improves judgment and execution the next time.
Together, these conditions define a defensible product boundary. A stronger model lets the agent accept larger tasks, but a model upgrade does not automatically solve organizational context, system permissions, or accountability for the final state.
The application opportunity is not a small feature the model happens to lack. It is a task that previously required a team to coordinate—and can now, for the first time, be completed by people and agents working together.
Models will absorb scaffolding, but not every domain
The question “Will models eat applications?” often confuses model capability with product responsibility.
Models will continue to absorb general capabilities. Summarization, rewriting, code completion, simple retrieval, and routine tool use may all become standard model features. Thin wrappers around those capabilities will remain under pressure.
But a general-purpose model cannot independently cover every real task. An enterprise refund agent must do more than understand a request. It has to identify the customer and order, retrieve the current policy, check refund authority, invoke the payment system, handle failures, and leave an audit trail. An engineering agent must do more than generate code. It has to understand repository constraints, run tests, observe deployment results, and recover after failure.
The value comes not from knowing slightly more than the model, but from bringing general intelligence into a specific environment and taking responsibility for the final state. As models improve, these applications can accept more complex tasks. Their market boundary may expand rather than disappear.
Electricity did not consume the appliance industry, and cloud computing did not consume SaaS. Infrastructure standardizes a capability while reducing the cost of creating applications on top of it. Application teams should not fear stronger models. They should avoid building their value on temporary model defects.
Application moats will evolve from data flywheels to intelligence flywheels
Traditional internet products emphasized data flywheels: more users generated more data, more data improved recommendations or matching, and a better product attracted more users.
Agent products need an intelligence flywheel. Each execution produces not only behavioral data but also a goal, context, an action trajectory, tool results, human corrections, and a final outcome. That evidence can directly improve the next run by changing context selection, tool boundaries, evals, exception handling, or the division of labor between people and agents.
More real tasks
↓
More execution traces and outcome evidence
↓
Better context, tools, evals, and workflows
↓
Higher success rates and greater trust
↓
More important and more complex real tasks
This is why an application company can build a technical moat without training a general-purpose foundation model. Models can be replaced. The task system’s accumulated judgment criteria, execution feedback, and customer trust cannot be copied by switching an API endpoint.
A strong agent company does not merely possess more data. It turns each delivery into a system that can act more reliably next time.
New platforms must match capability with demand
The previous generation of internet platforms matched information with demand. Search engines ranked relevant pages, marketplaces connected products with buyers, and social networks organized the distribution of people and content.
An agent platform faces a harder object. It must not only discover an agent but determine whether that agent can complete the task, access the necessary resources, operate safely under current permissions, and remain accountable when it fails. When several agents collaborate, the platform must also handle decomposition, state handoffs, conflicting results, and cost control.
This helps explain the interview’s claim that the next platform is unlikely to come automatically from the old one. Incumbent platforms own traffic and relationship graphs, but they are also constrained by existing information structures, business models, and organizational metrics. Turning a directory of websites into a directory of agents does not create a trustworthy capability market.
An agent-era entry point needs at least three infrastructure layers:
- Capability description: The system can understand what an agent does well, what input it needs, and where its boundaries lie.
- Invocation and collaboration standards: Agents, tools, and data sources can exchange tasks, state, and results.
- Trustworthy judgment: The platform uses execution history, evals, permissions, and real outcomes to decide which capability should be invoked.
A future platform may accept an intent and then coordinate several agents from the market to produce the result. Users will no longer select every tool themselves; they will delegate part of that judgment to the entry point. The platform position belongs to whoever can establish that trust.
This means the competition for the entry point will not be won by distribution alone. The platform must first develop judgment: which agent can complete which task under which conditions, whether a success can be repeated, and where responsibility lies when execution fails. Without evals, execution evidence, and a permission system, an agent store is only a directory of unverified claims.
As agents enter production, tasks begin to replace jobs
Most modern companies are organized around jobs. They hire for a position, define reporting lines, assign stable responsibilities, and evaluate performance on a cycle. Hierarchy solved the industrial problem of coordination at scale, but it assumes that capability primarily belongs to people and that information and commands must travel through managers.
Agents weaken both assumptions.
Some capabilities can now be invoked as software services. Company state can increasingly be read by systems without passing through layers of manual reporting. When tasks can be assigned dynamically to the most suitable person or agent, the basic organizational unit begins to shift from “Who occupies this job?” to “Who is responsible for this outcome?”
When people say “the company will disappear,” the more precise interpretation is not that the legal entity vanishes. It is that the traditional hierarchy loses its monopoly on coordination. Future organizations may look more like laboratories, sports teams, or partner networks. Participants assemble around tasks, contribution becomes legible through delivery records, and relationships change with the problem instead of remaining fixed on an org chart.
Management does not automatically disappear. Task-based organizations require a stronger management system. A job description can absorb ambiguity through departments and reporting lines. A task system must make the goal, inputs, permissions, completion criteria, and responsibility boundary explicit. Without them, dynamic coordination becomes continuous firefighting.
An AI-native organization is therefore not simply a company with fewer people or fewer middle managers. It converts coordination mechanisms that once lived inside managerial experience into an executable, observable, and improvable task system.
The harness is the operating foundation of a task-based organization
Models make capability callable. A harness makes that capability reliable inside an organization.
For an individual agent, the harness manages context, tools, permissions, state, verification, and recovery. At the organizational level, it assumes work previously distributed across job descriptions, operating procedures, managers, and quality systems.
A production task must answer at least six questions:
- What is the goal? Which external state must change, rather than which content must be generated?
- Where does context come from? Which facts are trustworthy, and which information is still current?
- Which actions are allowed? Which tools can the agent invoke, and where are the permission boundaries?
- Who is accountable? Which steps can run automatically, and which decisions must be escalated to a person?
- How is completion verified? Through tests, database state, human acceptance, or business metrics?
- What happens after failure? Can the task retry, resume, roll back, or stop safely?
This system preserves continuity and accountability as an organization moves from fixed jobs toward dynamic tasks. Context lets participants understand the situation. Tools provide the ability to act. Evals determine whether the result is acceptable. Execution traces reveal where the process failed.
The model determines how intelligent a single decision might be. The harness determines whether the work can be delivered repeatedly. The enduring advantage of an organization may come not from which model it calls, but from the quality of its task definitions, verifiable workflows, and real feedback.
Strategy will become a continuously generated system
When technology and competition change month by month, static strategy produced in annual meetings becomes increasingly inadequate. Zeng uses the idea of “strategy generation” to argue that a great direction cannot be fully planned. It must emerge through continuous interaction between the organization and reality.
The interview offers a practical cadence: look ten years ahead, think three years ahead, and execute for one year. The cycle can become shorter, but all three horizons remain necessary. A long-term view identifies a problem worth pursuing. Medium-term thinking defines nonlinear milestones. Short-term execution produces real feedback.
A team that looks only at the short term will extrapolate the future from current ARR or feature growth and miss changes in the technological track or the size of the market. A team that looks only at the long term turns mission into a slogan that cannot guide action. Creativity happens in the middle horizon, where the team must explain which steps can connect the present to a long-term direction when the current consensus offers no path.
AI makes it possible for strategy generation to become a continuously operating system. Agents can read customer feedback, product data, engineering state, and decision records to identify where assumptions diverge from reality. The team still chooses goals, explains tradeoffs, and decides which anomalies justify changing direction.
This does not hand strategy to the model. Models are good at processing large volumes of information and finding relationships. People must still formulate the problem, own the choice, and decide what is worth pursuing over time. The advantage of an AI-native organization is that both can participate in the same feedback loop:
Long-term direction
↓
Medium-term hypotheses → Task portfolio → Execution evidence → Evals
↑ ↓
└──────────── Context updates and strategic correction ──┘
Traditional organizations compress information through layers of reporting before a small group periodically revises the plan. A new organization can return execution evidence directly to strategic context. Strategy becomes not a document but a mechanism for continuous learning.
The next advantage is how efficiently an organization converts intelligence into action
Industrial history cannot tell us which company will win, and it cannot prove that every technology cycle repeats mechanically. Its real value is a discipline of observation: do not mistake the leading indicators of the current stage for the final score of the entire era.
As model supply matures, differentiation moves beyond the model. The teams that find important tasks, organize high-quality context, build reliable harnesses, and use evals to turn execution outcomes into evidence for the next decision are the teams most likely to accumulate a compounding advantage.
Agent application teams can reduce the judgment to four questions:
- Does the product produce an answer, or complete a task that changes external state?
- As models improve, does the product boundary disappear—or expand into more complex work?
- Does every execution leave verifiable evidence that improves the next execution?
- Are users willing to entrust the system with progressively more important work and greater authority?
Together, the answers determine whether a product can grow from a feature into a capability, and from a capability into an entry point.
From this perspective, an agent is not another feature added to traditional software, and task-based organization is not a new label for flat management. They point to a deeper reconstruction. Companies will no longer preserve capability mainly through fixed jobs; they will continuously compose the capabilities of people and agents through systems. Strategy will no longer depend only on periodic planning by a few leaders; it will be generated from real work.
Models provide general intelligence. What becomes scarce next is the ability to convert intelligence into reliable action—and to make every action improve the organization that follows.
Related reading
- Why AI agents need a harness, not just a better model — How context, tools, constraints, verification, and correction determine reliability beyond the model.
- AI-native companies reorganize context — How organizations turn fragmented information into context an agent can use for the task at hand.
- Building an AI-native company — How closed loops, queryable companies, and new organizational roles change the way companies operate.
- Build an agent-readable environment — How explicit state and observable execution make an environment legible to agents.
Sources
- A conversation with Zeng Ming on industrial history, companies that disappear, and why excellence is not greatness — Video interview by Zhang Xiaojun with Zeng Ming, September 3, 2026. (Chinese)