[]│Developer Philosophy│24 min read

$ Agent-First Development: Build the Capability Before the Interface

Mobile-first taught us to design for the most constrained screen first. Agent-first asks a more fundamental question: what should our software be able to do when the user never opens it? This post looks at assistants that live in messaging apps, at capabilities an agent can discover and verify, and at what changes for marketing, discovery, and small open-source projects when the evaluator is an agent.

Anyone who built software through the early 2010s remembers the moment mobile stopped being the version you squeezed down at the end.

For years the order was fixed. You designed for the desktop, where there was room for everything — navigation, secondary actions, three competing calls to action — and then somebody had to figure out how to fit all of it onto a screen the size of a hand. The results were exactly what you would expect from a design that was never meant to live there.

Mobile-first flipped the order, and the argument was never simply that phones were important. It was that starting with the most constrained interface forced you to understand what mattered. Luke Wroblewski’s original case put it plainly: mobile forces a team to focus on the most important data and actions, because there is no room for anything else. The practice that grew out of it was to start there and then expand into environments with more space.

I think we are approaching a similar shift, and the constraint is different enough that it changes what “interface” even means.

What happens when the user doesn’t need to visit your application?

Not because they stopped needing what it does. Not because a model somehow replaced the service. But because the assistant they already have can discover it, understand its capabilities, and use it on their behalf.

The application still exists. Its capabilities still matter. Its reliability probably matters more than before. But its dedicated user interface is no longer necessarily the front door.

That is the idea I want to poke at: agent-first development.


Agent-first does not mean adding a chatbot

The phrase is already overloaded — I have used it myself to describe development tools built around agents rather than an editor — so let me separate this from a few adjacent ideas first.

I am not talking about using an agent to write your application. That is about the development process, and I have written about it at length; harness-driven development is my name for it. This post is about the way the resulting software is designed to be used.

I am also not talking about putting a chat window inside an existing product. That can be useful, but it usually preserves the exact assumption I want to question: the user must come to your application, learn its boundaries, and talk to your assistant inside them.

The more interesting question is almost the opposite. How does your product become useful to the assistant the user already has?

My working definition:

Agent-first development means designing a product’s core capabilities so that an agent, acting on human intent, can discover, evaluate, use, and verify them without depending on the product’s dedicated graphical interface.

The human remains the person being served. The agent becomes the operator.

Notice that this does not require the product to contain a model at all. A deterministic service that performs one narrow task reliably can be an excellent agent-first product. A product stuffed with AI features can be nearly impossible for any external agent to use. The distinction isn’t how much AI is inside the application. It is whether the application can participate in an agent-mediated workflow.

Instead of beginning with “what screens do we need?”, we begin with “what can someone accomplish through this service, and what must an agent understand to accomplish it correctly?”

The screens become one possible expression of the product. They stop being its definition.

An enormous capability space inside a constrained host

At first the mobile-first analogy looks backward.

Mobile-first starts with less space. An agent appears to offer much more: describe an intent, combine capabilities, coordinate several services, and produce an outcome that would otherwise require a person to navigate five applications. There is no obvious limit to the workflows you can imagine, and that is precisely where the thinking goes wrong.

An imaginable workflow is not an executable one.

The agent still needs usable tools, sufficient context, appropriate permissions, and feedback it can interpret. Anthropic’s guidance on writing tools for agents is blunt about this: good tools need deliberate descriptions, useful responses, and evaluation against real agent behavior — not merely more operations exposed. Everything your backend happens to do is not a capability surface. It is a pile. The important capabilities have to be understandable, bounded, and composable.

Meanwhile, the human-facing side of the interaction may be more constrained than anything we designed for on mobile.

Perhaps the assistant lives in a messaging app. Perhaps the interaction starts with a voice note. Perhaps the user only ever sees a short notification asking them to approve an action. We have gained a potentially unbounded capability space while giving up control over almost all of the presentation.

That is the tension worth sitting with. In mobile design we asked how much would fit inside a known viewport. In agent-first design we have to ask what the host, meaning the app the assistant lives in, lets us communicate, and what the agent can reliably understand.

Can the host show a useful comparison, or only prose? Can the user approve a specific action, or only reply “yes”? Can a long-running task report progress without flooding a conversation? Can an attachment carry the result? What happens when the user changes their mind halfway through?

These are design questions, even though none of them involve drawing a screen. The constraint has moved from pixels to context, attention, permissions, and certainty.

The absence of a traditional interface does not remove the need for interface design. It changes where that design happens.

Instinct makes the shift easier to see

Instinct, a recently launched personal assistant, is a useful example because its positioning makes the interface argument unusually explicit. Its own description is almost aggressively plain: it connects to your applications and devices, it is trained to use a phone and a computer, you can text or call it, and there are no new interfaces. In practice, according to early users and press coverage, that means it lives in iMessage, WhatsApp, or plain SMS, the messaging apps you already open a hundred times a day, and it is invite-only as I write this. I have only looked at it from the outside, so treat this as a reading of its positioning rather than a review.

What interests me is not the list of things such an assistant can do. It is the abstraction it offers.

Once an assistant is just another conversation in iMessage or WhatsApp, the user does not need a separate application experience for every capability behind it. And the choice of host is strategic rather than technical. Telegram could be another host — for a team that already speaks to two messaging platforms, a third is not the hard part. Where users already communicate, what interactions that environment supports, and what operational constraints come with it: that is the decision. In North America, iMessage and WhatsApp cover most of the people such a product wants to reach, so those two are enough to start.

There is a second, quieter difference in how these assistants organize work.

ChatGPT’s projects give you explicit containers: a place to group chats, files, and instructions. That structure is valuable when you want deliberate boundaries around a body of work. But it also asks you to decide, before you ask for help, whether a thought belongs in a new chat, an existing session, or a project. That is a small tax, paid constantly.

The Instinct-style abstraction doesn’t have chats and projects at all. It is one ongoing relationship, and continuity is the assistant’s job rather than yours.

Hermes Agent, from Nous Research, explores another corner of the same space: a messaging gateway, conversations that continue across platforms, persistent memory, and reusable skills. It is a different product and a different operating model — something you run yourself rather than a service — but it makes the same point. An assistant’s capabilities don’t have to be tied to one dedicated interface.

I don’t think these approaches are interchangeable, and I am not predicting which one wins. What they share matters more: each is a different way of mediating between human intent and software. And that creates a question for everyone building the software underneath.

Are you trying to become the assistant, or become something the assistant can use?

Most products should not try to be the assistant

My expectation is that people will not maintain a dozen equally important personal assistants.

They may use different ones in different contexts: one for work, one for personal life. Preferences will differ around privacy, control, convenience, and how much organizing they want to do themselves. But being trusted with someone’s ongoing context is a demanding position to earn, and there will be room for only a few winners.

Instinct is competing for that relationship. So are ChatGPT, Claude, Grok Bot, and the self-hosted harnesses people like me keep building. There will not be many of them.

For most product builders there is a more practical opportunity than competing for the entire relationship: become a capability the user’s chosen assistant can depend on.

A scheduling service does not need to become a personal assistant. A reporting service does not need to own the conversation. A specialized developer tool does not need to become another general-purpose coding environment. It needs to do something useful, expose it clearly, and return a result the assistant can work with.

Drawn out, the agent-first stack looks like this. Most of it isn’t yours.

flowchart TD
  H[Human<br/>intent and approvals]
  S[Host surface<br/>iMessage, WhatsApp, a chat app]
  A[Assistant<br/>ChatGPT, Claude, Instinct, a harness]
  C[Your capability<br/>API, MCP server, skills]
  E[Evidence<br/>stable identifiers, current state, artifacts]
  U[Dedicated UI<br/>optional, where it earns its place]

  H <--> S <--> A --> C --> E
  E --> A
  C -.-> U
  U -.-> H

  classDef strong fill:#0A0A0A,stroke:#00FFFF,color:#E0E0E0
  classDef weak fill:#0A0A0A,stroke:#666,color:#999
  class C,E strong
  class H,S,A,U weak

In plain text: a human expresses intent and approves actions through a host surface such as iMessage, WhatsApp, or a chat app. The assistant behind that surface — ChatGPT, Claude, Instinct, or a harness you run yourself — calls your capability, exposed as an API or an MCP server, often with a skill describing how to use it. Your capability returns evidence: stable identifiers, current state, artifacts. That evidence flows back to the assistant and, through it, to the human. A dedicated UI hangs off the capability as a dotted, optional branch. The two highlighted boxes, capability and evidence, are the only ones you control.

Whether a particular assistant lets you plug in directly is a platform-specific question, and it is changing month to month. Not every assistant has an open extension system, and no single integration works everywhere. But the strategic direction holds regardless: build the capability independently, then provide the appropriate ways for supported assistants and harnesses to reach it.

The service should not have to change its meaning every time the user changes their assistant.

An API is not the finish line

There is an obvious objection here: haven’t we already separated frontends from backends? Haven’t we been building headless services and APIs for fifteen years?

Yes. Agent-first development should build on those ideas, not pretend to have invented them. I made a related prediction earlier this year in The Back Office / Front Office Split: the operational layer of software wants to become agnostic infrastructure, operated by agents, with the interfaces optional. Agent-first is what that prediction looks like from inside the team building one service.

What changes is the consumer we design for, and the completeness of the experience we expose.

A human developer integrating an API can read the documentation, resolve ambiguities, write an adapter, and encode their assumptions into a stable application. In an agent-mediated workflow, some of that interpretation happens while the task is being performed, by a system that has never seen your service before. Unclear semantics stop being an annoyance. They become expensive.

Consider a hypothetical reservation service.

The human-facing product walks someone through choosing a date, reviewing availability, entering details, accepting conditions, and receiving a confirmation. In an agent-first design, those meanings need to exist independently of the sequence of screens.

The agent must be able to distinguish availability from a temporary hold, know when that hold expires, and tell a hold from a confirmed reservation. It must be able to retrieve the price and the cancellation conditions before committing. Afterward, it must be able to verify the reservation rather than infer success from a pleasant sentence. And if a request times out, the retry has to be safe: either the agent sends a client-generated idempotency key with the booking call, so the service recognizes the repeat, or it can look the reservation up by a reference it supplied, because the server’s own identifier never arrived. Otherwise a perfectly reasonable retry creates a duplicate booking.

These are not cosmetic details. They are the interface.

An API provides the underlying operations. An MCP server exposes tools and their schemas to compatible clients. A skill describes how to combine capabilities into a workflow. Those are different responsibilities — the MCP specification defines tool interfaces and structured results, while the Agent Skills format packages procedural instructions and supporting resources — and none of those labels, by itself, makes a product agent-first.

OpenAI’s plugin documentation is a concrete example of the separation. A plugin for ChatGPT and Codex can contain skills, an MCP server, or both, and the docs state that custom UI is not required: use model responses or structured results when they communicate the outcome. The important word is optional. You should be able to establish the product’s usefulness before you need a custom screen to express it.

That does not mean shipping every integration on day one. I would rather have one coherent capability surface and one thoroughly tested integration than five wrappers around ambiguous behavior. The Magic Prompt argued the same thing for the entry point: a small invocation into a workflow that has actually been tested end to end beats an impressive surface that leaves the agent guessing at step four.

Evidence becomes part of the product

One implication deserves more attention than it usually gets: an agent-first product has to make its outcomes verifiable.

There is a difference between a service saying “done” and a service returning enough information to establish what happened.

For the reservation example, I would expect a stable identifier, the confirmed details, the applicable terms, and a way to retrieve the current state. For a generated report, I would expect the actual artifact and a clear account of the data it covers. For a data update, I would expect a way to inspect the resulting state.

Structured results make those checks possible. MCP supports structured tool results and output schemas that clients are expected to validate against. But a response matching a schema is not proof that the underlying task succeeded. That still depends on the service’s behavior and the evidence it exposes. I keep running into this rule from the other side — in the Latch story, “the command succeeded” was never verification — and agent-first design is where a service owes that evidence to its callers.

This changes how I would define a finished feature.

Not just: can a person click through the happy path?

Also: can an unfamiliar agent understand the capability, use it with the right authority, detect failure, and verify the outcome?

I would test that directly. Give an agent a realistic intent without privately explaining how the product works, then watch where the interface leaves it guessing. A feature that only works because its creator knows the undocumented sequence is not finished for this audience.

A button used to be how an action was controlled. In an agent-first product, an action is controlled by intent, backed by data, and settled by evidence.

UI-less should not mean invisible or unaccountable

The natural shorthand for all this is “UI-less,” and the phrase needs a boundary.

Usually we are not eliminating every human interface. We are eliminating the requirement for a dedicated product interface.

The user still needs to understand consequential decisions. They still need ways to approve, interrupt, inspect, and recover. Authentication may need a trusted handoff. Comparing visual work may genuinely need a visual surface. Agent-first should never become a reason to force everything into prose.

For the reservation service, a concise comparison card may be better than a long message. A clear confirmation may be appropriate before an expensive commitment. An operational dashboard may remain useful for resolving an exceptional failure. Those interfaces should exist because there is an actual need, not because we assumed every product had to begin with an application shell.

The integration plumbing can be opaque to the user. The consequences cannot.

I should not need to know which adapter or protocol my assistant used. I should know which provider was selected, what I authorized, what it cost, and whether it succeeded.

The interface can get thinner without accountability getting thinner.

Marketing moves toward evaluation

This architectural shift leads to a commercial one, and I find it at least as interesting.

What happens when the person choosing a product delegates most of the discovery and evaluation to an agent?

Say I ask my assistant to find a service for a specific task, compare the alternatives, and try the most suitable one. In that workflow I am not spending much time on anyone’s landing page. I may never experience the sequence of branding, visual hierarchy, testimonials, and sales calls the vendor designed for me. Instead, I expect the assistant to investigate suitability.

Does the service meet my requirements? What does it cost? Which limitations matter? Can it be tested? What evidence supports the claims? Can the assistant actually integrate it into the workflow I asked for?

We lived through a long era where content marketing was king and discovery was organic search. We are already in the transition where discovery happens through an agent’s web search instead. But I think the bigger shift is in what persuades. Palettes, branding, a beautiful landing page, a quick call with a friendly founder: those work on us because of biases we carry. An agent doing a functional comparison has no reason to be moved by most of that. It works from whatever it can retrieve and read, which is closer to data than to design, and it walks straight past the shortcuts we used to take.

I want to be careful here. This is a hypothesis about the workflow, not a claim that agents are immune to presentation or that branding disappears. Aesthetic preferences can be real requirements. Human relationships carry real information and accountability. An agent trained to weigh vision heavily in its decisions might see what we see. And an assistant should respect a user’s stated brand preferences rather than decide that being “objective” means ignoring them. But for a functional decision, a beautiful landing page should no longer be able to substitute for a missing capability — and the gaps some companies used to bridge with design, messaging, and relationships are going to be harder to bridge.

Content discovery is already halfway through this transition. Research on generative engine optimization studies visibility inside generated answers rather than placement in a conventional list of results. At the same time, Google is explicit that its generative search features are rooted in its core ranking and quality systems, and that the same SEO best practices still apply. This is an evolution of discovery, not a clean replacement of everything that came before.

Agent-first products add a further step. Being mentioned in an answer is not the same as being usable in a workflow.

A product can be easy to discover and impossible for the assistant to evaluate without a sales call. It can have excellent documentation and no safe trial. It can have an integration that takes longer to set up than the task itself.

The opportunity is to shrink the distance between discovery and evidence. Let the agent inspect the relevant capabilities. Give it honest limitations, current information, representative examples, and a bounded way to try the product. Make the result inspectable. Put as little resistance as possible between the agent and a real trial.

The strongest demonstration may not be a polished video of the product working. It may be the user’s own assistant making it work.

Agents have biases too

There is an optimistic version of this argument that I don’t think we should accept uncritically.

It goes like this: humans are swayed by appearances and persuasion; agents evaluate data; therefore agent-mediated markets will reward the objectively best products.

That conclusion does not follow.

Research on LLM-based recommendation has shown that models can be biased by the position of candidates in the prompt, and more recent work treats that position bias as a standing problem rather than a quirk. That doesn’t tell us how every deployed assistant behaves, but it is enough to reject the assumption that machine-generated recommendations are neutral. And we need to distinguish the model from the system around it. Which sources were retrieved? Which candidates were never found? Which tools were available? What did the host permit? Did the assistant actually test the product, or only summarize what the vendor said? Even a careful comparison is bounded by what entered the comparison.

There is an adversarial side too. An assistant consuming external content can encounter attempts to redirect its behavior. Prompt injection is not a metaphor for persuasive marketing. It is a documented and actively benchmarked problem in browser-using agents, and the browser is exactly where an assistant goes to evaluate your competitor.

This is where a genuinely new field opens up.

Marketing has spent a century understanding human attention and decision-making — what affects a choice, how to stand out in a crowded space. Agent-mediated discovery creates another set of selection behaviors to study: when a system notices an offering, treats it as credible, decides to investigate, and ultimately recommends or uses it. We have our biases, and models have theirs. There will be people who study them for the same reasons people studied ours.

Some of that work will improve communication. Some of it will try to exploit weaknesses — to trick the agent into believing a product matches the intent, or to talk its way into a conversation that should have been between the user and a bigger, better-suited player.

We should keep the distinction clear. Making real capabilities easier to discover and verify is product work. Manufacturing evidence, misrepresenting suitability, or trying to override the user’s intent is not a better form of it.

For builders of assistants, this creates a responsibility: preserve the user’s criteria, and distinguish claims from observations. For builders of products, it creates a choice about how to compete. My preference is to make the product easier to evaluate honestly — not merely easier to recommend.

The opening for small and niche projects

One reason I find this shift exciting is what I’ve been noticing in my own searches.

When I ask an agent to investigate a specific engineering problem, it sometimes surfaces a small repository, a recently released project, or a narrow discussion thread that I would never have found through my usual browsing. That is an observation, not evidence that AI search systematically favors small projects. But it changes how I think about discovery.

Discovery for open-source projects used to be hard. GitHub does a wonderful job, and it still wasn’t enough: you needed intense marketing through the dev community to land on a platform or a list that would showcase your work. And on the other side, in my own evaluation habits, a low star count or an unfamiliar name was a perfectly good reason to move on. Popularity was a convenient shortcut for deciding where to spend attention. Incomplete, niche, recently released projects were invisible almost by definition.

An intent-driven investigation asks a different question: does this particular project solve this particular problem?

A broad framework can be well known and poorly suited to a narrow constraint. A small project can address that constraint directly. Popularity doesn’t stop mattering. The ranking research I cited above found LLM rankers biased by popularity as well as position. A precise match simply gets another route into consideration. Google’s description of query fan-out illustrates one mechanism: to develop a response, its AI features may issue multiple related searches across subtopics and data sources. That doesn’t guarantee comprehensive or fair discovery, but it helps explain how a conversation ends up somewhere the user’s first search phrase never would have.

For an open-source maintainer, that makes specificity valuable. Explain the actual problem. Show a reproducible example. State what the project does not support. Make installation, testing, and failure understandable. Give a potential evaluator enough evidence to decide whether the project fits. A small project does not have to present itself as a universal platform to deserve attention.

Discovery is still not endorsement. A relevant repository needs evaluation for maintenance, licensing, security, and suitability, and an assistant should not confuse “I found something unusually specific” with “this is safe to adopt.” What excites me is the possibility of a fairer chance to demonstrate relevance — not skipping scrutiny.

A different starting point for building software

This connects directly to the threads I’ve been pulling on all year.

Harness-driven development asks what an agent needs around it to do useful work: context, tools, instructions, execution boundaries, feedback. Agent-first development looks at the software that environment operates. What must the product expose so a harness can use it without relying on hidden knowledge or a human navigating its screens?

The Magic Prompt explores a related entry point: how a small invocation leads into a complete, tested workflow instead of leaving the user to assemble the missing pieces.

Disposable software raises another possibility. Some interfaces may not need to be permanent at all. A capable service can back a temporary report, a one-off comparison view, or a task-specific application generated for the moment and discarded afterward. The capability stays durable while its presentation changes.

Taken together, these ideas make me less interested in treating the application shell as the natural starting point for every product.

I would rather begin with a human intent, identify the capability that satisfies it, define the authority required to use it, and decide what evidence establishes success. Then make that capability reachable through the assistants and harnesses the intended users actually choose. After that, ask where a dedicated interface adds value.

This is not a claim that every application should lose its UI, or that every decision should be delegated. It is a change in the default question. Instead of asking how to add agent access to the application we already designed, ask what the product would look like if agent access were a primary way of using it.

Build for the outcome, not the visit

The lesson I take from mobile-first is not that software should keep moving toward smaller screens.

It is that a new interaction model exposes assumptions we stopped noticing.

Mobile-first challenged the assumption that we could begin with abundant space and solve constraints later. Agent-first challenges the assumption that the user must enter our product’s interface to receive its value.

We went from a full desktop to a small viewport. Now the surface is getting thinner still, all the way to no UI at all — and the applications underneath must be served through agents and harnesses, with a breadth of integrations that a personal assistant can adopt easily and that the human requesting the service never has to see.

I think that is a substantial opportunity for software engineers. We can build services that participate in workflows without owning the whole experience. We can make capabilities portable across assistants. We can let products prove their usefulness through evidence rather than presentation. But we have to design for it deliberately.

Before drawing the first screen, I would ask one question:

Could someone’s assistant discover this product, use it correctly, and prove that it accomplished what the person wanted — without requiring them to open our application?

If the answer is yes, the interface becomes a choice.

That is where agent-first development begins.

//WAS THIS HELPFUL?