InnerCartography Field Notes · Generative UI
The Interface Layer

The Screen
Was Never
Finished

What generative UI actually is, why it matters now, and where the field is going once chat stops being the only surface agents can touch.

The interface you are looking at was written before you arrived. Every button, every field, every pathway was authored weeks or months before your intent was known. That is the fundamental constraint of software. Generative UI is the attempt to break that constraint at runtime — to let an agent read the situation and assemble the right surface for this moment, this task, this person.

01

The Confusion Is Structural

+

Everyone is calling different things by the same name. When someone says "generative UI," they might mean a component catalog the model selects from, a declarative JSON schema the agent emits, a streaming rendering protocol, or a tool that bundles its own interface. Conflating them produces bad architecture and useless debates.

Spec Layer
A2UI / Open-JSON-UI
What UI the agent wants to show — shape, structure, auditable
Transport Layer
AG-UI
How those instructions move — event-based protocol between agent and app
Tool Layer
MCP Apps
When the tool itself brings UI — dashboards, forms, workflows returned from a call

A2UI answers: what should appear? AG-UI answers: how do agent and app stay in sync? MCP Apps answers: how can a tool bring its own interface?

Once you see the three layers, the ecosystem stops looking like competing standards and starts looking like collaborating layers with different jobs — which is what it actually is.

02

Three Families of Implementation

+

The field is converging on three patterns, each with a different answer to how much freedom the model gets.

Pattern 01
Static Generative
Frontend owns a fixed catalog of trusted components. Agent selects among them or fills them with data — cannot invent arbitrary UI, only assemble known building blocks.
Safest
Pattern 02
Declarative
Agent returns a structured spec — cards, forms, lists, state — and the frontend renders it. Flexible enough to adapt, constrained enough to audit. Territory of A2UI and Open-JSON-UI.
Flexible
Pattern 03
Open-Ended
Agent returns more arbitrary payloads — HTML surfaces, iframe-style fragments. MCP Apps sits here. Maximum flexibility, maximum governance challenge.
Governed Carefully

The question is not how adaptive the interface can be. The question is how adaptive it can be while remaining governable.

03

Why Now

+

Agents now plan, call tools, maintain state, and stream intermediate progress. Once that is true, a plain chat window becomes a bottleneck. The agent has richer internal structure than the surface can express.

There is now enough protocol pressure to standardize the seams. AG-UI formalizes the backend/frontend event seam. A2UI formalizes UI description. MCP Apps formalizes UI-returning tools. When multiple actors standardize adjacent seams simultaneously, the underlying product pattern is real.

Trust concerns are forcing disciplined forms of flexibility. Google's A2UI explicitly emphasizes native rendering without executing arbitrary code. That is not a technical preference — it is a governance stance.

Watch the seams, not the demos. When independent organizations standardize adjacent layers at the same time, the pattern is load-bearing.

04

Who Is Building What

+

CopilotKit is betting on the event stream and synchronization contract between agents and applications. AG-UI is the durable thing they are building toward; generative UI is a consequence of it working.

LangChain starts from agents and orchestration, then extends into the frontend so the user can inspect, intervene, approve, or debug. Their docs emphasize human-in-the-loop, state, forking, and streaming — not decoration.

Google's A2UI is becoming an organizing point by providing an open declarative spec optimized for updatable agent-generated interfaces designed for native rendering.

MCP Apps changes the tool call contract fundamentally. The tool surface is no longer "call tool, get text back" — it is "call tool, get an interface back."

Once a tool can return a useful interface, tool quality is no longer just accuracy or API breadth — it is how well the tool guides the user through decisions.

05

Seven Directions

+
1
From chat-first to task-surface-first
Chat becomes the coordinator. Task-specific surfaces appear when needed. Chat doesn't disappear — it stops being the only container.
2
From monolithic apps to ephemeral surfaces
Smaller, temporary interfaces created for a specific workflow step. Software assembled on demand around context rather than navigated through a fixed menu tree.
3
From frontend logic to negotiated contracts
Frontend and agent co-manage the interaction through a protocol. Neither owns it alone.
4
Trusted catalogs over model freedom
Most serious products will converge on constrained approaches — model chooses from trusted primitives or emits declarative specs mapping to trusted renderers.
5
Human-in-the-loop at risk boundaries
Generative UI will create the right intervention points for humans in enterprise, finance, security, and healthcare workflows.
6
Tool ecosystems compete on UI
Once a tool can return an interface, quality is also about guidance depth. MCP Apps makes this shift explicit.
7
The durable asset moves from screens to context
If interfaces can be generated dynamically, persistent value shifts toward context, memory, policies, component registries, and protocol-compatible state. The protocols don't say this directly. Their architecture implies it.
06

Freedom vs. Control: The Real Equation

+

SaaS is being disrupted. Custom agentic software is possible now at a fraction of the previous cost, and software is becoming disposable — assembled for a context, then dissolved. The deeper question this surfaces is not how agents reshape the interface. It's when they should — and how much.

Freedom = latent space. Control = projection into reality. The interface decides when something stays fluid and when it collapses into form.

Every generative UI system is answering three questions, whether it knows it or not:

1
Where can the user safely explore?
2
Where must the system force clarity?
3
Where does it transition between the two?

The axis is determined by three variables. Cost of being wrong — annoying errors allow freedom, lawsuits demand control. Reversibility — undo-able actions allow freedom, one-way doors need control. Clarity of intent — a user who knows what they want benefits from control; someone exploring needs room to move.

More risk + less reversibility + clearer intent → more control
Less risk + more reversibility + exploration → more freedom
07

How This Plays Out by Domain

+

Different domains land at different points on the freedom/control axis. This is not a preference — it is a function of risk, reversibility, and intent clarity.

💰 Finance / Legal / Payments
Control: Very High
Structured UI only — forms, confirmations, hard constraints. Agents assist; they do not decide. Errors are expensive and irreversible.
Agents here are copilots, not drivers.
🏥 Healthcare / Mental Health
Control: High, Nuanced
A split model: the front (conversation) needs freedom to explore toward diagnosis; the back (decisions, actions) demands control. Exploration helps; premature commitment harms.
Front = freedom. Back = control.
🏢 Enterprise Ops
Control: Medium–High
Workflows and audit trails matter. Inputs are messy. The best pattern: agent proposes, UI forces validation. This is where generative UI actually shines — structured suggestions over rigid forms.
Guided autonomy.
🎨 Creative / Design / Writing
Freedom: Very High
UI should morph constantly. No fixed structure. Exploration is the product — control kills value here. The interface itself is part of the creative material.
Agents are collaborators, not tools.
🧠 Research / Knowledge Work
Freedom → Control
Start messy (exploration), end structured (insight, output). The system should tighten over time. Most tools fail here — they stay either rigid or chaotic, never transitioning.
The system should tighten over time.
🏠 Consumer Apps
Hybrid
Discovery needs freedom — open-ended agent UI. Checkout needs control — strict confirmation flows. Same session, different surfaces. The system shifts dynamically between modes.
"Find me a place" → agent UI. "Confirm payment" → strict UI.

The three-phase model makes this concrete. Every meaningful generative UI interaction moves through exploration, convergence, and commitment — and the interface should shift with it.

Phase 01
Exploration
Vague inputs. Generative UI. Suggestions. High freedom — agent reshapes the surface.
Phase 02
Convergence
Options narrow. Structured UI emerges. Freedom and control in tension.
Phase 03
Commitment
Validation. Confirmation. Execution. Agent steps back — control takes over.

Most tools don't do this. They stay either rigid or chaotic, never transitioning. The ones that learn to shift across phases will define the next generation of agent interfaces.

08

What Generative UI Is Not

+

It is not "LLM writes React." That exists, but it is the least mature and least trustworthy form. The more serious branch is about controlled dynamism: event streams, schemas, renderers, component registries, explicit lifecycle and state handling.

It is also not "chat with prettier widgets." In the stronger form, generative UI changes the division of labor between backend and frontend. The agent decides which control surface is appropriate, streams updates into it, accepts user edits back through it, and continues the task with that structured input.

The interface stops being a destination and starts being a conversation partner — one with structure, memory, and the capacity to reconfigure itself around what the task actually needs.

The Sharpest Definition

Generative UI means interfaces that are no longer fully authored ahead of time, but assembled by agents at runtime — and the real design problem is not how to make interfaces adaptive, but when to let them stay fluid and when to force them into form. Freedom is latent space. Control is projection into reality. Every domain, every workflow, every moment in a session lands somewhere on that axis. The interfaces that learn to shift across it will define what software feels like next.