← Back to “You already know how to direct it”
InnerCartography · Technical Companion

The thirteen labs

A day spent on the nuts and bolts of giving an agent a narrow, accountable surface to act through — each lab mapped to a spatially-aware social club inside a building.

This is the under-the-hood version of the shorter note. You don't need it to get the point — but if you want to see how an agent is actually given typed, discoverable, scoped, accountable ways to act, here's the whole arc, one lab at a time.

The throughline: the first labs feel like programming with extra steps. Around the middle, the subject quietly changes from "how do I wire up a tool" to "how do I direct an intelligence I didn't write and can't fully predict." Memory, context, coordination, self-verification — that's the part that's actually a new literacy. Each lab below is one idea; open the ones you're curious about.

the arc  ·  typed results → server → schemas → discovery → auth → memory → context → prompt-compiling → A2A → fleets → closed-API wrapping → ship & prove
01Typed results▸
Return outcomes, not stack traces
Give an agent back a typed result it can branch on — a clear "what happened" — instead of a wall of error text it has to interpret and guess at.
In the buildingA booking returns room booked — or RATE_LIMIT: 3rd booking today. The member's agent knows exactly what happened, and what to do next. No guessing.
02A callable surface (MCP server)▸
Expose what you can do, as verbs
Wrap the system as a set of named actions an agent can actually call — a clean surface, not a pile of internal functions.
In the buildingThe club becomes a set of agent-callable verbs: book_pod, checkout_headset, rsvp_event. Things a visiting agent can invoke without knowing the building's insides.
03Schemas & descriptions▸
Describe verbs sharply enough to pick the right one
An agent chooses an action from its description. Vague descriptions get the wrong action chosen; precise ones make the right call obvious.
In the buildingDescriptions sharp enough that an agent reliably picks reserve_pod, not reserve_whole_floor — the difference between a member getting a desk and a member accidentally booking the building.
04Discovery▸
Publish your capabilities; no manual required
A system announces what it can do, so any agent can learn its capabilities on arrival instead of being hand-configured.
In the buildingThe building publishes a /.well-known card. Any member's agent walks in and learns what's possible here — book, check out, RSVP — without anyone writing it a manual.
05Auth & scopes▸
Deny before you do work; gate on claims, not URLs
Check what someone is allowed to do before doing it — and gate on a verifiable claim, never on whether they happened to find the address.
In the buildingMembers can book; only staff can purge a record or force_return a headset. The door checks the badge before doing the work, not after. This is the first lab that's secretly about security.
06Durable + fresh memory▸
Memory is only useful if it's fresh and trusted
Store facts that persist — but distrust stale ones. A fact that was true yesterday can quietly wreck a decision today if nothing ages it out. This is the heart of spatial memory.
In the buildingOne fact per place: "Pod 3's mic is dead." "The east lounge seats twelve." A freshness guard so a fixed mic stops getting flagged — and a stale "this room is free" never double-books a real evening.
07Memory types▸
Different memories do different jobs
Working memory (now), episodic (what happened), semantic (what's generally true), procedural (how things are done) — each is assembled and trusted differently. And what you say last lands hardest, so order matters.
In the buildingWorking = tonight's mixer in progress. Episodic = what happened at last month's mixer in that room. Semantic = this member always takes the quiet pod. Procedural = how check-in runs. The building recites the open plan last, where attention is reliable.
08Context engineering▸
Control the attention, not just the data
The model isn't the bottleneck — what you put in front of it is. Don't dump everything; assemble the few thousand tokens that actually decide the answer.
In the buildingDon't hand the concierge agent the building's entire history. Isolate the context for this room and this request, prune the stale middle, keep the pinned facts. Attention is the scarce resource.
09Prompt-compiling▸
Stop hand-tuning; optimize against a metric
Treat an agent's instructions like any experiment — compiled from real interactions and refined against what it gets wrong, instead of hand-tuned by superstition.
In the buildingThe concierge's instructions are compiled from real member interactions, and self-correct on the cases it misses — the "is this a booking or a complaint?" reflex sharpening itself over time.
10Agent-to-agent (A2A)▸
Trust is the eighty percent
When two agents talk, the message format is the easy part. Signing the message and budgeting the sender — proving who it's from and capping what it can do — is the actual work.
In the buildingA member's personal agent RSVPs to Friday's social with a signed message and a receive-budget — so one misbehaving bot can't flood the mixer with four hundred fake yeses. Trust, made mechanical.
11Orchestrate a fleet▸
A gate is what makes it a team, not a scramble
Many agents working in parallel only become a team when a quality gate checks each piece before the whole thing proceeds. This is org design, in miniature.
In the buildingEvent prep fans out: one agent on catering, one on AV, one on badges. A gate checks every piece — AV confirmed? badges printed? — before the doors open. Without the gate, it's a scramble; with it, it's a team.
12Wrap a closed API▸
Quarantine the fragile part
Most real systems weren't built for agents. You wrap the messy, brittle thing in clean layers — and quarantine the part most likely to break — so agents can use it safely.
In the buildingThe building's 2009 keycard/booking system has no API? Wrap it in four layers, isolate the fragile piece, and now agents can use it cleanly — without touching the brittle original.
13Ship & prove▸
A green build is not a correct deploy
"It worked on my machine" isn't proof. Prove it live — poll the real state, smoke-test the flow end to end, make retries safe — or it didn't happen. Researchers already live by this.
In the buildingAfter "the event room is set up," don't trust it. Poll the live room state until it actually reflects the change, smoke-test the check-in flow, and make every RSVP idempotent — so a retry never double-charges a member or double-counts a seat.

The whole thing, in one line

A spatially-aware social club is this stack made physical: a building that exposes typed, discoverable, scoped actions, remembers each place freshly, coordinates a fleet of agents to run an event, and proves the room is actually ready before it tells a member “you're in.” Every lab was a rehearsal for that.

The five that matter

If you only keep five

  1. 01
    Return outcomes, not stack traces.
    A typed result an agent can branch on beats a wall of error text every time.
  2. 02
    Be discoverable and scoped.
    Publish what you can do; deny before work; gate on what someone may claim, never on the URL.
  3. 03
    Memory is only useful if it's fresh and ordered.
    Distrust stale facts; assemble by relevance instead of dumping everything and hoping.
  4. 04
    Trust is the eighty percent.
    Between agents, the format is easy. Signing the message and budgeting the sender is the work.
  5. 05
    A green build is not a correct deploy.
    Prove it live — sentinel, smoke test, idempotency — or it didn't happen.
Housekeeping, because it counts

After the thirteen, we decoupled the environment — a dedicated .venv, pinned requirements.txt, and a reusable audit sweep — all green and CVE-clean. The unglamorous part is the part that makes the rest trustworthy.