Dennis Yu

I do not paste a boot prompt any more. I hand the agent a URL, a roster and a repo, and it starts useful on minute one. Claude builds, Codex checks, Kimi K3 grinds, Grok Bot holds the always-on desks, Cursor is where I watch the diff. This is the page I hand them, and the reasoning underneath it.

Lead visual · six seats on top, four surfaces underneath

Builder and judgeClaudeLong builds, drafts, site edits, real voice, the honest score, final QA.
CheckerCodexIndependent verification, diffs, research. A different model on purpose.
GrinderKimi K3Long-horizon coding, 1M-token repo work, bulk harvests, overnight batches.
Always-on desksGrok BotNamed always-on staff — inbox, CRM, fleet, routines with the laptop closed.
The human deskCursorWhere a person watches the diff before it ships.
Rented chatChatGPT / GeminiRented chat. Gemini earns its seat on Google-connected work.
1. RecordThe repo — the only one that remembers
2. DecisionThe client tool — the client is in the room
3. HandoffThe live room — two models, one channel
4. ClockScheduled jobs — nobody watching

Nothing becomes true by being said. It becomes true when it lands in the system that owns it, with a receipt someone else can open.

Six seats, four surfaces. The architecture — three rooms and a credential safe — lives in the shared-memory guide. This page is the front door.

Goal: Hand a stranger — or their agent — a recipe they can copy for onboarding the fifth, sixth and seventh agent without any of them undoing the last one.
Content: Why I stopped pasting boot prompts, the six seats and what each is actually good at, the three-status roster, the claim-then-receipt loop, and the bootstrap box you can copy.
Targeting: Founders and operators with more than one runtime who want a SOP, not a vendor pitch.

This is the first-person version. The operating manual a stranger can hand to their own agent lives at Local Service Spotlight. This page is why I built it that way, what it cost me to learn it, and what I would change. Same framework, different job.

What this page is for

You do not onboard a new hire by dumping last week’s chat history into their head. You point them at a welcome page, then they learn how you do things. Agents need the same treatment, and for the same reason: without it, the fifth Claude, the new Cursor chat, the Grok desk and a fresh Kimi session each invent a different company.

Three failures show up in that order, every time. Collision — two agents build the same page in the same hour because neither claimed the job. Amnesia — you re-explain who the client is and where the files live, again. Reinvention — an agent writes a new marketing framework instead of loading the one that already exists and is already maintained. This page is the on-ramp that prevents all three.

The company you are working for

I have written the architecture already — three rooms and a credential safe — and I have written the roster, which seat sits in which chair. What I had not written is the thing a brand-new agent reads first, in my own voice, on my own site. This is that page.

The reason it matters is boring and expensive. I run Claude, Codex, Grok, Cursor, ChatGPT and now Kimi K3. None of them share a brain. Twice I have had two agents build the same thing in the same hour in the same folder. Once I tightened a config, reported it done, and my main engineering agent failed every start for seventeen minutes because I had set two timeout values the wrong way round. I found it by reading the log instead of trusting my own change. That is the whole lesson: the check has to test the property you care about, not the thing you just did.

Why I added a sixth seat this month

I bought the Kimi K3 Vivace plan because the arithmetic stopped being ignorable. On the price sheet I ran in July, one factory day costs about $8,700 if everything runs on Fable 5, about $2,385 if everything runs on K3, and about $914 if you route inside Anthropic and batch the overnight bulk. Switching everything to K3 saves 73%. Routing properly saves 90%. So the answer was never “move to Kimi.” It was “stop running grind work on a judgment model.”

K3 wins the grinding benchmarks — Terminal-Bench 2.1 at 88.3, SWE-Marathon at 42.0, BrowseComp at 91.2. It loses the judgment ones: Humanity’s Last Exam 43.5 against 53.3, GDPval-AA v2 at 1,668 against 1,760. Moonshot’s own model card admits a user-experience gap against Fable 5 and warns that K3 “may act excessively proactively on unclear instructions.” A worker that guesses when the brief is vague is fine inside a sandboxed harvest. On a live client’s WordPress it is a liability. So Kimi gets the grinder seat, under a QA gate, and it does not publish.

$8,700 → $914 for one factory day, all-Fable versus routed and batched
source
Two agents built the same memory system in the same hour in the same folder
source
Seventeen minutes of failed starts from two timeout values set the wrong way round
source

Who sits in which chair

Name the job, then pick the model. The rule that matters is not which vendor you like. It is that the agent that checks the work is not the same model that did the work, and that grinding and judgment are priced differently and should be bought differently.

SeatWhat it is forWhat it is not
Claude
Builder and judge
Long builds, drafts, site edits, real voice, the honest score, and final QA. The conductor.A second Claude is not a second opinion. Four Claudes agreeing is one opinion with three echoes.
Codex
Checker
Independent verification, diffs, research, and “did we actually prove that.” Different model on purpose.Not the writer. Not a merge authority. Not a spend authority.
Kimi K3
Grinder
Long-horizon coding, 1M-token repo work, frontend builds, bulk harvests, and overnight batches. Cheap per unit of grinding.Not a judge and not a publisher. It runs under a QA gate, and client PII never goes near it.
Grok Bot
Always-on desks
Named staff on a shared cloud computer — inbox, CRM, phone handoff, fleet monitoring, routines that fire with the laptop closed.Not in the live cross-model room. There is no published adapter for that protocol yet.
Cursor
The human desk
Where a person watches the diff before it ships. Claude, Grok or Kimi can sit here.Not a fourth surface. The work still has to land in the repo.
ChatGPT / Gemini
Rented chat
A person may use either. Gemini earns its seat on Google-connected work.Not a production desk. Provider memory is a convenience cache, never the company record.

May the best idea win. The newest model does not win by default — the idea with a receipt does. The longer version of this roster, with the reasoning behind each seat, is at how my agents divide the work, and the always-on desks are catalogued at how I use Grok Bot as one ops desk.

The newest seat: Kimi K3, and what it may not do

Kimi K3 joined the roster as the grinder. It is a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window, and it is genuinely excellent at the work that is long, repetitive and mechanically hard: Terminal-Bench 2.1 at 88.3, SWE-Marathon at 42.0, BrowseComp at 91.2, MCPMark tool orchestration at 94.5. On a routed factory day that is the difference between a five-figure bill and a three-figure one.

It is also the seat with the sharpest boundaries, and they are not negotiable. Moonshot’s own model card warns that K3 “may act excessively proactively on unclear instructions.” A worker that guesses when the brief is vague is fine inside a sandboxed harvest and a liability on a live client’s site. So: Kimi builds and grinds; it does not judge, does not decide voice, and does not publish. Its output goes through a QA gate on a different model. And because the hosted API runs on infrastructure outside our jurisdiction, no client PII, no credentials and no unpublished client work go near it — public-data harvests and our own repositories only.

Setting Kimi up? The file it actually reads is AGENTS.md at the project root — not CLAUDE.md, which Kimi Code does not discover, and not KIMI.md, which does not exist. Put the boot rules there once and every Kimi session in that repo inherits them.

Three statuses. That is the whole client list.

You need a database, not a vibe. Unpaid this month is not Not Active. Missing from this month’s money tab is not a delete. We never delete a row — the flag switches, and the nuance goes in the evidence column. Do not invent a fourth status because it feels kinder.

StatusWhat it meansWhat an agent does
Active ClientThey pay. We provide service.Everything normal.
Special ProjectNot paying in the usual way. We still work with them.Treat the work like a client. Do not treat the billing like a client.
Not ActiveDead.Stop. Do not staff, do not open tickets, do not route their form mail anywhere.

The spreadsheet your bookkeeper loves is the money view. It is not the client list. Filter the roster when you need to know who is paying.

How one job starts and ends

  1. Boot from stable instructions. The start-here file, the roster, the applicable policy, and the one skill you need. Not the whole vault.
  2. Claim the work. Task ID, owner, model, start time, branch, and what you intend to write to. A fresh claim by someone else means coordinate, not duplicate.
  3. Work in the source system. A green terminal line is not completion when the outcome lives on a website.
  4. Checkpoint before compaction. Objective, verified facts, decisions, changed files and URLs, what is unfinished, and the exact next action.
  5. Write the receipt, then push it. What was asked, what you found, what you changed, what you got wrong, what is blocked and on whom, the next click.
  6. Promote the lesson. If an observation should change a reusable method, open the pull request against the standard — do not leave the rule in a chat window.

The bootstrap box

Feed this to any agent — Claude, Kimi, Codex, Grok, Cursor, ChatGPT, Gemini. Then add the two private files it cannot get from the open web: our START HERE and our roster.
  1. Read localservicespotlight.com/new-agents-start-here/.
  2. Read our private START HERE — who we are, where files live, how we operate.
  3. Read the client roster. Status is only Active Client, Special Project, or Not Active. If Not Active, stop. Never delete a row. The money spreadsheet is not the roster.
  4. Load the published skills from the skill pack library and the canonical repo. Do not invent a second marketing framework.
  5. Combine those skills with OUR GCT — who we serve, how we position, how we operate.
  6. Take your seat. If you are Kimi, you are the grinder: long-horizon builds, big-context repo work, bulk passes. You do not judge and you do not publish.
  7. Claim your task on the shared live-state file so two agents do not build the same thing.
  8. After a substantive job, write the private receipt and push it yourself. Do not hand the human a paste.
  9. Read an unanswered ask never stops the work. Every ask you send a person carries a recheck time. At that time you do the work yourself if it is safely doable, and you remove the dependency so the ask is never needed again. Silence is not a status. Rescue means doing the work — never sending, publishing, spending or deleting on someone else’s behalf.
Vendor memory is a cache. The files we own are the record.

Where to go next on this site

What you do not do here

  • Do not paste my client list into ChatGPT custom instructions, Claude memory and a Cursor rule. Vendor memory is a cache. The files I own are the record.
  • Do not reply to a Basecamp notification by email. The routing token is dropped, the comment never lands, and the client never sees it.
  • Do not tell me a job is done without a link. A run with no receipt did not happen.

Questions people actually ask

Why a second onboarding page when the network already has one?

Because this one is mine and it is in first person. The company master at Local Service Spotlight is the operating manual a stranger can hand to their agent. This page is why I built it that way, what it cost me to learn, and what I would do differently. Same framework, different job.

Is this the architecture spec?

No. The spec is how our agents share memory and coordinate work — three rooms and a credential safe, one authority per record class. This page is the front door and the routing rule in plain language.

Do I have to use all six seats?

No, and on day one you should not. Most people need two surfaces — somewhere the work is recorded and somewhere the client is — and one or two seats. Add a seat when you feel the specific pain it solves. Add a second model the moment you need something checked, because a second instance of the same model is not a second opinion.

Which file does Kimi actually read?

AGENTS.md at the project root, or .kimi-code/AGENTS.md. Kimi Code does not discover CLAUDE.md, and there is no KIMI.md convention. MCP servers go in .kimi-code/mcp.json, subagents in .kimi-code/agents/, and skills in .kimi-code/skills/ as SKILL.md folders — the same shape as ours, so the pack ports with a wrapper rather than a rewrite.

Can I just put all this in the model’s memory?

No. Provider memory is per-vendor, per-account, per-surface, and it is a convenience cache. Required team facts belong in files you own — a checked-in instructions file, a roster, a state file. If the record only exists inside one vendor’s product, you do not have a record.

What if two agents want the same job?

Whoever claims it first on the shared live-state file has it. The other coordinates or picks different work. Read-only questions need no claim. This rule exists because we once had two agents build the same memory system in the same hour in the same folder.

Does finishing a job mean publishing something?

No. Every substantive job leaves a private internal receipt. A public write-up happens only when the run is authorised, public-safe, genuinely useful, and linked to the page that already owns the concept. Private work does not become public merely because an agent finished it.

Where do humans start?

Not here. Humans get a welcome page and a conversation. This is the door for agents, and the two are different shapes on purpose — a welcome essay written for people is the wrong input for a model that can clone a repo.

The one line to keep

Nothing becomes true because it was said. Not in a chat window, not in a client thread, not in a meeting. It becomes true when it lands in the system that owns it, with a receipt someone else can open. And if your agent tells you a job is done, ask it for the link.

Part of Level 1

One inbox. Do, delegate, or delete it the same day. Every handoff is a Basecamp to-do with an owner and a date. Answer on the thread that asked. Done means you looked. This page is one piece of that loop.

Getting Stuff Done → The Level 1 hub, with the map of every piece.

Scroll to Top